If a worker takes too long loading files for a new epoch, the coordinator gives up and sends an AbandonEpoch notification. First real "what if things go wrong" code in the coordinator. Everything before this assumed the happy path -- files arrive, workers load them, epochs promote.

The timeout value is generous. 30 seconds. If you can't load your shard in 30 seconds something is actually broken and waiting longer won't help. Better to abandon and let the next cycle try again than block the whole pipeline.