Distributed batch¶
yt-uniq queue + yt-uniq worker let multiple machines share a single
work queue without any external coordinator. Coordination is just
os.rename(2) between two directories on a shared filesystem.
Filesystem requirements¶
The contract: cross-directory rename must be atomic — exactly one worker wins each lease. Compatible setups:
| FS / mount | Status |
|---|---|
NFSv4 with noac mount option |
supported |
ZFS shared via NFSv4 (sharenfs=on) |
supported |
| ext4 on a single host (multi-process on one machine) | supported |
| APFS on macOS (single-host testing) | supported |
NFSv3 without noac |
NOT supported — attribute caching makes rename appear non-atomic |
s3fs-fuse, goofys (S3 over FUSE) |
NOT supported — rename is implemented as copy+delete |
| SMB1 | NOT supported |
| SMB2/3 | partial; verify with init step |
yt-uniq queue init performs a real atomic-rename test against the chosen
directory and fails fast with a clear error if the filesystem can't honour
the contract.
NFS mount options (Linux)¶
On the client:
noac disables the attribute cache that makes rename look non-atomic to
concurrent readers. Yes, it slows down metadata operations — that's the
trade-off for correct leasing.
Workflow¶
# On one machine (or any one machine that mounts the share):
yt-uniq queue init /shared/queue
yt-uniq queue add /shared/queue /shared/sources/movie1.mp4 \
/shared/sources/movie2.mp4 \
/shared/sources/movie3.mp4
# On worker hosts A and B (both mount /shared):
yt-uniq worker /shared/queue \
--profile /shared/profiles/cid_aware.yaml \
--out-dir /shared/uniq/ \
--encoder libx264 \
--workers 4 \
--heartbeat-sec 30 &
# Monitor from anywhere:
yt-uniq queue status /shared/queue --json
# {"pending": 0, "in_progress": 2, "done": 1, "failed": 0}
Heartbeat and recovery¶
Each worker touches <host>.alive in in_progress/ every
--heartbeat-sec (default 30s). If the file's mtime is older than
stale_sec (default 300s, configurable via yt-uniq queue reset
--stale-sec), any other worker's next lease() will move that host's
leased files back to pending/.
Worker that died mid-encode loses its segment progress (each worker's
work_dir is local). The new worker re-encodes from scratch — acceptable
trade-off for not requiring a distributed work_dir.
Per-machine encoder variation¶
By design, two workers may use different encoders (one libx264 on
Linux, one hevc_videotoolbox on a Mac). The resulting variants of the
same source therefore differ slightly in H.264 / HEVC bitstream params.
This is usually fine for the CID-divergence use case (variants are
supposed to differ). If you need consistency, pin all workers to the same
encoder via --encoder libx264.
Output collisions¶
Workers write to out_dir/<input.stem>.uniq.mp4. If two workers happen to
lease two inputs with the same stem (e.g. movie.mp4 from two folders
that got flattened in pending/), the writes race. Avoid by:
- Ensuring unique basenames in
pending/, or - Setting per-worker
--out-dirand merging later.
Docker Compose example¶
services:
worker:
image: yt-uniquifier:0.3
volumes:
- /mnt/shared:/shared:rw # NFSv4 noac mount on host
environment:
YT_UNIQ_PROFILE: /shared/profiles/cid_aware.yaml
command:
- yt-uniq
- worker
- /shared/queue
- --profile=/shared/profiles/cid_aware.yaml
- --out-dir=/shared/uniq
- --encoder=libx264
- --workers=4
deploy:
mode: replicated
replicas: 4 # 4 workers in this stack
Cleanup¶
done/ is just a marker — files accumulate. Periodically:
(A yt-uniq queue prune --done-older-than 30d subcommand is deferred to
v0.4 — there's no work going into it before the real-CID validation
harness lands. Use the find snippet above in cron until then.)