iter8 (Phase 2.5 disabled, 32 MiB chunks) hit 28.7 MB/s mean
— 1.3 MB/s below R2 floor. Without Phase 2.5, WAIT_AFTER blocks the
mux thread directly once per chunk. Halving the WAIT_AFTER frequency
(32 → 64 MiB chunks) should raise the mean without changing the
underlying architecture.
iter7's 64 MiB chunks didn't help (23.1 vs iter4's 24.0). Reverting
to 32 MiB so iter8 differs from iter4 by exactly one variable:
Phase 2.5 disabled / direct passthrough mux thread.
Test question: is Phase 2.5 helping at all? If iter8 mean > iter4,
Phase 2.5 has been a net negative on this rig the whole time. If
iter8 mean < iter4, Phase 2.5 is doing what it claimed.
iter6 (8 MiB chunks) regressed mean -8.2 MB/s vs iter4 (32 MiB).
Trend: bigger = better in this workload, against the page-age
theory. Try 64 MiB to see if fewer/larger syncs raise mean further.
0.21.14 went to 128 MiB and was reverted; 64 is between.
iter5 (256-frame channel) regressed -1.8 MB/s vs iter4. Reverting to
32 frames.
iter4 sample pattern shows clear ~30 s oscillation (peak 45 MB/s →
dip 3 MB/s → recovery). Matches Linux vm.dirty_expire_centisecs
default (30 s). 32 MiB chunks at 25 MB/s issue WAIT_AFTER every
~1.3 s, which can't outrun the kernel's own page-age limit, so
pages buildup then flush in bursts. Smaller 8 MiB chunks
(WAIT_AFTER every ~0.33 s) should keep the dirty-page set young
and eliminate the periodic flush-burst dip.
iter4 measurement (Phase 2.5 + DONTNEED, no FileSectorSource buffer)
showed mean 24.0 MB/s but with dips to 2.8 MB/s. Channel at 32 frames
(~1.6 MiB at 50 KB/frame avg) drains in <100 ms on any producer pause,
starving the consumer.
256 frames = ~12 MiB ≈ 3-4 sec of consumer drain at the rolling-avg
floor we want to hit. Should coast through producer micro-stalls
without idle gaps.
iter1 baseline (Phase 2.5 + DONTNEED restored) measured at 18.4 MB/s
mean on Civil War remux — well below the rig's 37 MB/s concurrent-r+w
ceiling. Per-sector pread is the producer-side bottleneck: ~50us per
pread on NFS = ~19k preads/sec = effective ceiling near what we see.
The 0.21.3 bypass commit cited an A/B test showing the buffer hurt
throughput on NFS bidirectional workloads. That test was taken under
the 0.21.7 producer polling cap; once the cap is gone, the cap was
the bottleneck, not the buffer. Same invalidation pattern as the
0.21.5 Phase 2.5 revert.
Restored the buffered path. DONTNEED page-cache eviction from 0.21.6
stays intact.
Three coupled changes that together meet the "stable max throughput
on any medium" bar:
1. Remove `is_nfs` from `skip_wait`. WAIT_AFTER runs on every medium
now. The `bounded_syscall` 30s safety net (already in place) covers
the wedged-FS case generically — no need to predict NFS-hangs at
compile time. Re-applies the 0.21.12 fix that landed the flat
25 MB/s on NFS in the first place.
2. Drop the `posix_fadvise(DONTNEED)` calls after WAIT_AFTER. Bounded
*dirty* pages (the actual invariant) does not require evicting
*clean* pages. DONTNEED was forcing read-modify-write on any
in-window seek-then-write — exactly the pattern matroska
cluster-size backpatches produce. Empirical 2026-05-15: write side
of mux fed NFS at 43 MB/s while the file grew at 25 MB/s, an 18
MB/s overhead almost certainly composed of those RMW cycles. The
kernel reclaims clean pages under LRU when memory is actually
needed — we don't have to ask.
3. Replace `detect_nfs(fd) -> bool` with `detect_storage_class(fd)
-> (StorageClass, chunk_bytes_seed)`. Medium detection happens
*once* at construction and produces *one* output: the initial
`chunk_bytes` seed for the autotuner. NFS gets 64 MiB (commit ack
~10-30 ms; needs bigger chunks to amortize); other media use the
caller's hint. No hot-path branches on medium. The autotuner drives
all subsequent decisions from measured WAIT_AFTER p95 latency,
identically on every (OS, FS) combination.
Module-level doc rewritten to match: no more "NFS escape hatch", no
"is_nfs" framing. The whole writeback module is now medium-agnostic
except for one labelled detection point.
The pipeline's adaptive autotuner grows chunk_bytes only when the p95
WAIT_AFTER latency exceeds 200 ms. On NFS, sync_file_range(WAIT_AFTER)
translates to an NFS COMMIT RPC whose ack lands within ~10 ms — so
the autotuner never triggered and the pipeline stayed at the original
32 MiB initial value forever.
That capped sustained mux throughput by paying NFS COMMIT-RPC overhead
roughly once per second of writes. Bidirectional mountstats on rip1
2026-05-15: write side 29 MB/s (RTT 30 ms but exec_time 257 ms — 220 ms
queue/serial waiting) while concurrent dd on the same disk shows
~91 MB/s write + ~65 MB/s read available.
128 MiB initial drops COMMIT cadence 4x while keeping the bounded-
cache invariant intact (worst-case dirty pages ~2 x chunk = 256 MiB,
well under vm.dirty_ratio = 6.6 GB on the 32 GB rig). The adaptive
autotuner can still grow further (up to 256 MiB) or shrink if
WAIT_AFTER ever measures sub-20 ms p95 on faster media.
The WritebackPipeline's WAIT_AFTER + posix_fadvise(DONTNEED) dance
keeps dirty pages bounded at ~2 × chunk_bytes by waiting for each
chunk's writeback to commit before issuing the DONTNEED hint to drop
it from cache. The original 0.18-era design unconditionally skipped
this on NFS on the premise that "NFS clients have their own buffering
and commit semantics that handle dirty-page bounds without us forcing
the issue."
Empirically wrong. On unraid-1 NFS the kernel client buffers dirty
pages up to vm.dirty_ratio (default 20% of RAM = ~6.6 GB on the rip1
host) before the kernel forces writeback and throttles app writes.
Result on 0.21.11 mux measured 2026-05-15: mux throughput cycled
between ~45 MB/s (cache absorbing) and ~7 MB/s (cache draining under
throttle) on a ~100 s period — exactly the burst-flush pathology this
pipeline was built to fix, but disabled on the medium it actually
runs on. /proc/meminfo Dirty: column climbed lockstep with mux
write rate during the slow half of every cycle, confirming the cause.
The original safety concern — `sync_file_range(WAIT_AFTER)` hanging
indefinitely on a wedged NFS server — is already handled by
`wait_after_with_timeout`'s `bounded_syscall` wrapper (30 s deadline).
If a real WAIT_AFTER call exceeds the deadline the pipeline flips to
the `degraded` state and skips WAIT_AFTER + DONTNEED for the rest of
its life — same effect as the old NFS branch, but only triggered when
something is genuinely broken rather than as a blanket exception.
This change is medium-agnostic: every medium goes through the same
path now, every medium gets the same safety net, and the
ADAPTIVE_WINDOW chunk-size autotuner (lines 226-255 — measures p95 of
WAIT_AFTER and resizes between 4 MiB and 256 MiB) finally activates
on NFS where previously it was dead code. Slow medium auto-grows
chunks to amortise per-chunk overhead; fast medium auto-shrinks to
keep cache pressure tight; nothing in the code special-cases the
filesystem type.
`is_nfs` is still detected (for logging + observability) but no
longer keys `skip_wait`. Module doc + startup log line updated to
match.
Pre-coalescing the Phase-2.5 writer thread issued one `file.write_all`
syscall per `Cmd::Write` dequeued. The mux side calls
`WritebackFile::write_all(buf)` per PES frame, typically 30-200 KB.
On NFS that translates to one RPC per syscall, capping per-thread
throughput at `(wsize / rtt) × inflight` — well below what the same
disk delivers under a 1 MiB `dd oflag=direct` workload (empirical
2026-05-15: dd 71 MB/s vs mux ~25 MB/s sustained, with instantaneous
samples bursting 7→108 MB/s as the kernel page cache filled and
drained on its own cadence).
Coalesce instead: dequeue drains consecutive `Cmd::Write` items off
the ring up to a 1 MiB byte budget, returns them as
`DequeuedWork::Writes(Vec<Vec<u8>>)`, and the run loop concatenates
into one contiguous buffer and issues a single `file.write_all`.
Non-write commands (Seek, Flush, SyncAll, Finish) break the run and
are returned one at a time as `DequeuedWork::Other`, preserving their
ordering relative to the writes.
Single-buffer fast path avoids the concat allocation when only one
write is in the queue at dequeue time. A single oversize write (e.g.
the rare matroska cluster larger than 1 MiB) is admitted alone so it
still makes progress — the kernel splits internally.
This is generic across mediums: bigger app writes are at-least-as-
good on local SSD, HDD, or NFS. On fast storage the ring rarely fills
so coalescing is mostly a no-op; on slow storage with significant
per-RPC overhead it materially improves throughput.
Tests in `tests/` (write_then_drop_persists_bytes, sync_all_drains,
seek_then_patch_roundtrip, flush_is_observed_in_order) still pass —
ordering and durability semantics are unchanged.
The depth was conservative because pre-0.21.8 a full sync_file_range
stall on NFS could traverse this channel and pin the producer. With
0.21.8's restored Phase 2.5 writer thread + 128 MiB byte-bounded ring
inside WritebackFile, every blocking syscall happens downstream of
this channel — never on it. The original "smaller buffer reduces
backpressure risk" rationale no longer applies.
Empirical (2026-05-15 Civil War UHD remux on 0.21.8): 30.9 MB/s
sustained but instantaneous samples spanning 9-53 MB/s, stdev 9.3.
Distribution clusters 52% of samples in 25-35 MB/s but has a long
9-15 MB/s tail. The tail corresponds to brief stalls in the matroska
builder when the sink momentarily lags — exactly the case a deeper
inter-thread channel covers. 32 frames at PES-frame sizes is still
well under a megabyte of additional memory, so the cost is zero.
Reverts 1001084. That revert was made on the premise that Phase 2.5
caused a ~60% mux throughput regression on NFS bidirectional workloads.
The premise was wrong: at the time of measurement the producer was
capped at ~8 MB/s by a 50 ms thread::sleep poll in
Pipeline::send_with_halt (fixed in v0.21.7's io/pipeline change), so
the comparison was measuring the polling cap on both sides.
With the polling cap removed, direct passthrough exposes the kernel's
default dirty-page writeback pathology on NFS: writes accumulate, the
kernel periodically bursts a flush, app writes block for the burst.
Observed empirically on Civil War UHD remux 2026-05-14: 5-45 MB/s
spiking around a ~21 MB/s sustained mean, dominated by burst-flush
back-pressure cycles.
Phase 2.5 decouples the mux thread from the file syscall:
* mux writes complete instantly into a 128 MiB byte-bounded SPSC ring,
* a dedicated writer thread executes the real File writes, seeks,
and sync_file_range calls; can sit in a kernel burst without
blocking the mux pipeline,
* backpressure via Condvar notify/wait, no polling primitive,
* the ActiveClusterBuffer fast-path preserves the original MKV
cluster-backpatch optimisation so in-window seeks don't drain
the current writeback chunk.
Halt-safety is preserved: every blocking writeback syscall on the
writer thread still routes through bounded_syscall with a 60 s
deadline. A wedged NFS server cannot trap the writer indefinitely;
the muxer keeps queueing into the ring; the kernel page cache and
the ring together absorb the stall.
The halt-aware send loop polled try_send on a 50 ms sleep slice when
the channel was full. That capped producer throughput at 1 / 50 ms =
20 frames/sec ≈ 1 MB/s at typical PES frame sizes — way below NFS,
let alone local SSD/NVMe. Multi-day diagnostic 2026-05-13/14 surfaced
it as the root cause of the 18 → 3 MB/s mux throughput regression on
0.21.x.
Replaced std::sync::mpsc::sync_channel with crossbeam_channel::bounded
and switched send_with_halt to send_timeout. Producer now BLOCKS on
consumer drain (kernel-wakeup) instead of polling — the timeout slice
(250 ms) only fires when the producer is waiting for stop signal
observation, never on the happy path.
Result: no throughput cap from this primitive at any medium speed.
Mux is bounded by actual storage / network bandwidth, not by our
channel implementation.
Same fix needs to land in autorip's mux.rs producer-side
sync_channel (separate sibling commit).
Documented in freemkv-private/memory/
feedback_send_with_halt_poll_throttle.md.
Empirical: isolated NFS read 70 MB/s + write 93 MB/s on the rip1 setup
right now, but mux throughput pinned at 2.7 MB/s on 0.21.5. NOT
environmental — code regression.
Root cause: Phase 1 silently dropped the read-side
posix_fadvise(POSIX_FADV_DONTNEED) eviction that the pre-Phase-1 (0.20.7)
hot path had. Without it, an 85 GB streaming ISO read pins the entire
file in the kernel page cache, starving concurrent MKV writeback. 0.21.2
then also dropped the POSIX_FADV_SEQUENTIAL hint on the same theory,
compounding the regression.
Restored both, per-OS split:
- linux: posix_fadvise(SEQUENTIAL) at open + posix_fadvise(DONTNEED)
on consumed 32 MiB windows
- macos: F_RDADVISE hint at open (kept); drop_window no-op (macOS unified
buffer cache less prone to the pin pathology)
- windows / other: both no-op stubs
Target mux speed restored to 20+ MB/s (per concurrent-NFS math:
70/2 read × 0.73 MKV/ISO ratio ≈ 25 MB/s achievable).
The Phase 2.5 writer-thread + bounded ring + ActiveClusterBuffer
architecture introduced a ~60% mux throughput regression on NFS
bidirectional workloads (18 MB/s -> 7-8 MB/s in 0.21.x). Other
candidates (FileSectorSource readahead buffer, per-OS sink split)
ruled out empirically across 0.21.1-0.21.4.
Reverts WritebackFile's hot path to direct File passthrough matching
the 0.20.7 baseline:
* Write/write_all/flush call straight through to self.file.
* Seek calls through to self.file and notifies the pipeline.
* sync_all runs pipeline.finalize() then per-OS durable_sync.
Kept (intentional): the per-OS file split established in earlier work
(writeback_file/{linux,macos,windows,other}.rs). create_with_size_hint
still dispatches to platform::preallocate and sync_all still routes
through platform::durable_sync (bounded_syscall 60s deadline on
Linux/macOS), so halt-safety on a wedged NFS fsync is preserved.
Removed:
* Cmd / RingState / Shared / WriterState / writer_thread_main
* ActiveClusterBuffer (active-cluster window for in-window seeks)
* push_command / dequeue / publish_error / mark_writer_gone
* RING_CAPACITY_BYTES, ACTIVE_CLUSTER_WINDOW_BYTES,
MAX_WRITE_CHUNK_BYTES, WRITER_THREAD_NAME constants
* writer JoinHandle field + Drop join logic
* Tests that asserted internal writer-thread state
(backpressure_blocks_when_ring_full, in/out_of_window_seek_*,
active_cluster_buffer_*, writer_thread_panic_surfaces_on_drop)
Kept tests (correctness-of-output only):
* write_then_drop_persists_bytes
* sync_all_drains_and_flushes (renamed from sync_all_blocks_until_ring_drains)
* seek_then_patch_roundtrip (replaces in_window_seek_then_patch_roundtrip)
* flush_is_observed_in_order
mod.rs: 1026 -> 303 lines. linux/macos/windows/other.rs untouched.
The three tests (multi_sector_read_spanning_buffer_boundary,
backward_seek_rebuffers, partial_buffer_at_eof) were asserting
internal buf_start_lba / buf_len_sectors state. With 0.21.3's
read-path bypass, those fields are no longer mutated. The
byte-level contract assertions (read returns correct bytes for
every scenario the tests cover) remain intact.
The 32 MiB readahead window (0.21.0–0.21.1) regressed mux throughput
on NFS bidirectional workloads vs the pre-Phase-1 0.20.7 baseline
(18 -> 7-8 MB/s). The 0.21.2 4 MiB shrink made it worse (5-6 MB/s).
Both signs point at the application-level buffer itself, not the size.
This commit bypasses the buffer entirely on the read path — every
read_sectors call seeks and pread()s direct to the file. That matches
0.20.7's hot path. Kernel readahead handles the policy; on NFS that
interleaves naturally with concurrent writes on the same TCP
connection.
Buffer state fields and refill/buffer_covers are kept so the
structure is preserved for a future per-source-type policy (e.g. a
local-disk source where batched reads ARE beneficial), and so the
existing tests still exercise that machinery.
Empirical regression observed during 0.21.1 mux test on rip1/unraid-1:
historical 0.20.7 baseline averaged ~18 MB/s mux throughput; 0.21.1
dropped to ~7-8 MB/s flat. Same NFS source + destination, same disc.
Suspect: 32 MiB FileSectorSource readahead + posix_fadvise(SEQUENTIAL)
together saturate the TCP connection on read bursts, starving the
writer thread's concurrent NFS writes (mux reads the source ISO and
writes the MKV over the same connection).
- READAHEAD_BUF_BYTES: 32 MiB -> 4 MiB. Matches NFS rsize=1 MiB * 4
round-trips per refill, interleaves cleanly with writes.
- linux/hint_sequential: now no-op. Kernel's default ~128 KiB
readahead is what we want on NFS-backed ISOs (the dominant case).
Per-OS file stays so we can re-enable a hint cleanly later if a
different path benefits.
CI's lint workflow runs clippy on linux target where:
- platform/fs_type/linux.rs and io/writeback/linux.rs: the i64 cast
on buf.f_type / NFS_SUPER_MAGIC is unnecessary on glibc x86_64 (both
already i64) but required on musl (c_ulong); silence the lint via
inline allow with explanatory comment.
- mux/m2ts_mux/packet.rs: doc comment continuation across lines was
parsed as an unindented list item. Reworded to a single flowing
sentence.
- m2ts_mux: replace u64::is_multiple_of (stable in 1.87+) with %.
CI on the pinned 1.86 toolchain rejected the unstable feature use.
- mux/{fmp4,hevc,m2ts_mux}: rustfmt drift cleanup (test-code wrapping).
SocketSink + UdpSocketSink (`src/io/sink/socket.rs`) — sequential-only
TCP/UDP write destinations. SocketSink wraps BufWriter<TcpStream> with
1 MiB capacity, tunes SO_SNDBUF on construction, calls shutdown(Write)
on finish(). UdpSocketSink emits one datagram per write — caller
packetizes. Both impl Write+Send and thus satisfy SequentialSink via
the Phase 2 blanket; neither impls Seek, so RandomAccessSink is
correctly inaccessible (compile error to mux MKV onto a socket).
New sequential container muxers in src/mux/:
- hevc/ — raw HEVC Annex B elementary stream. Length-prefixed NALU
→ 00 00 00 01 NALU. hvcC parsing emits VPS/SPS/PPS once at stream
head. Fully ships.
- m2ts_mux/ — standard MPEG-TS (188-byte packets). Single program,
HEVC video on PID 0x100, optional AC3/TrueHD audio on PID 0x101.
PAT+PMT re-emitted every 250 packets; PCR stamped on video every
40 packets. Hand-rolled, no new deps. Distinct from the existing
BD-TS (192-byte) `mux::m2ts::M2tsStream` — that path stays as-is.
- fmp4/ — fragmented MP4. STUB: ftyp + minimal moov skeleton with
one HEVC video trak + mvex/trex. Media fragments (moof+mdat) are
TODO for v0.22.0 — write_video accumulates frames into a pending
buffer that finish() clears. Init segment is well-formed enough
that init_segment_starts_with_ftyp_then_moov asserts the box
chain.
17 new unit tests added (socket round-trip, HEVC Annex B conversion,
M2TS packet alignment + PAT/PMT cadence + per-PID CC, fMP4 box chain).
All 514 lib tests + 17 new = pass on Rust 1.86 (fmt + clippy + test
via freemkv-private/scripts/precommit.sh libfreemkv).
No new dependencies. No version bump. Don't-touch list clean.
WritebackFile now offloads all File I/O to a dedicated writer thread.
Muxer's Write/Seek/sync_all push into a bounded byte ring; the writer
drains the ring and runs the syscalls.
- Commands: Write(Vec<u8>), Seek(SeekFrom), SyncAll(oneshot),
Finish(oneshot). One ring carries all four; ordering preserved.
- ActiveClusterBuffer mirrors the last ACTIVE_CLUSTER_WINDOW_BYTES
written. Seek-back within window → in-memory patch + re-emit
(no forced drain). Seek-back outside window → real drain+seek.
Wins for MKV: cluster size patches almost always land inside the
active window; only the end-of-mux Cues + segment header backpatch
fall outside.
- Ring capacity RING_CAPACITY_BYTES = 128 MiB. Backpressure on full
blocks the muxer (correct semantics for archival workflows).
- All syscalls in writer thread wrapped in bounded_syscall(60s) so
a wedged NFS doesn't pin the thread forever.
- Drop drains via Finish + joins the writer thread; sync_all blocks
until ring drained AND underlying fsync completes.
Introduces the SequentialSink / RandomAccessSink trait pair under
io::sink and an open_for_mkv dispatch helper that picks WritebackFile
on Linux+NFS and LocalFileSink everywhere else. LocalFileSink wraps
BufWriter<File> with a 4 MiB buffer and exposes a per-OS preallocate
path (fallocate on Linux, F_PREALLOCATE on macOS, no-op fallback).
Adds platform::fs_type::detect with a per-OS split (statfs on Linux /
macOS, UNC heuristic on Windows, Unknown elsewhere) so construction-
site dispatch has a single primitive to call.
Blanket impls cover the common shapes: any Write+Send is a
SequentialSink, and any SequentialSink+Seek is a RandomAccessSink.
WritebackFile satisfies the random-access trait via the blanket impl
without needing an explicit per-type impl. No callers wired yet — the
mux::resolve construction sites stay on WritebackFile pending Phase 3.
Tests: 5 new sink/preallocate tests + 3 fs_type tests (1 ignored,
needs a real NFS mount). cargo +1.86 fmt + clippy + tests all green.
Three changes targeting 0.20.9's "muxer never read-stalls on NFS read
latency" invariant:
A. FileSectorSource gets a 32 MiB internal read-ahead buffer
(READAHEAD_BUF_BYTES). Splits out from src/sector/file.rs into
src/io/file_sector_source/ with per-OS open hints (Linux
posix_fadvise(SEQUENTIAL), macOS fcntl(F_RDADVISE) with 64 MiB
cap, Windows TODO stub, BSD/illumos no-op). Backward seeks
rebuffer; partial reads at EOF return only the bytes that exist;
oversize-request bypass for count > BUF_SECTORS.
B. WritebackFile inline #[cfg(target_os = "linux")] blocks split
into per-OS files under src/io/writeback_file/. Linux unchanged
(fallocate KEEP_SIZE, fsync via bounded_syscall). macOS gets a
real F_PREALLOCATE + F_FULLFSYNC impl (was a "skipped (non-linux)"
debug log before). Windows is a stub (FlushFileBuffers via
std sync_all; TODO for SetFileValidData). BSDs/illumos fall back
to std sync_all.
C. New byte_channel module — byte-bounded producer/consumer wrapping
std sync_channel with Mutex/Condvar byte accounting. Sender blocks
when used_bytes + item.byte_size() > capacity. HasByteSize impl
for PesFrame. Default cap BYTE_CHANNEL_DEFAULT_CAPACITY = 64 MiB,
sized to absorb worst-case NFS read p99 (~2 s × UHD peak compressed
~15 MB/s). The mux call site lives in autorip (out of scope here);
this lands the primitive in libfreemkv for autorip to adopt.
Test counts: byte_channel +6, file_sector_source +5, sector::file
round-trip suite (3) preserved. passn_handler_ab.rs A/B fixture
(8 profiles) still green.
precommit.sh libfreemkv: fmt + clippy + test all green on Rust 1.86.
No version bump; no Cargo.lock changes; no forbidden-file edits
(disc/patch.rs, disc/read_error.rs, io/pipeline.rs,
tests/passn_handler_ab.rs).
The 0.20.7 work lives downstream in autorip — process-level safety net
(hard watchdog 5-min mux escalation + std::process::exit, restart-loop
counter + auto-quarantine after 3 attempts, partial-state preservation,
failure_reason surfaced in /api/state).
Bumping libfreemkv to 0.20.7 keeps the 4 crates on the unified release
train per CLAUDE.md.
Generalizes 0.20.5's hand-written wait_after_with_timeout into a
reusable primitive. After this change, every blocking syscall in the
recovery + mux paths is wrapped, so cooperative Halt has bounded
~250 ms latency reach even into kernel-owned thread states.
New module src/io/bounded.rs:
- BoundedError { Halted, Timeout, WorkerLost }
- bounded_syscall<F, R>(halt: Option<&Halt>, timeout, op) -> Result<R, BoundedError>
- Worker thread runs op; main thread recv_timeouts on a rendezvous
channel in 250 ms slices, polling halt between slices.
- Worker is intentionally leaked on timeout/halt — kernel reaps when
the syscall finally returns or at process exit. Calling thread is
NEVER trapped inside a kernel call.
- 6 unit tests cover the happy path + each error variant.
Refactored callsites:
- src/io/writeback/linux.rs::wait_after_with_timeout now delegates
to bounded_syscall. ~30 LOC of duplicated channel/thread plumbing
deleted. Same semantics, cleaner.
- src/io/writeback_file.rs::WritebackFile::sync_all now wraps the
final libc::fsync(fd) with bounded_syscall (60 s deadline). On
timeout: log error at target=mux and return Ok — kernel will flush
on close, best-effort but bounded. Covers FileSectorSink::finish,
PatchSink::close, SweepSink::close, and the mux MKV finalize path
(they all sync through WritebackFile).
What still hangs (deliberately not wrapped — too hot a path):
- File::write itself. Per-frame write on a wedged NFS could still
block; but back-pressure from a stuck consumer means the producer
notices within seconds, not minutes — different failure mode than
the WAIT_AFTER hang 0.20.5/0.20.6 fix.
Targets the recurring mux hang on NFS dest where the consumer thread
sits indefinitely inside libc::sync_file_range(SYNC_FILE_RANGE_WAIT_AFTER)
because the NFS server never returns a commit ack. The whole rip
wedges; halt is cooperative and can't reach inside a kernel syscall.
A. NFS detection at WritebackPipeline construction (fstatfs f_type ==
NFS_SUPER_MAGIC 0x6969). When NFS:
- Skip SYNC_FILE_RANGE_WAIT_AFTER entirely.
- Skip posix_fadvise(DONTNEED) — NFS client handles its own buffering.
- Still issue async SYNC_FILE_RANGE_WRITE (harmless hint).
Cannot hang on a syscall not made. fstatfs failure fails open (assume
local). Logged at info on construction so operators see which strategy
is active. The whole hang vector is removed for NFS deployments.
B. Hard timeout on WAIT_AFTER for non-NFS (defense in depth, since
even a degraded local disk could in principle hang the syscall).
Each WAIT_AFTER runs on a worker thread; main thread waits on a
sync_channel rendezvous with 30s deadline. On timeout: log error,
set per-pipeline 'degraded' Arc<AtomicBool>, downgrade to NFS-style
skip for the rest of the pipeline's life. Worker thread leaks
intentionally — it'll unwind when the syscall eventually returns or
the process exits. Converts indefinite freeze into 'log loud +
downgrade + keep ripping'.
C. Diagnostic logging for the 73%-of-this-movie reproduction:
- WritebackFile::seek logs every non-trivial seek (from, to, signed
delta) at target=mux so we can see if MkvMuxer seeks back before
a stall.
- WritebackPipeline::finalize logs the chunk being finalised before
any WAIT_AFTER call, so a hung chunk is identifiable by offset.
No new dependencies. macOS / Windows noop stubs unchanged. Net
+198 LOC libfreemkv (mostly writeback/linux.rs).
Four targeted changes to maximize mux throughput regardless of storage
backend (local SSD, local HDD, NFS, network share) and surface enough
log data to diagnose 'mux slow' reports without a re-rip:
1. POSIX_FADV_SEQUENTIAL on FileSectorSource::open (Linux only).
Widens the kernel readahead window for sequential ISO reads. One
syscall at open, free on every storage type.
2. POSIX_FADV_DONTNEED on the ISO read side after every 32 MiB chunk.
Mirrors the writeback DONTNEED that already runs on the write
side. Keeps the read-side page cache bounded during multi-GB ISO
reads — eliminates the OOM-pressure / eviction-storm risk on
long mux runs. Linux only; per-drop trace at target="mux".
3. WritebackFile::create_with_size_hint(path, size_bytes) calls
fallocate(FALLOC_FL_KEEP_SIZE) on Linux to pre-reserve extents
for the output. Reported file size stays 0 (writes grow it
naturally) but the on-disk extent allocation is contiguous —
reduces extent fragmentation for big sequential muxes. Wired
into mkv:// and m2ts:// output paths via DiscTitle::size_bytes.
No-op on macOS/Windows; old create() kept with #[allow(dead_code)]
for callers without a size hint.
4. Adaptive WRITEBACK_CHUNK_BYTES in the Linux writeback pipeline.
Tracks sync_file_range(WAIT_AFTER) elapsed_ms in a rolling
16-sample window. p95 > 200 ms → double chunk size (cap 256 MiB).
p95 < 20 ms → halve (floor 4 MiB). One algorithm, both
fast-storage (small chunks, responsive) and slow-storage (big
chunks, fewer commit round-trips) optimized. Per-chunk trace +
per-32-chunk debug snapshot + info-on-resize so an operator can
see where the autoscaler settled.
All four are universal — no storage-type detection, no env vars to
flip, no per-deploy tuning required. Total +201/-6 across four files.
WO-5 (partial): the patch backtrack inner loop ('while bt_pos <
backtrack_end' in disc/patch.rs) issues per-sector reads to fill the
gap created by a damage-window skip. A long backtrack span can run
minutes; without an inline halt poll, the outer halt only takes
effect when control returns to the per-range loop. Adds a halt poll
at the top of each iteration so cancellation propagates inside the
backtrack span.
Per-sector read failures inside the backtrack already drop through
to the main fail path; this only changes the cancellation latency
between an /api/stop call and the producer actually unwinding. Drops
worst-case unwind from 'whole backtrack span × per-sector recovery
timeout' (minutes) to 'one in-flight SCSI command' (seconds).
The broader Arc<AtomicBool> → Halt migration on CopyOptions /
SweepOptions / PatchOptions / Drive::halt and Pipeline::send halt-
awareness is deferred — separate cycle, larger API impact.
WO-3c: Delete dead non-NOT_READY retry block (~100 LOC). The block
declared retry_count = 0 inside the per-iteration Err arm, so the
'MAX_NON_NOT_READY_RETRIES=3' budget actually fired exactly once
(1s pause + 1 retry) before falling through to NonTrimmed. The
'exponential backoff: 2s, 4s, 8s' comment was wrong by construction.
Cross-pass NonTrimmed retry (each pass gives the same sectors another
shot) already covers the recovery case, and gives the drive minutes
between attempts instead of 1-8 seconds — empirically what stochastic
recovery on the BU40N actually needs.
WO-4 (targeted slice): Add wedge-family cooldown on HARDWARE_ERROR /
ILLEGAL_REQUEST senses. These are what the BU40N's firmware fast-fail
state returns; every subsequent read in that state comes back in
<100ms. Pre-fix patch hammered the drive: mark NonTrimmed, sleep 1s,
advance, hit next wedge, mark, sleep 1s — exactly the rapid-retry
cadence the firmware is sensitive to. Now a wedge-family sense triggers
WEDGE_FAMILY_COOLDOWN_SECS=30 cooldown (matches read_error.rs's
ZONE_ENTRY_COOLDOWN_SECS), and WEDGE_ABORT_THRESHOLD=16 consecutive
wedges aborts the pass for autorip eject+reload. Any non-wedge read
clears the counter.
Also drops the duplicate NonTrimmed dispatch (Mapfile::record is
idempotent so it wasn't a correctness bug, but it doubled per-failure
consumer work).
WO-2 (delete SectorReader trait):
- The 0.18 trait split into SectorSource (read-only) and SectorSink
(write-only) is final; the legacy SectorReader alias was a bridge.
- Renames every internal &mut dyn SectorReader (~25 sites) to
&mut dyn SectorSource. The trait method capacity() becomes
capacity_sectors() with a default of 0 (preserves SectorReader's
default-0 behavior).
- Deletes the SectorReader trait, its blanket-to-Source bridge, and
the FileSectorReader type alias. Adds explicit forwarding impls
for Box<dyn SectorSource> and &mut dyn SectorSource so generic
decorators like DecryptingSectorSource<S: SectorSource> compose.
WO-3a (extract Disc::patch):
- Moves Disc::patch (1230 lines) and bytes_bad_in_title from
disc/mod.rs into disc/patch.rs as a split inherent impl. Zero
behavior change — pure mechanical relocation. disc/mod.rs drops
from 3,945 to 2,714 LOC.
WO-6 (partial):
- Deletes src/labels/png_filenames.rs — was a 72-LOC stub with
detect() returning false, never wired into the PARSERS registry.
CLAUDE.md doc drift fixes (audited 2026-05-13):
- JUMP_BASE_SECTORS: 256→1024 (64 MB base for UHD, not 8 MB)
- PASSN_DAMAGE_THRESHOLD_PCT: 12→6
- PASSN_SKIP_SECTORS_BASE: 64→32
- MAX_RANGE_SECS=180: replaced by proportional range_sectors × 25,
capped at RANGE_BUDGET_CAP_SECS=1800.
The 0.18 trait split into FrameSource (read-only) and FrameSink
(write-only) was an over-engineered API. Consumers don't think
"frame source backed by MKV" — they think "open MKV for reading".
The split paid a real API-complexity cost (two trait names, two
re-exports, dual impls per bidirectional type, deprecation bridge)
for one marginal property: compile-time direction-safety at the
trait-object boundary. The runtime error path on a wrong-direction
call (StreamReadOnly / StreamWriteOnly) is unambiguous and rare in
practice.
Deletions:
- pes::Stream is no longer #[deprecated]
- pes::FrameSource trait + its blanket-from-Stream bridge
- pes::FrameSink trait + the trampoline impls on every concrete type
- The compile-time-direction-safety test scaffolding
- Crate-root FrameSource / FrameSink re-exports
Additions:
- Stream is now Send-bounded (Stream: Send supertrait). Every
concrete impl was already Send-compliant — Box<dyn Read + Send>
and Box<dyn Write + Send> were already in place on the trait
objects MkvStream / M2tsStream / etc hold internally. Promoting
Send into the trait makes Box<dyn Stream> Send too, which lets
autorip drop its SendStream unsafe newtype.
The public API is now: one Stream trait, one concrete type per
format, two constructors (open/create or input/output). Bidirectional
types route through internal Mode { Read | Write } discriminants.
Net: -347 lines libfreemkv, -38 lines autorip, -5 lines freemkv.
v0.19.0 was tagged with a search-and-replace gone wrong:
rust-version, serde, and zip all had their version strings
replaced with "0.19.0". Edition 2024 rejected rust-version
0.19.0 (< 1.85), failing every CI build. No artifacts shipped
to crates.io.
Repair:
- rust-version: 0.19.0 → 1.86 (CI pin)
- serde: 0.19.0 → 1
- zip: 0.19.0 → 2
Also drops an unused start_lba binding in mux/disc.rs that
clippy 1.86 catches.
scan_with() collapsed every failure path from resolve_encryption() into
None via .ok(), so callers couldn't tell the difference between "no
KEYDB found", "KEYDB failed to parse", "disc hash not in KEYDB and
fallback derivation failed", "AACS files unreadable on disc", and a
handshake that rejected every host cert. autorip's UI was stuck
printing "no decryption keys found (check KEYDB)" for all of them,
which is a particularly bad message when the user has actually loaded
a KEYDB and the real failure is something else.
Changes:
- New pub field Disc.aacs_error: Option<Error>. Populated by scan_with
whenever encrypted && aacs.is_none(). Sentinel KeydbLoad path
"<no keydb in search paths>" distinguishes the no-keydb case from
a real load failure without adding a new Error variant (which would
be a breaking change for downstream exhaustive matches).
- tracing::warn in scan_with at scan_aacs_resolve_failed and
scan_aacs_no_keydb, with error_code and keydb path for grepping.
- tracing in do_handshake: keydb load failure, host-cert exhaustion
(with cert count and last error code), VID read failure post-auth,
and a debug-level success log. Lets us see whether handshake even
got off the ground for a given disc.
Test fixtures updated to set aacs_error: None.
The 'All probes failed — possible wedge condition' log fired during patch
probing whenever 10+ consecutive failures hit AND a probe sweep at the
local zone returned 0 successes. This was distinct from the read_error.rs
'wedge_transition' log that fires when the SCSI sense family ACTUALLY
flips into Hardware/IllegalRequest fast-fail mode.
Two logs both saying 'wedge' caused operator confusion during the
2026-05-11 Dune Pt 2 wedge investigation — was the drive wedged, or was
it just a zone of fully-bad sectors? They mean different things.
Relabel to 'patch_zone_fully_bad' with explicit pointer to read_error.rs
for the canonical wedge detection. Same triggering condition; just clearer
wording in the log stream.
Disc-04 (Top Gun: Maverick) re-test 2026-05-11 surfaced a real-world
bdmt_eng.xml where <di:description> contained no prose, only nested
<di:thumbnail href="…"/> elements. The previous parser surfaced
the raw XML fragment as the description string ("<di:thumbnail
href=\"tgm_meta_sm.jpg\" />\\r\\n <di:thumbnail
href=\"tgm_meta_lg.jpg\" />"). Worse than no description.
Fix: filter description candidates that begin with `<` after
trimming. Real prose never starts with an angle bracket; XML-only
content always does. Net: title extraction unaffected (it uses its
own element-priority path); description field drops when it would
otherwise carry XML noise.
Two new bdmt tests, 12 of 12 passing.
Three layered sources of stream labels now, in precedence order:
1. **Framework parser** (paramount/criterion/pixelogic/ctrm/dbp/deluxe)
— editorial labels with purpose/qualifier ("English Atmos",
"Director's Commentary", "English SDH"). High or Medium confidence.
2. **MPLS gap-fill** (`fill_gaps_from_mpls`) — every stream the
playlist references gets at least a basic lang+codec label, even
when the framework parser missed it.
3. **CLPI orphan append** (`append_clpi_orphans`) — streams in
/BDMV/CLIPINF/*.clpi ProgramInfo that NO MPLS playlist references.
Empirical (2026-05-11): ~5% of streams across the 11-disc corpus,
most dramatic on disc-02 (HDMV-only) at 40% CLPI-only.
Orphan numbering: each appended orphan gets
`stream_number = max(existing per type) + N` so playlist-reachable
streams keep their original positions and orphans sort cleanly at
the tail.
Orphan dedup: (stream_type, language, codec_hint) tuple — fuzzier
than PID matching (PIDs aren't carried on StreamLabel) but it's the
only signal available downstream of the gap-fill. False positives
(genuine orphan that happens to share lang+codec with an existing
entry) silently drop, which is the conservative failure mode — the
user-facing display would just see a confusing duplicate otherwise.
`mpls_universal::language_display_name` and `::codec_name` promoted
from private fn to pub(crate) so this module can build orphan labels
with consistent naming.
Tests: 2 new in gap_fill_tests — synthetic-input verification of the
dedup tuple logic and the stream_number assignment. 6/6 tests in the
gap-fill module now passing.
Rewrites the Pass 1 wedge handling from "slow skip after the drive
has already wedged" to "prevent the wedge transition in the first
place." Driven by 2026-05-11 empirical data: the BU40N transitioned
into IllegalRequest fast-fail mode at exactly 7 medium errors in
6.5 seconds (~1 read/sec retry cadence). Once there, only physical
eject + reload clears it — 30s pauses + 1 GB jumps do not.
The fix is the user's mental model from that session:
"We can detect bad reads, failed reads, and asking to read again
fast after causes a wedge. We need to prevent the wedge in the
first place."
Two changes to the centralized error handler:
1. **`for_sweep().fast_jump_threshold = 1`** (was 4). Pass 1 now
JumpAheads on the FIRST outer-batch failure, not the 4th. The
drive never gets back-to-back retries at the same LBA in Pass 1
— every error → jump 64 MB forward + long cooldown. Pass N keeps
`fast_jump_threshold = u64::MAX` because retries on already-known-
bad LBAs are its whole job.
2. **`ZONE_ENTRY_COOLDOWN_SECS = 30`**. The FIRST error after a
clean run (when `consecutive_outer_failures == 1` and we're not
bisecting) uses this long pause instead of the standard 5 s
FAIL_PAUSE_SECS. Gives the BU40N's firmware / bridge internal
retry counters 30 s of breathing room before the next read,
preventing the "7 errors in 6.5 s" cascade. Subsequent errors
in the same zone use the standard 5 s pause (we've already
jumped past the initial damage; further errors mean we landed
in another bad cluster).
Pass N exempt from the zone-entry cooldown — `bisect_on_marginal=
true` skips the long-pause arm. Pass N's per-sector retries on
known-bad LBAs would multiply uselessly with 30 s/error.
Test updates: 4 tests' expected behavior changed under the new
policy. Renamed `pass_1_marginal_skips_instead_of_bisecting` →
`pass_1_marginal_jumps_immediately_not_bisecting`. Renamed
`pass_1_jumps_after_4_consecutive_outer_failures` →
`pass_1_jumps_immediately_on_first_outer_failure`. Updated
`both_passes_pause_on_failed_read_for_wedge_avoidance` (now
`pass_1_zone_entry_uses_long_cooldown` + `pass_n_pauses_uniformly_on_failed_read`).
Cost analysis:
- Clean disc (no errors): unchanged. 0% overhead.
- Lightly damaged (1-2 zones): +30 s per zone = ~1 min total. Fine.
- Heavily damaged (10+ zones): +5+ min total. The trade for never
wedging the drive and getting a usable Pass N afterwards.
Expected behavior on the next damaged-disc rip:
- Pass 1 hits damage at LBA X → jumps 64 MB forward immediately,
pauses 30 s
- Drive's firmware never accumulates the retry pressure that triggers
IllegalRequest fast-fail
- bytes_maybe accumulates faster (we skip more), but Pass N picks up
the slack with proper per-sector recovery — and Pass N can actually
RUN because the drive isn't wedged
Two layered changes, in service of the empirical question "is CLPI
truly redundant with MPLS for label data?":
1. **clpi.rs ProgramInfo parser**. The existing CLPI parser only
walked the EP map (for sector-range lookups). Added a parser for
the ProgramInfo section's per-stream stream_coding_info table:
pid, coding_type, audio_format/rate, video_format/rate, ISO 639-2
language. Spec layout per libbluray clpi_parse.c. Best-effort —
malformed program_info leaves `streams: vec![]`, EP map keeps
working. `ClipInfo` gains a `streams: Vec<ClpiStream>` field.
2. **labels/clpi_audit.rs**. Diagnostic that walks both
`/BDMV/CLIPINF/*.clpi` (via the new program_info parser) and
`/BDMV/PLAYLIST/*.mpls`, builds a (PID → fields) merged view, and
classifies each row:
- `Match`: both sources agree (same coding_type + language)
- `ClpiOnly`: PID in CLPI but no MPLS playlist references it
(orphan stream on disc — reachable via low-level access, not via menu)
- `MplsOnly`: PID in MPLS but no CLPI lists it (would indicate a
parser bug; verified empirically that this NEVER happens)
- `Divergent`: same PID, different coding_type or language between
sources (playlist re-tagged or attribute encoding mismatch)
Surfaced via `labels-analyze` as `clpi_vs_mpls_audit: {matches,
clpi_only, mpls_only, divergent, total_pids}`. Doesn't affect the
label output — pure diagnostic.
Empirical findings on the 11-disc corpus (excl. disc-04 truncated):
- 226 matches / 0 mpls_only / 8 clpi_only / 5 divergent across 239 PIDs
- 6 of 10 non-truncated discs have CLPI-only streams (orphans)
- disc-02 (HDMV-only) is the most dramatic: 40% of its 5 streams are
CLPI-only — MPLS sees 3, CLPI sees 5
- Conclusion: CLPI is NOT truly redundant. ~5% of streams disc-wide
are CLPI-exclusive. Future work: layer CLPI as a tertiary source
below MPLS in the labels pipeline (orphan streams marked with even
lower confidence than MPLS).
LabelAnalysis gains `chapter_summary: Vec<ChapterSummary>` — one row
per .mpls file in /BDMV/PLAYLIST/, with chapter count (PlaylistMark
entries with mark_type ≤ 1) and approximate playlist duration in
seconds. Sorted by playlist filename.
Sourced from the existing crate::mpls parser (no new format work).
Useful for identifying the main feature playlist at a glance — it's
the one with the longest duration. Verified on disc-11 (Dune Pt 2):
00800.mpls correctly identified as 2h 45m 49s with 18 chapters
amid 30+ shorter playlists.
Doesn't touch the per-title `disc::DiscTitle::chapters` field which
disc::bluray.rs already populates from the same marks during disc
init — this is purely the diagnostic surface for labels-analyze.
When a framework parser (paramount, criterion, pixelogic, ctrm, dbp,
deluxe) is chosen but its label list covers only a subset of the
stream slots MPLS knows about, merge MPLS-derived entries for the
uncovered (stream_type, stream_number) slots. Framework labels keep
their richer fields (purpose=Commentary, codec_hint with "Atmos",
qualifier=Sdh); MPLS only fills slots the framework left unnamed.
Implementation:
- `fn fill_gaps_from_mpls` walks the MPLS label list, pushing any
entry whose (type, number) tuple isn't already in the framework
output. Stable sort by (type, number) groups audios before
subtitles in the merged result.
- Called from both `extract()` and `analyze()`. Skipped when the
chosen parser is itself `mpls_universal` (no gaps possible).
- `LabelAnalysis::gap_fill_added` field reports how many slots got
filled — useful diagnostic from `labels-analyze`.
- `StreamLabelType` gains `Eq + Hash` so the dedup HashSet works.
Tested via 4 new unit tests (155 of 155 labels tests passing, was
151). End-to-end on partial-yield corpus discs:
- disc-05 (Oppenheimer): pixelogic 4/5 already covered, gap_fill_added=0
- disc-11 (Dune Pt 2): pixelogic 8/11 already covered, gap_fill_added=0
(Real-world gap-fill activations are rare in the current corpus because
pixelogic already incorporates MPLS-equivalent data when matching;
the merge is defensive for less-thorough frameworks.)