Read the vendor forced-subtitle field as the enumeration it is

One vendor's playlists.xml carries a per-subtitle-slot cell that looks
like a boolean, and it was parsed as one: value 1 meant forced, anything
else meant not forced. Across every image in the corpus that uses this
format the cell takes four values, and 1 is not the forced one.

Decoding three of those discs and counting every PGS display set:

  * 1 marks a FULL dialogue track that additionally contains some
    forced-narrative signs. All nine cells bearing it on one disc are
    full tracks of 949-1411 display sets; all seven on another are full
    tracks of 1602-1651. Neither disc has a small track among them.
  * 2 and 3 mark a DEDICATED forced-narrative track, in its own trailing
    stream slot, duplicating a language that already holds a full track.
    The two 2 slots measured are 15 and 10 display sets with every one
    flagged forced; the four 3 slots are 7, 14, 23 and 59 against
    1216-2655 on the tracks they duplicate.

So the old reading was wrong in both directions — it flagged full
dialogue tracks forced, which is how one language came to present as two
identical full subtitle tracks with one of them marked forced, and it
threw away the cells naming the real forced tracks.

Content could not have corrected this afterwards. Clearing a wrong
forced label needs a disc whose authoring sets forced_on_flag, and on
the measured disc carrying four genuine forced tracks not one display
set anywhere sets it — there, the vendor cell is the only evidence there
is. The classification is now explicit: only a dedicated forced slot
earns the flag, an unrecognised value never does, and the
contains-forced-signs value is dropped rather than weakened into a
forced label, since a wrong forced flag on a full dialogue track is the
user-visible defect while a missing hint costs nothing.

Four of the crate's own tests had been asserting the boolean reading;
their subject was positional alignment, so they keep it and now use a
real forced value. The other two parsers that emit a forced qualifier
from vendor metadata were audited and are structurally immune — in both,
the forced marker names a slot of its own rather than hanging off a full
track's entry, so the failure has no encoding there — and each is now
pinned by a test saying so.
This commit is contained in:
Matthew Jackson
2026-08-02 18:21:32 -07:00
parent 0d4aab99df
commit 6d2ff4d1fc
4 changed files with 278 additions and 16 deletions
+28
View File
@@ -4,6 +4,34 @@
### Fixed
- **A vendor label's forced-subtitle field was read as a boolean when it is an
enumeration, mislabelling full dialogue tracks and discarding the real forced
tracks.** One vendor's `playlists.xml` carries a per-subtitle-slot cell that
looks like a flag; the parser treated the value `1` as "this track is forced"
and every other value as "not forced". Measured across every image in the
corpus that uses this format — seven distinct discs — the cell takes four
values, and `1` is not the forced one: it marks a FULL dialogue track that
additionally contains some forced-narrative signs. Decoding three of those
discs and counting every PGS display set: all nine tracks bearing `1` on one
disc are full tracks of 949-1411 display sets, all seven on another are full
tracks of 1602-1651, and neither disc's `1` tracks are anything but full —
which is how a language ended up presenting as two identical full subtitle
tracks with one of them flagged forced. The values that DO name a dedicated
forced-narrative track, `2` and `3`, were being thrown away: they take their
own trailing stream slots, one per localized language, and measure 7 to 59
display sets against the 1216-2655 of the full tracks they duplicate. So the
reading was wrong in both directions. The cell is now classified as the enumeration it is; only a
dedicated forced slot earns the flag, an unrecognised value never does, and
the "contains forced signs" value is dropped rather than weakened into a
forced label. This is not something content could have corrected afterwards:
clearing a wrong forced label requires a disc whose authoring sets
`forced_on_flag`, and measured discs in this format do not set it — so on
those discs the vendor cell was, and remains, the only evidence there is.
Four of the crate's own tests had been asserting the boolean reading. The
other two parsers that emit a forced qualifier from vendor metadata were
audited and are structurally immune — in both, the forced marker names a slot
of its own rather than hanging off a full track's entry — and are now pinned
by tests saying so.
- **Content-based forced-subtitle detection never observed anything on a
feature-length disc.** The PGS probe spent its entire 256 MiB budget on the
first sectors of a title, where a feature has no subtitles at all — it hit