Skip to content

[i18n] Polish eval set: frozen intents.pl.jsonl + speaker-holdout gap (issue #79 remainder) #135

Description

@grzanka

Background

Spun out of #79 ("Define sentence/speaker counts for the Polish eval set"), which proposed three
independently-scoped tracks. Two of the three are done — Track 1 (human recordings,
eval/RECORDING.pl.md, issue #84) and Track 3 (TTS-generated batch, PR #97, later exceeded by
#106/PR #105's Chatterbox batches). This issue tracks the two genuinely open items left over.

1. Track 2 — the frozen Polish text eval set (eval/intents.pl.jsonl)

Not started — the file doesn't exist yet. Per #79's original proposal:

This is the actual missing deliverable from #79 — the PL equivalent of eval/intents.jsonl,
CI-gated the way the English set already is.

2. Track 1's unmet speaker-count gap

eval/RECORDING.pl.md's 50 sentences exist as audio in eval/audio/lg/, but — confirmed via
docs/tts-chatterbox-pl-clone.md:26all 50 clips are one speaker (the project owner). #84
asked for 2–3 volunteers; only the bootstrap speaker actually delivered.

This is the same shape of gap #20 already tracks for the English set (a tuning-set vs. holdout-set
distinction — rules can't be validated as generalizing past one speaker). Not blocking Track 2
above, and not urgent on its own — recording is human-dependent, not code-dependent, same as #20's
current status. Listed here so it isn't lost, not as a call to action:

Related

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions