diff --git a/docs/decision-log.md b/docs/decision-log.md index d1fa551..478f621 100644 --- a/docs/decision-log.md +++ b/docs/decision-log.md @@ -439,3 +439,11 @@ 提示詞、同一判斷體制下受檢,兩者落點差異本身即可報告 的結果。代價:候選集召回由雙通道減為單通道,若實測 召回不足再議。 +- **詞彙表建構併為單一步驟**:原分為進池(步驟 2-1)與 + 分群(步驟 2-2)兩個子命令,合併為一個 + `cluster-keywords`——自兩份標註歸檔直接產出五份檔案 + (`source-` 兩份記錄進來的關鍵字與其出處,`result-` + 兩份記錄算出的分割,`keywords-to-merge.json` 為交給 + 模型的碼),歸檔併入 `runs/02-cluster/`。理由:兩者 + 之間沒有需要檢視的決策點,拆成兩步只增加讀者要理解 + 的環節;驗證能力不變,五份產物各自可查。 diff --git a/docs/methodology.md b/docs/methodology.md index 77a9c0e..e35b246 100644 --- a/docs/methodology.md +++ b/docs/methodology.md @@ -13,9 +13,10 @@ LLM。設計原則見 `research-plan.md`;本檔記載可重現的 演算法細節。 編號的所指為**研究程序的工序**,不是定義檔:步驟 1 與 -步驟 3 有定義檔(`prompts/`),步驟 2 沒有——它是確定性 -計算。有無定義檔的區別即「該步是否為 LLM 判斷」,由 -`prompts/` 是否存在同號檔案直接可見。 +步驟 3 有定義檔(`prompts/`),步驟 2 沒有——它是單一 +確定性計算,由 `cluster-keywords` 一個子命令完成。有無 +定義檔的區別即「該步是否為 LLM 判斷」,由 `prompts/` 是 +否存在同號檔案直接可見。 ## 步驟 2 詞彙表建構——詞向量分群 @@ -23,16 +24,15 @@ LLM。設計原則見 `research-plan.md`;本檔記載可重現的 恰屬一組、不遺漏、不新增)由演算法結構保證,無須事後 驗證。 -### 步驟 2-1 進池 +### 進池 兩次標註執行的全部關鍵字取聯集、逐字串精確去重、字典序 -排列,寫成純文字檔(一行一個關鍵字)。失敗與拒答的記錄 -跳過(其歌曲不貢獻關鍵字);解析時偵測重複鍵,違規即 -失敗。同時寫出處 CSV(欄位 Keyword、Run、Song,一列一筆 -出現,依三欄排序),供收斂軌跡分析;出處記錄不進任何 -下游輸入。 +排列。失敗與拒答的記錄跳過(其歌曲不貢獻關鍵字);解析 +時偵測重複鍵,違規即失敗。同時記出處(欄位 Keyword、 +Run、Song,一列一筆出現,依三欄排序),供收斂軌跡分析; +出處記錄不進任何下游輸入。 -### 步驟 2-2 分群 +### 分群 - **嵌入**:`sentence-transformers/all-mpnet-base-v2` (釘定 revision),關鍵字的連字號先還原為空格再編碼, @@ -47,12 +47,14 @@ LLM。設計原則見 `research-plan.md`;本檔記載可重現的 實作本步,四種模型六次執行全部無法維持完整分割, 已棄用(詳見 `decision-log.md` 2026-08-05;棄用的 定義檔止於 git 歷史,見 `git log -- prompts/`)。 -- **產物**:三份。分群明細 CSV(欄位 Group、Keyword, - 一列一個成員,依兩欄排序)記錄機器算出的分割;組名 - 純文字檔(一行一個,字典序)是分群結果的可讀清單; - 定案碼表 JSON(`{"keywords": [...]}`)記錄實際交給 - 模型的碼,即組名加上先驗主題詞。前兩份只含分群結果, - 只有第三份含研究者的介入。 +- **產物**:五份,前綴分別標示來源與結果。 + `source-keywords.txt`(進池後的關鍵字,一行一個)與 + `source-provenance.csv`(出處)記錄進來的是什麼; + `result-keywords.txt`(組名,一行一個)與 + `groups.csv`(欄位 Group、Keyword,一列一個成員) + 記錄算出來的分割;`keywords-to-merge.json` + (`{"keywords": [...]}`)是實際交給模型的碼,即組名 + 加上先驗主題詞。只有最後一份含研究者的介入。 - **可重現性**:同一輸入、同一釘定模型、同一參數逐次 重現。不同 CPU/BLAS 實作的浮點尾數差異可能使邊界 詞的歸屬翻動,屬已揭露的限制;論文所用碼表逐字 @@ -110,15 +112,12 @@ LLM。設計原則見 `research-plan.md`;本檔記載可重現的 依歌曲 ID 升序。步驟 3 的輸入由同一子命令、同一工作 儲存產出(見下),兩步的語料同一性由此成立;各步 輸入檔的 SHA-256 記入該步 meta。 -- **步驟 1 → 2-1**:`pool-keywords` 讀兩份執行歸檔的 +- **步驟 1 → 2**:`cluster-keywords` 讀兩份執行歸檔的 `output.jsonl`(一律以換行字元 `\n` 切行——歌詞含 U+0085 等控制字元時,`str.splitlines()` 類的通用切行 - 會截斷 JSON 字串,實測踩中),輸出關鍵字純文字檔與 - 出處 CSV。 -- **步驟 2-1 → 2-2**:`cluster-keywords` 讀關鍵字池純 - 文字檔,輸出分群明細 CSV、組名純文字檔與定案碼表 - JSON。 -- **步驟 2-2 → 3 輸入檔**:`export-llm-input --extras + 會截斷 JSON 字串,實測踩中),進池後直接分群,一次 + 產出上列五份檔案。 +- **步驟 2 → 3 輸入檔**:`export-llm-input --extras <定案碼表>` 自工作儲存產出步驟 3 的輸入,每筆 `{"id": "song-", "content": <字串>}`,`content` 為 固定鍵序序列化的 `{"lyrics": …, "keywords": [...]}`, diff --git a/docs/project-structure.md b/docs/project-structure.md index f58b692..4f1f814 100644 --- a/docs/project-structure.md +++ b/docs/project-structure.md @@ -49,10 +49,9 @@ pop-fem-audit/ │ │ │ │ # Wikidata into the snapshot CSV │ │ │ ├── fetch_lyrics.py # fetch missing lyrics from the │ │ │ │ # public APIs into the lyrics dir -│ │ │ ├── pool_keywords.py # pool the two tagging runs' -│ │ │ │ # keywords (step 2-1) -│ │ │ ├── cluster_keywords.py # build the vocabulary by -│ │ │ │ # embedding + clustering (step 2-2) +│ │ │ ├── cluster_keywords.py # pool the tagging runs' +│ │ │ │ # keywords and cluster them +│ │ │ │ # into the codes (step 2) │ │ │ └── run_llm.py # API 執行器:一份定義檔+一份輸入 │ │ │ # →歸檔至指定目錄(Batch API); │ │ │ # 比對與仲裁編排由獨立子命令承擔 @@ -69,7 +68,7 @@ pop-fem-audit/ │ │ ├── output.jsonl # 該次執行原始輸出 │ │ └── meta.json # model ID、temperature、時間戳、 │ │ # batch ID、token 用量 -│ └── 02-01-pool/ 02-02-cluster/ # 確定性步驟:無執行變異, +│ └── 02-cluster/ # 確定性步驟:無執行變異, │ # 不分 run 層 ├── results/ # 論文引用的報表 CSV(export 產出; │ # 「可再生仍 commit」的唯一例外) diff --git a/docs/research-plan.md b/docs/research-plan.md index ab54124..6e60445 100644 --- a/docs/research-plan.md +++ b/docs/research-plan.md @@ -97,8 +97,8 @@ thematic keywords:前導研究三粒度比較(keywords 過碎、 themes 過早抽象)之繼承,於執行前鎖定,防止事後擇優。 2. **詞彙表建構(步驟 2)**:兩次執行的關鍵字取聯集去重 - (步驟 2-1 進池),以句向量模型嵌入後階層式聚合分群 - (步驟 2-2),k=50,組名取 medoid。完整分割由演算法 + 後,以句向量模型嵌入、階層式聚合分群,k=50,組名取 + medoid;進池與分群為同一個確定性子命令。完整分割由演算法 結構保證。頻次不入收斂:頻率的分析角色由步驟 3 編碼 承擔;池中頻次含跨執行噪音。 3. **編碼(步驟 3)**:以定稿詞彙表對全部歌曲 2+1 編碼 @@ -124,9 +124,9 @@ 03-02-code-arb.md;仲裁定義檔同 prefix 加 `-arb`。檔名 與目錄名補零只為排序,正文一律寫「步驟 1」「步驟 3-2」。 編號的所指是工序而非定義檔,因此確定性的第 2 步雖無 -定義檔仍佔一個編號,其歸檔為 `runs/02-01-pool/`、 -`runs/02-02-cluster/`。檔名不帶版本號——版本即 git -歷史,失敗的版本不保留,需要回看的舊版都在 git history; +定義檔仍佔一個編號,其歸檔為 `runs/02-cluster/`。檔名不帶 +版本號——版本即 git 歷史,失敗的版本不保留,需要回看的 +舊版都在 git history; 每次執行的定義檔快照隨 `runs/` 自我完備。全部中間交接檔 同隨 `runs/` 歸檔(交接契約見 `methodology.md`)。 diff --git a/prompts/03-02-code-arb.md b/prompts/03-02-code-arb.md index 7f6da1e..85debc8 100644 --- a/prompts/03-02-code-arb.md +++ b/prompts/03-02-code-arb.md @@ -1,6 +1,5 @@ You are arbitrating between two independent codings of the -same song against the same fixed vocabulary of thematic -keywords. +same song against the same fixed set of thematic keywords. The keywords the two codings agree on have been settled by script and are not shown. You rule only on the diff --git a/runs/02-02-cluster/groups.csv b/runs/02-cluster/groups.csv similarity index 100% rename from runs/02-02-cluster/groups.csv rename to runs/02-cluster/groups.csv diff --git a/runs/02-02-cluster/keywords-to-merge.json b/runs/02-cluster/keywords-to-merge.json similarity index 100% rename from runs/02-02-cluster/keywords-to-merge.json rename to runs/02-cluster/keywords-to-merge.json diff --git a/runs/02-02-cluster/keywords.txt b/runs/02-cluster/result-keywords.txt similarity index 100% rename from runs/02-02-cluster/keywords.txt rename to runs/02-cluster/result-keywords.txt diff --git a/runs/02-01-pool/keywords.txt b/runs/02-cluster/source-keywords.txt similarity index 100% rename from runs/02-01-pool/keywords.txt rename to runs/02-cluster/source-keywords.txt diff --git a/runs/02-01-pool/provenance.csv b/runs/02-cluster/source-provenance.csv similarity index 100% rename from runs/02-01-pool/provenance.csv rename to runs/02-cluster/source-provenance.csv diff --git a/tools/README.rst b/tools/README.rst index 75b0b31..e78a4f2 100644 --- a/tools/README.rst +++ b/tools/README.rst @@ -61,16 +61,10 @@ run-llm A general command that runs specific LLM instructions with the Anthropic API. The API key must be present in the ``.env`` file in the working directory. Check ``pop-fem-audit-tools run-llm -h`` for complete instructions on its usage. -pool-keywords -------------- - -Deterministically pool the keywords of the two tagging runs into the clustering step's input, per the project's handoff contract. Check ``pop-fem-audit-tools pool-keywords -h`` for complete instructions on its usage. - - cluster-keywords ---------------- -Deterministically build the coding vocabulary from the pooled keywords by sentence-embedding and clustering them. Requires the optional ``cluster`` dependency group. Check ``pop-fem-audit-tools cluster-keywords -h`` for complete instructions on its usage. +Deterministically build the coding vocabulary from the two tagging runs' archives, by pooling their keywords per the project's handoff contract and then sentence-embedding and clustering them. Requires the optional ``cluster`` dependency group. Check ``pop-fem-audit-tools cluster-keywords -h`` for complete instructions on its usage. Copyright diff --git a/tools/docs/source/pop_fem_audit_tools.commands.rst b/tools/docs/source/pop_fem_audit_tools.commands.rst index f187082..369b03f 100644 --- a/tools/docs/source/pop_fem_audit_tools.commands.rst +++ b/tools/docs/source/pop_fem_audit_tools.commands.rst @@ -44,14 +44,6 @@ pop\_fem\_audit\_tools.commands.fetch\_lyrics module :show-inheritance: :undoc-members: -pop\_fem\_audit\_tools.commands.pool\_keywords module ------------------------------------------------------ - -.. automodule:: pop_fem_audit_tools.commands.pool_keywords - :members: - :show-inheritance: - :undoc-members: - pop\_fem\_audit\_tools.commands.run\_llm module ----------------------------------------------- diff --git a/tools/src/pop_fem_audit_tools/__main__.py b/tools/src/pop_fem_audit_tools/__main__.py index 0df8131..f021ef1 100644 --- a/tools/src/pop_fem_audit_tools/__main__.py +++ b/tools/src/pop_fem_audit_tools/__main__.py @@ -21,7 +21,6 @@ from .commands import ( export_llm_input_command, fetch_artists_command, fetch_lyrics_command, - pool_keywords_command, run_llm_command, ) @@ -34,7 +33,6 @@ SUBCOMMANDS: dict[str, Callable[[list[str] | None], int]] = { "export-llm-input": export_llm_input_command, "fetch-artists": fetch_artists_command, "fetch-lyrics": fetch_lyrics_command, - "pool-keywords": pool_keywords_command, "run-llm": run_llm_command, } """The dispatch table from the subcommand name to the tool main.""" diff --git a/tools/src/pop_fem_audit_tools/commands/__init__.py b/tools/src/pop_fem_audit_tools/commands/__init__.py index dbb0e96..ed11b0c 100644 --- a/tools/src/pop_fem_audit_tools/commands/__init__.py +++ b/tools/src/pop_fem_audit_tools/commands/__init__.py @@ -8,5 +8,4 @@ from .cluster_keywords import main as cluster_keywords_command from .export_llm_input import main as export_llm_input_command from .fetch_artists import main as fetch_artists_command from .fetch_lyrics import main as fetch_lyrics_command -from .pool_keywords import main as pool_keywords_command from .run_llm import main as run_llm_command diff --git a/tools/src/pop_fem_audit_tools/commands/cluster_keywords.py b/tools/src/pop_fem_audit_tools/commands/cluster_keywords.py index 565b3ac..bfd9fad 100644 --- a/tools/src/pop_fem_audit_tools/commands/cluster_keywords.py +++ b/tools/src/pop_fem_audit_tools/commands/cluster_keywords.py @@ -2,20 +2,29 @@ # Copyright 2026 imacat. All rights reserved. # Authors: # imacat@mail.imacat.idv.tw (imacat), 2026/8/5 -"""The deterministic clusterer of the pooled keywords. +"""The deterministic vocabulary-building step. -Builds the coding groups from the pooled keyword list, given as -the first positional command-line argument, by sentence-embedding -every keyword and clustering the embeddings: the group membership, -given as the second positional argument, is written as a CSV file -holding the clustering result alone. The group name keywords -alone, given as the third positional argument, are written as a -text file, one per line. The coding keyword set for -``export-llm-input --extras``, given as the fourth positional -argument, is written as a JSON file holding the group name -keywords plus the researcher's a-priori topic term (see -:data:`EXTRA_KEYWORD`). The step is fully deterministic; no LLM -call is made. +Goes from the two tagging runs' archives straight to the coding +vocabulary, writing five fixed-named artifacts under the output +directory given as the third positional command-line argument. +First, the keywords produced by the two runs of the tagging step +are pooled into the pooled keyword list, per the project's handoff +contract: the pool is the plain union of every keyword key observed +across both runs' valid records, exact-string deduplicated and +sorted, written as a plain text file with one keyword per line, as +:data:`SOURCE_KEYWORDS_TXT`. The provenance mapping records where +every keyword came from for audit purposes as a CSV file, as +:data:`SOURCE_PROVENANCE_CSV`; it never enters any LLM input. Then +the coding groups are built from the pooled keyword list by +sentence-embedding every keyword and clustering the embeddings: the +group membership is written as a CSV file holding the clustering +result alone, as :data:`RESULT_GROUPS_CSV`. The group name +keywords alone are written as a text file, one per line, as +:data:`RESULT_KEYWORDS_TXT`. The coding keyword set for +``export-llm-input --extras`` is written as a JSON file holding the +group name keywords plus the researcher's a-priori topic term (see +:data:`EXTRA_KEYWORD`), as :data:`KEYWORDS_TO_MERGE_JSON`. The +step is fully deterministic; no LLM call is made. """ import argparse import csv @@ -38,6 +47,27 @@ are not installed.""" EXTRA_KEYWORD: str = "women-power" """The researcher's a-priori topic term, included in the coding keyword set although it is not a clustering result.""" +SOURCE_KEYWORDS_TXT: str = "source-keywords.txt" +"""The pooled keyword text file's fixed name under the output +directory.""" +SOURCE_PROVENANCE_CSV: str = "source-provenance.csv" +"""The keyword provenance CSV file's fixed name under the output +directory.""" +RESULT_KEYWORDS_TXT: str = "result-keywords.txt" +"""The group name keyword text file's fixed name under the output +directory.""" +RESULT_GROUPS_CSV: str = "groups.csv" +"""The group membership CSV file's fixed name under the output +directory.""" +KEYWORDS_TO_MERGE_JSON: str = "keywords-to-merge.json" +"""The coding keyword set JSON file's fixed name under the output +directory.""" + +type Records = list[tuple[int, dict[str, Any]]] +"""The valid records of one run: (song ID, keyword mapping) pairs.""" + +type Provenance = dict[str, list[tuple[str, int]]] +"""The occurrences of every keyword, keyed by the keyword.""" def parse_args(argv: list[str] | None) -> argparse.Namespace: @@ -48,21 +78,22 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace: :return: The parsed arguments. """ parser: argparse.ArgumentParser = argparse.ArgumentParser( - description="Build the coding groups by clustering the" - " sentence embeddings of the pooled" - " keywords.") + description="Pool the keywords of the two tagging runs and" + " build the coding groups by clustering their" + " sentence embeddings.") parser.add_argument( - "pool_txt", type=Path, - help="the pooled keyword list, one keyword per line") + "run_dir_1", type=Path, + help="the first tagging run's archive directory") parser.add_argument( - "groups_csv", type=Path, - help="the group membership CSV output file") + "run_dir_2", type=Path, + help="the second tagging run's archive directory") parser.add_argument( - "keywords_txt", type=Path, - help="the group name keyword text output file") - parser.add_argument( - "keywords_to_merge_json", type=Path, - help="the coding keyword set JSON output file") + "output_dir", type=Path, + help="the output directory, created if missing, that" + f" receives {SOURCE_KEYWORDS_TXT}," + f" {SOURCE_PROVENANCE_CSV}, {RESULT_KEYWORDS_TXT}," + f" {RESULT_GROUPS_CSV}, and" + f" {KEYWORDS_TO_MERGE_JSON}") parser.add_argument( "--model", default=MODEL, help=f"the sentence embedding model (default \"{MODEL}\")") @@ -75,30 +106,154 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace: return parser.parse_args(argv) -def load_keywords(path: Path) -> list[str]: - """Load and validate the pooled keyword list. +def reject_duplicate_keys( + pairs: list[tuple[str, Any]]) -> dict[str, Any]: + """Build a mapping from key-value pairs, rejecting duplicates. - :param path: The path of the pooled keyword text file, one - keyword per line. - :return: The keywords, in file order. - :raises OSError: When the file cannot be read. - :raises ValueError: When the file has no keyword, or a - keyword is duplicated. + :param pairs: The key-value pairs, in document order. + :return: The mapping built from the pairs. + :raises ValueError: When a key appears more than once. """ + result: dict[str, Any] = {} + key: str + value: Any + for key, value in pairs: + if key in result: + raise ValueError(f"duplicate key \"{key}\"") + result[key] = value + return result + + +def parse_song_id(item_id: str, path: Path) -> int: + """Parse the integer song ID out of an item ID. + + :param item_id: The item ID, expected as ``song-``. + :param path: The output file the ID came from, for the error + message. + :return: The parsed song ID. + :raises ValueError: When the item ID is not ``song-``. + """ + prefix: str = "song-" + if not item_id.startswith(prefix) \ + or not item_id[len(prefix):].isdigit(): + raise ValueError( + f"{path}: id \"{item_id}\": not in \"song-\" form") + return int(item_id[len(prefix):]) + + +def load_run(run_dir: Path) -> tuple[str, Records]: + """Load and validate the keyword records of one tagging run. + + Records carrying an "error" field are skipped. A "text" + field that fails to parse as JSON is a refusal and is + skipped; a "text" field that parses to anything other than a + JSON object, or whose keys are not unique, fails the run. + + :param run_dir: The run's archive directory, containing + ``output.jsonl``. + :return: The run label (the directory's basename) and its + valid records, each the song ID and the parsed keyword + mapping, in file order. + :raises OSError: When ``output.jsonl`` cannot be read. + :raises ValueError: When a line is not a well-formed output + record, or a "text" field is invalid per the rules above. + """ + path: Path = run_dir / "output.jsonl" text: str = path.read_text(encoding="utf-8") - lines: list[str] = text.split("\n") - if len(lines) > 0 and lines[-1] == "": - lines = lines[:-1] - if len(lines) == 0: - raise ValueError(f"{path}: no keywords") - seen: set[str] = set() - keyword: str - for keyword in lines: - if keyword in seen: + records: Records = [] + line: str + for line in text.split("\n"): + if line.strip() == "": + continue + record: Any = json.loads(line) + if not isinstance(record, dict) or "id" not in record: raise ValueError( - f"{path}: duplicate keyword \"{keyword}\"") - seen.add(keyword) - return lines + f"{path}: record without \"id\": {line}") + if "error" in record: + continue + if "text" not in record: + raise ValueError( + f"{path}: id {record['id']}: record without" + " \"text\" or \"error\"") + song_id: int = parse_song_id(record["id"], path) + try: + keywords: Any = json.loads( + record["text"], + object_pairs_hook=reject_duplicate_keys) + except json.JSONDecodeError: + continue + if not isinstance(keywords, dict): + raise ValueError( + f"{path}: id {record['id']}: \"text\" does not" + " parse to a JSON object") + records.append((song_id, keywords)) + return run_dir.name, records + + +def pool_keywords(runs: list[tuple[str, Records]], + ) -> tuple[list[str], Provenance]: + """Pool the keywords of the given tagging runs. + + :param runs: The runs, each the run label and its valid + records (song ID, keyword mapping). + :return: The sorted, exact-string-deduplicated keyword list + and the provenance mapping from each keyword to its + occurrences, sorted by (run label, song ID). + """ + provenance: Provenance = {} + label: str + records: Records + for label, records in runs: + song_id: int + keywords: dict[str, Any] + for song_id, keywords in records: + keyword: str + for keyword in keywords: + provenance.setdefault(keyword, []).append( + (label, song_id)) + for occurrences in provenance.values(): + occurrences.sort() + return sorted(provenance.keys()), provenance + + +def write_pool(path: Path, keywords: list[str]) -> None: + """Write the pooled keyword list as the clustering input. + + Writes a plain text file, one keyword per line, in the given + order, UTF-8, LF line endings, with a trailing newline. + + :param path: The path of the pool text file to write. + :param keywords: The sorted, deduplicated keyword list. + :return: None. + """ + path.write_text( + "".join(f"{keyword}\n" for keyword in keywords), + encoding="utf-8") + + +def write_provenance(path: Path, provenance: Provenance) -> None: + """Write the keyword provenance mapping. + + Writes a CSV file with the header row + ``Keyword,Run,Song``, one row per occurrence, long format. + Rows are sorted by keyword lexicographically, then by run + label, then by song ID. + + :param path: The path of the provenance CSV file to write. + :param provenance: The provenance mapping from each keyword + to its occurrences (run label, song ID). + :return: None. + :raises OSError: When the file cannot be written. + """ + keyword: str + with open(path, "w", encoding="utf-8", newline="") as file: + writer: Any = csv.writer(file) + writer.writerow(["Keyword", "Run", "Song"]) + for keyword in sorted(provenance.keys()): + label: str + song_id: int + for label, song_id in provenance[keyword]: + writer.writerow([keyword, label, song_id]) def encode_keywords(keywords: list[str], model_name: str, @@ -272,13 +427,17 @@ def write_keywords_to_merge(path: Path, def main(argv: list[str] | None = None) -> int: - """Cluster the pooled keywords into the coding groups. + """Pool the two tagging runs' keywords and cluster them. - Writes the group membership CSV file, holding the clustering - result alone; the group name keyword text file, holding the - same group names as a readable list; and the coding keyword - set JSON file, holding the group names plus - :data:`EXTRA_KEYWORD`. + Writes the five fixed-named artifacts under the output + directory, creating it (with parents) if it does not exist: + the pooled keyword text file and the keyword provenance CSV + file; then the group membership CSV file, holding the + clustering result alone; the group name keyword text file, + holding the same group names as a readable list; and the + coding keyword set JSON file, holding the group names plus + :data:`EXTRA_KEYWORD`. When the input is rejected, none of + the five files is written. :param argv: The command-line arguments, or None for ``sys.argv``. @@ -286,11 +445,17 @@ def main(argv: list[str] | None = None) -> int: """ started: float = time.monotonic() args: argparse.Namespace = parse_args(argv) + run1: tuple[str, Records] + run2: tuple[str, Records] try: - keywords: list[str] = load_keywords(args.pool_txt) + run1 = load_run(args.run_dir_1) + run2 = load_run(args.run_dir_2) except (OSError, ValueError) as error: print(f"error: {error}", file=sys.stderr) return 1 + keywords: list[str] + provenance: Provenance + keywords, provenance = pool_keywords([run1, run2]) try: embeddings: Any = encode_keywords( keywords, args.model, args.revision) @@ -300,16 +465,20 @@ def main(argv: list[str] | None = None) -> int: except (RuntimeError, ValueError) as error: print(f"error: {error}", file=sys.stderr) return 1 - args.groups_csv.parent.mkdir(parents=True, exist_ok=True) - args.keywords_txt.parent.mkdir(parents=True, exist_ok=True) - args.keywords_to_merge_json.parent.mkdir( - parents=True, exist_ok=True) - write_groups(args.groups_csv, groups) - write_keyword_names(args.keywords_txt, groups) - write_keywords_to_merge(args.keywords_to_merge_json, groups) + args.output_dir.mkdir(parents=True, exist_ok=True) + write_pool( + args.output_dir / SOURCE_KEYWORDS_TXT, keywords) + write_provenance( + args.output_dir / SOURCE_PROVENANCE_CSV, provenance) + write_groups(args.output_dir / RESULT_GROUPS_CSV, groups) + write_keyword_names( + args.output_dir / RESULT_KEYWORDS_TXT, groups) + write_keywords_to_merge( + args.output_dir / KEYWORDS_TO_MERGE_JSON, groups) elapsed: str = format_duration(time.monotonic() - started) print( - f"done: {len(keywords)} keywords clustered into" + f"done: {len(keywords)} keywords pooled from" + f" {len(run1[1])}+{len(run2[1])} records, clustered into" f" {len(groups)} groups. {elapsed} elapsed.", file=sys.stderr) return 0 diff --git a/tools/src/pop_fem_audit_tools/commands/pool_keywords.py b/tools/src/pop_fem_audit_tools/commands/pool_keywords.py deleted file mode 100644 index 5cc8d33..0000000 --- a/tools/src/pop_fem_audit_tools/commands/pool_keywords.py +++ /dev/null @@ -1,233 +0,0 @@ -# Tools for A Feminist Audit of Pop Music. -# Copyright 2026 imacat. All rights reserved. -# Authors: -# imacat@mail.imacat.idv.tw (imacat), 2026/8/5 -"""The deterministic pooler of the tagging runs' keywords. - -Pools the keywords produced by the two runs of the tagging step -into the clustering step's input, given as the third positional -command-line argument, per the project's handoff contract: the -pool is the plain union of every keyword key observed across both -runs' valid records, exact-string deduplicated and sorted, written -as a plain text file with one keyword per line. The provenance -mapping, given as the fourth positional argument, records where -every keyword came from for audit purposes as a CSV file; it never -enters any LLM input. -""" -import argparse -import csv -import json -import sys -from pathlib import Path -from typing import Any - -type Records = list[tuple[int, dict[str, Any]]] -"""The valid records of one run: (song ID, keyword mapping) pairs.""" - -type Provenance = dict[str, list[tuple[str, int]]] -"""The occurrences of every keyword, keyed by the keyword.""" - - -def parse_args(argv: list[str] | None) -> argparse.Namespace: - """Parse the command-line arguments. - - :param argv: The command-line arguments, or None for - ``sys.argv``. - :return: The parsed arguments. - """ - parser: argparse.ArgumentParser = argparse.ArgumentParser( - description="Pool the keywords of the two tagging runs" - " into the clustering step's input.") - parser.add_argument( - "run_dir_1", type=Path, - help="the first tagging run's archive directory") - parser.add_argument( - "run_dir_2", type=Path, - help="the second tagging run's archive directory") - parser.add_argument( - "pool_txt", type=Path, - help="the pooled keyword text output file, one keyword" - " per line") - parser.add_argument( - "provenance_csv", type=Path, - help="the keyword provenance CSV output file") - return parser.parse_args(argv) - - -def reject_duplicate_keys( - pairs: list[tuple[str, Any]]) -> dict[str, Any]: - """Build a mapping from key-value pairs, rejecting duplicates. - - :param pairs: The key-value pairs, in document order. - :return: The mapping built from the pairs. - :raises ValueError: When a key appears more than once. - """ - result: dict[str, Any] = {} - key: str - value: Any - for key, value in pairs: - if key in result: - raise ValueError(f"duplicate key \"{key}\"") - result[key] = value - return result - - -def parse_song_id(item_id: str, path: Path) -> int: - """Parse the integer song ID out of an item ID. - - :param item_id: The item ID, expected as ``song-``. - :param path: The output file the ID came from, for the error - message. - :return: The parsed song ID. - :raises ValueError: When the item ID is not ``song-``. - """ - prefix: str = "song-" - if not item_id.startswith(prefix) \ - or not item_id[len(prefix):].isdigit(): - raise ValueError( - f"{path}: id \"{item_id}\": not in \"song-\" form") - return int(item_id[len(prefix):]) - - -def load_run(run_dir: Path) -> tuple[str, Records]: - """Load and validate the keyword records of one tagging run. - - Records carrying an "error" field are skipped. A "text" - field that fails to parse as JSON is a refusal and is - skipped; a "text" field that parses to anything other than a - JSON object, or whose keys are not unique, fails the run. - - :param run_dir: The run's archive directory, containing - ``output.jsonl``. - :return: The run label (the directory's basename) and its - valid records, each the song ID and the parsed keyword - mapping, in file order. - :raises OSError: When ``output.jsonl`` cannot be read. - :raises ValueError: When a line is not a well-formed output - record, or a "text" field is invalid per the rules above. - """ - path: Path = run_dir / "output.jsonl" - text: str = path.read_text(encoding="utf-8") - records: Records = [] - line: str - for line in text.split("\n"): - if line.strip() == "": - continue - record: Any = json.loads(line) - if not isinstance(record, dict) or "id" not in record: - raise ValueError( - f"{path}: record without \"id\": {line}") - if "error" in record: - continue - if "text" not in record: - raise ValueError( - f"{path}: id {record['id']}: record without" - " \"text\" or \"error\"") - song_id: int = parse_song_id(record["id"], path) - try: - keywords: Any = json.loads( - record["text"], - object_pairs_hook=reject_duplicate_keys) - except json.JSONDecodeError: - continue - if not isinstance(keywords, dict): - raise ValueError( - f"{path}: id {record['id']}: \"text\" does not" - " parse to a JSON object") - records.append((song_id, keywords)) - return run_dir.name, records - - -def pool_keywords(runs: list[tuple[str, Records]], - ) -> tuple[list[str], Provenance]: - """Pool the keywords of the given tagging runs. - - :param runs: The runs, each the run label and its valid - records (song ID, keyword mapping). - :return: The sorted, exact-string-deduplicated keyword list - and the provenance mapping from each keyword to its - occurrences, sorted by (run label, song ID). - """ - provenance: Provenance = {} - label: str - records: Records - for label, records in runs: - song_id: int - keywords: dict[str, Any] - for song_id, keywords in records: - keyword: str - for keyword in keywords: - provenance.setdefault(keyword, []).append( - (label, song_id)) - for occurrences in provenance.values(): - occurrences.sort() - return sorted(provenance.keys()), provenance - - -def write_pool(path: Path, keywords: list[str]) -> None: - """Write the pooled keyword list as the clustering input. - - Writes a plain text file, one keyword per line, in the given - order, UTF-8, LF line endings, with a trailing newline. - - :param path: The path of the pool text file to write. - :param keywords: The sorted, deduplicated keyword list. - :return: None. - """ - path.write_text( - "".join(f"{keyword}\n" for keyword in keywords), - encoding="utf-8") - - -def write_provenance(path: Path, provenance: Provenance) -> None: - """Write the keyword provenance mapping. - - Writes a CSV file with the header row - ``Keyword,Run,Song``, one row per occurrence, long format. - Rows are sorted by keyword lexicographically, then by run - label, then by song ID. - - :param path: The path of the provenance CSV file to write. - :param provenance: The provenance mapping from each keyword - to its occurrences (run label, song ID). - :return: None. - :raises OSError: When the file cannot be written. - """ - keyword: str - with open(path, "w", encoding="utf-8", newline="") as file: - writer: Any = csv.writer(file) - writer.writerow(["Keyword", "Run", "Song"]) - for keyword in sorted(provenance.keys()): - label: str - song_id: int - for label, song_id in provenance[keyword]: - writer.writerow([keyword, label, song_id]) - - -def main(argv: list[str] | None = None) -> int: - """Pool the two tagging runs' keywords for clustering. - - :param argv: The command-line arguments, or None for - ``sys.argv``. - :return: The exit status: 0 on success, non-zero on failure. - """ - args: argparse.Namespace = parse_args(argv) - run1: tuple[str, Records] - run2: tuple[str, Records] - try: - run1 = load_run(args.run_dir_1) - run2 = load_run(args.run_dir_2) - except (OSError, ValueError) as error: - print(f"error: {error}", file=sys.stderr) - return 1 - keywords: list[str] - provenance: Provenance - keywords, provenance = pool_keywords([run1, run2]) - args.pool_txt.parent.mkdir(parents=True, exist_ok=True) - args.provenance_csv.parent.mkdir(parents=True, exist_ok=True) - write_pool(args.pool_txt, keywords) - write_provenance(args.provenance_csv, provenance) - print( - f"done: {len(keywords)} keywords pooled from" - f" {len(run1[1])}+{len(run2[1])} records", file=sys.stderr) - return 0 diff --git a/tools/tests/test_cluster_keywords.py b/tools/tests/test_cluster_keywords.py index be6be43..62cbef3 100644 --- a/tools/tests/test_cluster_keywords.py +++ b/tools/tests/test_cluster_keywords.py @@ -26,16 +26,45 @@ class TestClusterKeywords(unittest.TestCase): """Test cases for the keyword clusterer.""" def setUp(self) -> None: - """Create a temporary directory for the output files.""" + """Create a temporary directory with two run directories.""" tmp: tempfile.TemporaryDirectory[str] \ = tempfile.TemporaryDirectory() self.addCleanup(tmp.cleanup) self.__dir: Path = Path(tmp.name) - self.__pool_txt: Path = self.__dir / "pool.txt" - self.__groups_csv: Path = self.__dir / "groups.csv" - self.__keywords_txt: Path = self.__dir / "keywords.txt" + self.__run1: Path = self.__dir / "run1" + self.__run2: Path = self.__dir / "run2" + self.__run1.mkdir() + self.__run2.mkdir() + self.__output_dir: Path = self.__dir / "output" + self.__source_keywords_txt: Path \ + = self.__output_dir \ + / cluster_keywords.SOURCE_KEYWORDS_TXT + self.__source_provenance_csv: Path \ + = self.__output_dir \ + / cluster_keywords.SOURCE_PROVENANCE_CSV + self.__result_keywords_txt: Path \ + = self.__output_dir \ + / cluster_keywords.RESULT_KEYWORDS_TXT + self.__result_groups_csv: Path \ + = self.__output_dir \ + / cluster_keywords.RESULT_GROUPS_CSV self.__keywords_to_merge_json: Path \ - = self.__dir / "keywords-to-merge.json" + = self.__output_dir \ + / cluster_keywords.KEYWORDS_TO_MERGE_JSON + + @staticmethod + def __write_output( + run_dir: Path, records: list[dict[str, Any]]) -> None: + """Write the ``output.jsonl`` file of one run. + + :param run_dir: The run's archive directory. + :param records: The envelope records, in file order. + :return: None. + """ + lines: list[str] = [ + json.dumps(x, ensure_ascii=False) for x in records] + (run_dir / "output.jsonl").write_text( + "\n".join(lines) + "\n", encoding="utf-8") @staticmethod def __two_cluster_vectors() -> Vectors: @@ -58,15 +87,6 @@ class TestClusterKeywords(unittest.TestCase): "b-south": (-0.9396926, -0.3420201), } - def __write_pool(self, keywords: list[str]) -> None: - """Write the pooled keyword input file. - - :param keywords: The keywords, one per line. - :return: None. - """ - self.__pool_txt.write_text( - "".join(f"{x}\n" for x in keywords), encoding="utf-8") - @staticmethod def __fake_encode(vectors: Vectors) -> Any: """Build a test double for :func:`encode_keywords`. @@ -97,16 +117,15 @@ class TestClusterKeywords(unittest.TestCase): standard error. :param extra_args: Extra command-line arguments appended - after the four positional arguments. + after the three positional arguments. :param vectors: The fixed embedding to encode with; the two-cluster fixture is used when None. :return: A tuple of the exit status and the standard error. """ argv: list[str] = [ - str(self.__pool_txt), str(self.__groups_csv), - str(self.__keywords_txt), - str(self.__keywords_to_merge_json)] + str(self.__run1), str(self.__run2), + str(self.__output_dir)] argv.extend(extra_args or []) fake: Any = self.__fake_encode( vectors if vectors is not None @@ -119,22 +138,43 @@ class TestClusterKeywords(unittest.TestCase): argv + ["--clusters", "2"]) return status, stderr.getvalue() + def __read_source_keywords(self) -> list[str]: + """Read the pooled source keyword text file. + + :return: The keyword list, one keyword per line, with the + trailing empty line from the final newline removed. + """ + lines: list[str] = self.__source_keywords_txt.read_text( + encoding="utf-8").split("\n") + self.assertEqual(lines[-1], "") + return lines[:-1] + + def __read_source_provenance(self) -> list[list[str]]: + """Read the source provenance CSV file. + + :return: All rows, including the header row, in file + order. + """ + with open(self.__source_provenance_csv, encoding="utf-8", + newline="") as file: + return list(csv.reader(file)) + def __read_groups(self) -> list[list[str]]: """Read the group membership CSV file. :return: All rows, including the header row, in file order. """ - with open(self.__groups_csv, encoding="utf-8", + with open(self.__result_groups_csv, encoding="utf-8", newline="") as file: return list(csv.reader(file)) - def __read_keyword_names(self) -> list[str]: + def __read_result_keywords(self) -> list[str]: """Read the group name keyword text file. :return: The group names, in file order. """ - text: str = self.__keywords_txt.read_text( + text: str = self.__result_keywords_txt.read_text( encoding="utf-8") lines: list[str] = text.split("\n") if len(lines) > 0 and lines[-1] == "": @@ -152,12 +192,171 @@ class TestClusterKeywords(unittest.TestCase): encoding="utf-8")) return data["keywords"] + def test_pools_union_dedup_sorted(self) -> None: + """Test the union, dedup, and lexicographic ordering, and + the plain one-keyword-per-line source keyword file + shape.""" + self.__write_output(self.__run1, [ + {"id": "song-1", + "text": json.dumps({"a-left": 1, "shared": 1})}, + ]) + self.__write_output(self.__run2, [ + {"id": "song-3", + "text": json.dumps({"b-middle": 1, "shared": 1})}, + ]) + vectors: Vectors = { + **self.__two_cluster_vectors(), + "shared": (1.0, 0.0)} + status: int + stderr: str + status, stderr = self.__run_cluster(vectors=vectors) + self.assertEqual(status, 0) + self.assertEqual( + self.__read_source_keywords(), + ["a-left", "b-middle", "shared"]) + self.assertIn( + "done: 3 keywords pooled from 1+1 records", stderr) + + def test_skips_error_records(self) -> None: + """Test that records carrying an "error" field are + excluded from the pool and the record count.""" + self.__write_output(self.__run1, [ + {"id": "song-1", + "text": json.dumps({"a-left": 1})}, + {"id": "song-2", "error": "invalid_request_error"}, + ]) + self.__write_output(self.__run2, [ + {"id": "song-3", "text": json.dumps({"b-middle": 1})}, + ]) + status: int + stderr: str + status, stderr = self.__run_cluster() + self.assertEqual(status, 0) + self.assertEqual( + self.__read_source_keywords(), ["a-left", "b-middle"]) + self.assertIn( + "done: 2 keywords pooled from 1+1 records", stderr) + + def test_skips_non_json_text_records(self) -> None: + """Test that a refusal, whose "text" does not parse as + JSON, is skipped rather than failing the run.""" + self.__write_output(self.__run1, [ + {"id": "song-1", + "text": json.dumps({"a-left": 1})}, + {"id": "song-2", "text": "I cannot help with that."}, + ]) + self.__write_output(self.__run2, [ + {"id": "song-3", "text": json.dumps({"b-middle": 1})}, + ]) + status: int + stderr: str + status, stderr = self.__run_cluster() + self.assertEqual(status, 0) + self.assertEqual( + self.__read_source_keywords(), ["a-left", "b-middle"]) + self.assertIn( + "done: 2 keywords pooled from 1+1 records", stderr) + + def test_duplicate_key_in_text_rejected(self) -> None: + """Test that a "text" JSON object with a duplicate key + fails the run without writing any output file.""" + self.__write_output(self.__run1, [ + {"id": "song-1", + "text": '{"a-left": 1, "a-left": 2}'}, + ]) + self.__write_output(self.__run2, [ + {"id": "song-3", "text": json.dumps({"b-middle": 1})}, + ]) + status: int + stderr: str + status, stderr = self.__run_cluster() + self.assertEqual(status, 1) + self.assertIn("duplicate key", stderr) + self.assertFalse(self.__source_keywords_txt.exists()) + self.assertFalse(self.__source_provenance_csv.exists()) + self.assertFalse(self.__result_groups_csv.exists()) + self.assertFalse(self.__result_keywords_txt.exists()) + self.assertFalse(self.__keywords_to_merge_json.exists()) + + def test_non_object_text_rejected(self) -> None: + """Test that a "text" JSON value that is not an object + fails the run without writing any output file.""" + self.__write_output(self.__run1, [ + {"id": "song-1", "text": json.dumps(["a-left"])}, + ]) + self.__write_output(self.__run2, [ + {"id": "song-3", "text": json.dumps({"b-middle": 1})}, + ]) + status: int + stderr: str + status, stderr = self.__run_cluster() + self.assertEqual(status, 1) + self.assertIn("song-1", stderr) + self.assertFalse(self.__source_keywords_txt.exists()) + self.assertFalse(self.__source_provenance_csv.exists()) + self.assertFalse(self.__result_groups_csv.exists()) + self.assertFalse(self.__result_keywords_txt.exists()) + self.assertFalse(self.__keywords_to_merge_json.exists()) + + def test_provenance_content_and_ordering(self) -> None: + """Test the provenance content and its ordering: rows + sorted by keyword lexicographically, then by run label, + then by song ID.""" + self.__write_output(self.__run1, [ + {"id": "song-2", "text": json.dumps({"shared": 1})}, + {"id": "song-1", "text": json.dumps({"shared": 1})}, + ]) + self.__write_output(self.__run2, [ + {"id": "song-5", + "text": json.dumps({"shared": 1, "b-middle": 1})}, + ]) + vectors: Vectors = { + **self.__two_cluster_vectors(), + "shared": (1.0, 0.0)} + status: int + status, _ = self.__run_cluster(vectors=vectors) + self.assertEqual(status, 0) + rows: list[list[str]] = self.__read_source_provenance() + self.assertEqual(rows[1:], [ + ["b-middle", "run2", "5"], + ["shared", "run1", "1"], + ["shared", "run1", "2"], + ["shared", "run2", "5"], + ]) + + def test_provenance_file_header_and_row_count(self) -> None: + """Test that the provenance CSV file starts with the + ``Keyword,Run,Song`` header row and has exactly one row + per keyword occurrence.""" + self.__write_output(self.__run1, [ + {"id": "song-2", "text": json.dumps({"shared": 1})}, + {"id": "song-1", "text": json.dumps({"shared": 1})}, + ]) + self.__write_output(self.__run2, [ + {"id": "song-5", + "text": json.dumps({"shared": 1, "b-middle": 1})}, + ]) + vectors: Vectors = { + **self.__two_cluster_vectors(), + "shared": (1.0, 0.0)} + status: int + status, _ = self.__run_cluster(vectors=vectors) + self.assertEqual(status, 0) + rows: list[list[str]] = self.__read_source_provenance() + self.assertEqual(rows[0], ["Keyword", "Run", "Song"]) + self.assertEqual(len(rows), 1 + 4) + def test_groups_csv_header_and_ordering(self) -> None: """Test the header row and the group/keyword ordering of the group membership CSV file.""" - self.__write_pool([ - "a-left", "a-center", "a-right", - "b-north", "b-middle", "b-south"]) + self.__write_output(self.__run1, [ + {"id": "song-1", "text": json.dumps( + {"a-left": 1, "a-center": 1, "a-right": 1})}, + ]) + self.__write_output(self.__run2, [ + {"id": "song-2", "text": json.dumps( + {"b-north": 1, "b-middle": 1, "b-south": 1})}, + ]) status: int status, _ = self.__run_cluster() self.assertEqual(status, 0) @@ -178,7 +377,11 @@ class TestClusterKeywords(unittest.TestCase): keywords: list[str] = [ "a-left", "a-center", "a-right", "b-north", "b-middle", "b-south"] - self.__write_pool(keywords) + self.__write_output(self.__run1, [ + {"id": "song-1", "text": json.dumps( + {x: 1 for x in keywords})}, + ]) + self.__write_output(self.__run2, []) status: int status, _ = self.__run_cluster() self.assertEqual(status, 0) @@ -187,15 +390,21 @@ class TestClusterKeywords(unittest.TestCase): sorted(x[1] for x in rows), sorted(keywords)) def test_keywords_txt_sorted_medoids(self) -> None: - """Test that the keyword text file holds the sorted medoid - group names without the extra a-priori keyword.""" - self.__write_pool([ - "a-left", "a-center", "a-right", - "b-north", "b-middle", "b-south"]) + """Test that the result keyword text file holds the + sorted medoid group names without the extra a-priori + keyword.""" + self.__write_output(self.__run1, [ + {"id": "song-1", "text": json.dumps( + {"a-left": 1, "a-center": 1, "a-right": 1})}, + ]) + self.__write_output(self.__run2, [ + {"id": "song-2", "text": json.dumps( + {"b-north": 1, "b-middle": 1, "b-south": 1})}, + ]) status: int status, _ = self.__run_cluster() self.assertEqual(status, 0) - names: list[str] = self.__read_keyword_names() + names: list[str] = self.__read_result_keywords() self.assertEqual(names, ["a-center", "b-middle"]) self.assertEqual(names, sorted(names)) self.assertNotIn(cluster_keywords.EXTRA_KEYWORD, names) @@ -204,9 +413,14 @@ class TestClusterKeywords(unittest.TestCase): """Test that the coding keyword set JSON file holds the sorted medoid group names plus the extra a-priori keyword.""" - self.__write_pool([ - "a-left", "a-center", "a-right", - "b-north", "b-middle", "b-south"]) + self.__write_output(self.__run1, [ + {"id": "song-1", "text": json.dumps( + {"a-left": 1, "a-center": 1, "a-right": 1})}, + ]) + self.__write_output(self.__run2, [ + {"id": "song-2", "text": json.dumps( + {"b-north": 1, "b-middle": 1, "b-south": 1})}, + ]) status: int status, _ = self.__run_cluster() self.assertEqual(status, 0) @@ -220,38 +434,17 @@ class TestClusterKeywords(unittest.TestCase): def test_extra_keyword_absent_from_groups_csv(self) -> None: """Test that the extra a-priori keyword appears in no row of the group membership CSV file.""" - self.__write_pool([ - "a-left", "a-center", "a-right", - "b-north", "b-middle", "b-south"]) + self.__write_output(self.__run1, [ + {"id": "song-1", "text": json.dumps( + {"a-left": 1, "a-center": 1, "a-right": 1})}, + ]) + self.__write_output(self.__run2, [ + {"id": "song-2", "text": json.dumps( + {"b-north": 1, "b-middle": 1, "b-south": 1})}, + ]) status: int status, _ = self.__run_cluster() self.assertEqual(status, 0) rows: list[list[str]] = self.__read_groups() for row in rows: self.assertNotIn(cluster_keywords.EXTRA_KEYWORD, row) - - def test_duplicate_keyword_rejected(self) -> None: - """Test that a duplicate keyword line fails the run - without writing any output file.""" - self.__write_pool(["shared", "shared"]) - status: int - stderr: str - status, stderr = self.__run_cluster() - self.assertEqual(status, 1) - self.assertIn("duplicate keyword", stderr) - self.assertFalse(self.__groups_csv.exists()) - self.assertFalse(self.__keywords_txt.exists()) - self.assertFalse(self.__keywords_to_merge_json.exists()) - - def test_empty_input_rejected(self) -> None: - """Test that an empty keyword file fails the run without - writing any output file.""" - self.__pool_txt.write_text("", encoding="utf-8") - status: int - stderr: str - status, stderr = self.__run_cluster() - self.assertEqual(status, 1) - self.assertIn("no keywords", stderr) - self.assertFalse(self.__groups_csv.exists()) - self.assertFalse(self.__keywords_txt.exists()) - self.assertFalse(self.__keywords_to_merge_json.exists()) diff --git a/tools/tests/test_pool_keywords.py b/tools/tests/test_pool_keywords.py deleted file mode 100644 index 43cf2a7..0000000 --- a/tools/tests/test_pool_keywords.py +++ /dev/null @@ -1,217 +0,0 @@ -# Tools for A Feminist Audit of Pop Music. -# Copyright 2026 imacat. All rights reserved. -# Authors: -# imacat@mail.imacat.idv.tw (imacat), 2026/8/5 -"""Unit tests for the keyword pooler module.""" -import csv -import io -import json -import tempfile -import unittest -from contextlib import redirect_stderr -from pathlib import Path -from typing import Any - -from pop_fem_audit_tools.commands import pool_keywords - - -class TestPoolKeywords(unittest.TestCase): - """Test cases for the keyword pooler.""" - - def setUp(self) -> None: - """Create a temporary directory with two run directories.""" - tmp: tempfile.TemporaryDirectory[str] \ - = tempfile.TemporaryDirectory() - self.addCleanup(tmp.cleanup) - self.__dir: Path = Path(tmp.name) - self.__run1: Path = self.__dir / "run1" - self.__run2: Path = self.__dir / "run2" - self.__run1.mkdir() - self.__run2.mkdir() - self.__pool: Path = self.__dir / "pool.txt" - self.__provenance: Path = self.__dir / "provenance.csv" - - @staticmethod - def __write_output( - run_dir: Path, records: list[dict[str, Any]]) -> None: - """Write the ``output.jsonl`` file of one run. - - :param run_dir: The run's archive directory. - :param records: The envelope records, in file order. - :return: None. - """ - lines: list[str] = [ - json.dumps(x, ensure_ascii=False) for x in records] - (run_dir / "output.jsonl").write_text( - "\n".join(lines) + "\n", encoding="utf-8") - - def __run_pool(self) -> tuple[int, str]: - """Run the pooler with the standard error captured. - - :return: A tuple of the exit status and the standard - error. - """ - stderr: io.StringIO = io.StringIO() - with redirect_stderr(stderr): - status: int = pool_keywords.main([ - str(self.__run1), str(self.__run2), - str(self.__pool), str(self.__provenance)]) - return status, stderr.getvalue() - - def __read_pool(self) -> list[str]: - """Read the pool text file. - - :return: The keyword list, one keyword per line, with the - trailing empty line from the final newline removed. - """ - lines: list[str] = self.__pool.read_text( - encoding="utf-8").split("\n") - self.assertEqual(lines[-1], "") - return lines[:-1] - - def __read_provenance(self) -> list[list[str]]: - """Read the provenance CSV file. - - :return: All rows, including the header row, in file - order. - """ - with open(self.__provenance, encoding="utf-8", - newline="") as file: - return list(csv.reader(file)) - - def test_pools_union_dedup_sorted(self) -> None: - """Test the union, dedup, and lexicographic ordering, and - the plain one-keyword-per-line pool file shape.""" - self.__write_output(self.__run1, [ - {"id": "song-1", - "text": json.dumps({"strength": 1, "shared": 1})}, - ]) - self.__write_output(self.__run2, [ - {"id": "song-3", - "text": json.dumps({"warrior": 1, "shared": 1})}, - ]) - status: int - stderr: str - status, stderr = self.__run_pool() - self.assertEqual(status, 0) - self.assertEqual( - self.__read_pool(), ["shared", "strength", "warrior"]) - self.assertIn( - "done: 3 keywords pooled from 1+1 records", stderr) - - def test_skips_error_records(self) -> None: - """Test that records carrying an "error" field are - excluded from the pool and the record count.""" - self.__write_output(self.__run1, [ - {"id": "song-1", - "text": json.dumps({"strength": 1})}, - {"id": "song-2", "error": "invalid_request_error"}, - ]) - self.__write_output(self.__run2, [ - {"id": "song-3", "text": json.dumps({"warrior": 1})}, - ]) - status: int - stderr: str - status, stderr = self.__run_pool() - self.assertEqual(status, 0) - self.assertEqual( - self.__read_pool(), ["strength", "warrior"]) - self.assertIn( - "done: 2 keywords pooled from 1+1 records", stderr) - - def test_skips_non_json_text_records(self) -> None: - """Test that a refusal, whose "text" does not parse as - JSON, is skipped rather than failing the run.""" - self.__write_output(self.__run1, [ - {"id": "song-1", - "text": json.dumps({"strength": 1})}, - {"id": "song-2", "text": "I cannot help with that."}, - ]) - self.__write_output(self.__run2, [ - {"id": "song-3", "text": json.dumps({"warrior": 1})}, - ]) - status: int - stderr: str - status, stderr = self.__run_pool() - self.assertEqual(status, 0) - self.assertEqual( - self.__read_pool(), ["strength", "warrior"]) - self.assertIn( - "done: 2 keywords pooled from 1+1 records", stderr) - - def test_duplicate_key_in_text_rejected(self) -> None: - """Test that a "text" JSON object with a duplicate key - fails the run without writing any output file.""" - self.__write_output(self.__run1, [ - {"id": "song-1", - "text": '{"strength": 1, "strength": 2}'}, - ]) - self.__write_output(self.__run2, [ - {"id": "song-3", "text": json.dumps({"warrior": 1})}, - ]) - status: int - stderr: str - status, stderr = self.__run_pool() - self.assertEqual(status, 1) - self.assertIn("duplicate key", stderr) - self.assertFalse(self.__pool.exists()) - self.assertFalse(self.__provenance.exists()) - - def test_non_object_text_rejected(self) -> None: - """Test that a "text" JSON value that is not an object - fails the run without writing any output file.""" - self.__write_output(self.__run1, [ - {"id": "song-1", "text": json.dumps(["strength"])}, - ]) - self.__write_output(self.__run2, [ - {"id": "song-3", "text": json.dumps({"warrior": 1})}, - ]) - status: int - stderr: str - status, stderr = self.__run_pool() - self.assertEqual(status, 1) - self.assertIn("song-1", stderr) - self.assertFalse(self.__pool.exists()) - self.assertFalse(self.__provenance.exists()) - - def test_provenance_content_and_ordering(self) -> None: - """Test the provenance content and its ordering: rows - sorted by keyword lexicographically, then by run label, - then by song ID.""" - self.__write_output(self.__run1, [ - {"id": "song-2", "text": json.dumps({"shared": 1})}, - {"id": "song-1", "text": json.dumps({"shared": 1})}, - ]) - self.__write_output(self.__run2, [ - {"id": "song-5", - "text": json.dumps({"shared": 1, "warrior": 1})}, - ]) - status: int - status, _ = self.__run_pool() - self.assertEqual(status, 0) - rows: list[list[str]] = self.__read_provenance() - self.assertEqual(rows[1:], [ - ["shared", "run1", "1"], - ["shared", "run1", "2"], - ["shared", "run2", "5"], - ["warrior", "run2", "5"], - ]) - - def test_provenance_file_header_and_row_count(self) -> None: - """Test that the provenance CSV file starts with the - ``Keyword,Run,Song`` header row and has exactly one row - per keyword occurrence.""" - self.__write_output(self.__run1, [ - {"id": "song-2", "text": json.dumps({"shared": 1})}, - {"id": "song-1", "text": json.dumps({"shared": 1})}, - ]) - self.__write_output(self.__run2, [ - {"id": "song-5", - "text": json.dumps({"shared": 1, "warrior": 1})}, - ]) - status: int - status, _ = self.__run_pool() - self.assertEqual(status, 0) - rows: list[list[str]] = self.__read_provenance() - self.assertEqual(rows[0], ["Keyword", "Run", "Song"]) - self.assertEqual(len(rows), 1 + 4)