Drop the source-provenance artifact from cluster-keywords
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -495,6 +495,13 @@
|
|||||||
代價是少數組名不副實,依前條據實報告。實驗歸檔不入
|
代價是少數組名不副實,依前條據實報告。實驗歸檔不入
|
||||||
repo,費用 $0.09(sonnet)與 $0.43(fable)記入
|
repo,費用 $0.09(sonnet)與 $0.43(fable)記入
|
||||||
`run-costs.md`。
|
`run-costs.md`。
|
||||||
|
- **刪除 `source-provenance.csv`,`cluster-keywords` 的產物
|
||||||
|
由六份減為五份**:該檔是兩份標註執行歸檔 `output.jsonl`
|
||||||
|
的攤平視圖,14,035 列 Keyword,Run,Song,未帶入任何庫裏
|
||||||
|
尚無的資訊,亦無下游消費者——軌跡對映診斷直接讀那兩份
|
||||||
|
歸檔即可。稽核軌跡的慣例保留的是原始資料與實質的分析
|
||||||
|
產物;「哪個碼出自哪份文件」在質性分析軟體裏本是一個
|
||||||
|
查詢,不是一份交付物。
|
||||||
- **仲裁定義檔改名為 `03-02-arbitration.md`**,2026-08-04
|
- **仲裁定義檔改名為 `03-02-arbitration.md`**,2026-08-04
|
||||||
訂下的「仲裁檔同 prefix 加 `-arb`」規則隨之作廢。理由:
|
訂下的「仲裁檔同 prefix 加 `-arb`」規則隨之作廢。理由:
|
||||||
`code` 與編號重複——`03-02` 已表明是步驟 3(編碼)之下
|
`code` 與編號重複——`03-02` 已表明是步驟 3(編碼)之下
|
||||||
|
|||||||
+12
-15
@@ -28,9 +28,7 @@ LLM。設計原則見 `research-plan.md`;本檔記載可重現的
|
|||||||
|
|
||||||
兩次標註執行的全部關鍵字取聯集、逐字串精確去重、字典序
|
兩次標註執行的全部關鍵字取聯集、逐字串精確去重、字典序
|
||||||
排列。失敗與拒答的記錄跳過(其歌曲不貢獻關鍵字);解析
|
排列。失敗與拒答的記錄跳過(其歌曲不貢獻關鍵字);解析
|
||||||
時偵測重複鍵,違規即失敗。同時記出處(欄位 Keyword、
|
時偵測重複鍵,違規即失敗。
|
||||||
Run、Song,一列一筆出現,依三欄排序),供收斂軌跡分析;
|
|
||||||
出處記錄不進任何下游輸入。
|
|
||||||
|
|
||||||
### 分群
|
### 分群
|
||||||
|
|
||||||
@@ -59,14 +57,13 @@ Run、Song,一列一筆出現,依三欄排序),供收斂軌跡分析;
|
|||||||
實作本步,四種模型六次執行全部無法維持完整分割,
|
實作本步,四種模型六次執行全部無法維持完整分割,
|
||||||
已棄用(詳見 `decision-log.md` 2026-08-05;棄用的
|
已棄用(詳見 `decision-log.md` 2026-08-05;棄用的
|
||||||
定義檔止於 git 歷史,見 `git log -- prompts/`)。
|
定義檔止於 git 歷史,見 `git log -- prompts/`)。
|
||||||
- **產物**:六份,前綴分別標示來源與結果。
|
- **產物**:五份,前綴分別標示來源與結果。
|
||||||
`source-keywords.txt`(進池後的關鍵字,一行一個)與
|
`source-keywords.txt`(進池後的關鍵字,一行一個)
|
||||||
`source-provenance.csv`(出處)記錄進來的是什麼;
|
記錄進來的是什麼;`result-keywords.txt`(組名,一行
|
||||||
`result-keywords.txt`(組名,一行一個)與
|
一個)與 `groups.csv`(欄位 Group、Keyword,一列一個
|
||||||
`groups.csv`(欄位 Group、Keyword,一列一個成員)
|
成員)記錄算出來的分割;`keywords-to-merge.json`
|
||||||
記錄算出來的分割;`keywords-to-merge.json`
|
|
||||||
(`{"keywords": [...]}`)是實際交給模型的碼,即組名
|
(`{"keywords": [...]}`)是實際交給模型的碼,即組名
|
||||||
加上先驗主題詞——六份中只有這一份含研究者的介入。
|
加上先驗主題詞——五份中只有這一份含研究者的介入。
|
||||||
`meta.json` 記錄執行本身:進池的兩份執行歸檔與其有效
|
`meta.json` 記錄執行本身:進池的兩份執行歸檔與其有效
|
||||||
筆數、嵌入模型與釘定 revision、分群參數與組數、外加
|
筆數、嵌入模型與釘定 revision、分群參數與組數、外加
|
||||||
的先驗詞、關鍵字總數,以及產生數字的套件版本。凡命令
|
的先驗詞、關鍵字總數,以及產生數字的套件版本。凡命令
|
||||||
@@ -117,10 +114,10 @@ Run、Song,一列一筆出現,依三欄排序),供收斂軌跡分析;
|
|||||||
|
|
||||||
## 軌跡對映(診斷用)
|
## 軌跡對映(診斷用)
|
||||||
|
|
||||||
沿收斂軌跡的機械對映:原始關鍵字 →(出處記錄)歌曲、
|
沿收斂軌跡的機械對映:原始關鍵字 →(兩份標註執行歸檔的
|
||||||
原始關鍵字 →(分群)組,純程式查表,決定性。以其結果
|
`output.jsonl`)歌曲、原始關鍵字 →(分群)組,純程式查表,
|
||||||
與步驟 3 直接編碼的差異率作為「收斂軌跡扭曲」的診斷量,
|
決定性。以其結果與步驟 3 直接編碼的差異率作為「收斂軌跡
|
||||||
不作主結果。
|
扭曲」的診斷量,不作主結果。
|
||||||
|
|
||||||
## 全管線的交接契約
|
## 全管線的交接契約
|
||||||
|
|
||||||
@@ -136,7 +133,7 @@ Run、Song,一列一筆出現,依三欄排序),供收斂軌跡分析;
|
|||||||
`output.jsonl`(一律以換行字元 `\n` 切行——歌詞含
|
`output.jsonl`(一律以換行字元 `\n` 切行——歌詞含
|
||||||
U+0085 等控制字元時,`str.splitlines()` 類的通用切行
|
U+0085 等控制字元時,`str.splitlines()` 類的通用切行
|
||||||
會截斷 JSON 字串,實測踩中),進池後直接分群,一次
|
會截斷 JSON 字串,實測踩中),進池後直接分群,一次
|
||||||
產出上列六份檔案。
|
產出上列五份檔案。
|
||||||
- **步驟 2 → 3 輸入檔**:`export-llm-input --extras
|
- **步驟 2 → 3 輸入檔**:`export-llm-input --extras
|
||||||
<定案碼表>` 自工作儲存產出步驟 3 的輸入,每筆
|
<定案碼表>` 自工作儲存產出步驟 3 的輸入,每筆
|
||||||
`{"id": "song-<ID>", "content": <字串>}`,`content` 為
|
`{"id": "song-<ID>", "content": <字串>}`,`content` 為
|
||||||
|
|||||||
File diff suppressed because it is too large
Load Diff
@@ -5,17 +5,15 @@
|
|||||||
"""The deterministic vocabulary-building step.
|
"""The deterministic vocabulary-building step.
|
||||||
|
|
||||||
Goes from the two tagging runs' archives straight to the coding
|
Goes from the two tagging runs' archives straight to the coding
|
||||||
vocabulary, writing six fixed-named artifacts under the output
|
vocabulary, writing five fixed-named artifacts under the output
|
||||||
directory given as the third positional command-line argument.
|
directory given as the third positional command-line argument.
|
||||||
First, the keywords produced by the two runs of the tagging step
|
First, the keywords produced by the two runs of the tagging step
|
||||||
are pooled into the pooled keyword list, per the project's handoff
|
are pooled into the pooled keyword list, per the project's handoff
|
||||||
contract: the pool is the plain union of every keyword key observed
|
contract: the pool is the plain union of every keyword key observed
|
||||||
across both runs' valid records, exact-string deduplicated and
|
across both runs' valid records, exact-string deduplicated and
|
||||||
sorted, written as a plain text file with one keyword per line, as
|
sorted, written as a plain text file with one keyword per line, as
|
||||||
:data:`SOURCE_KEYWORDS_TXT`. The provenance mapping records where
|
:data:`SOURCE_KEYWORDS_TXT`. Then the coding groups are built from
|
||||||
every keyword came from for audit purposes as a CSV file, as
|
the pooled keyword list by
|
||||||
:data:`SOURCE_PROVENANCE_CSV`; it never enters any LLM input. Then
|
|
||||||
the coding groups are built from the pooled keyword list by
|
|
||||||
sentence-embedding every keyword and clustering the embeddings into
|
sentence-embedding every keyword and clustering the embeddings into
|
||||||
the number of groups given by the required ``--clusters``
|
the number of groups given by the required ``--clusters``
|
||||||
command-line option: the group membership is written as a CSV file
|
command-line option: the group membership is written as a CSV file
|
||||||
@@ -55,9 +53,6 @@ are not installed."""
|
|||||||
SOURCE_KEYWORDS_TXT: str = "source-keywords.txt"
|
SOURCE_KEYWORDS_TXT: str = "source-keywords.txt"
|
||||||
"""The pooled keyword text file's fixed name under the output
|
"""The pooled keyword text file's fixed name under the output
|
||||||
directory."""
|
directory."""
|
||||||
SOURCE_PROVENANCE_CSV: str = "source-provenance.csv"
|
|
||||||
"""The keyword provenance CSV file's fixed name under the output
|
|
||||||
directory."""
|
|
||||||
RESULT_KEYWORDS_TXT: str = "result-keywords.txt"
|
RESULT_KEYWORDS_TXT: str = "result-keywords.txt"
|
||||||
"""The group name keyword text file's fixed name under the output
|
"""The group name keyword text file's fixed name under the output
|
||||||
directory."""
|
directory."""
|
||||||
@@ -74,9 +69,6 @@ directory."""
|
|||||||
type Records = list[tuple[int, dict[str, Any]]]
|
type Records = list[tuple[int, dict[str, Any]]]
|
||||||
"""The valid records of one run: (song ID, keyword mapping) pairs."""
|
"""The valid records of one run: (song ID, keyword mapping) pairs."""
|
||||||
|
|
||||||
type Provenance = dict[str, list[tuple[str, int]]]
|
|
||||||
"""The occurrences of every keyword, keyed by the keyword."""
|
|
||||||
|
|
||||||
|
|
||||||
def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||||
"""Parse the command-line arguments.
|
"""Parse the command-line arguments.
|
||||||
@@ -99,9 +91,8 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
|||||||
"output_dir", type=Path,
|
"output_dir", type=Path,
|
||||||
help="the output directory, created if missing, that"
|
help="the output directory, created if missing, that"
|
||||||
f" receives {SOURCE_KEYWORDS_TXT},"
|
f" receives {SOURCE_KEYWORDS_TXT},"
|
||||||
f" {SOURCE_PROVENANCE_CSV}, {RESULT_KEYWORDS_TXT},"
|
f" {RESULT_KEYWORDS_TXT}, {RESULT_GROUPS_CSV},"
|
||||||
f" {RESULT_GROUPS_CSV}, {KEYWORDS_TO_MERGE_JSON},"
|
f" {KEYWORDS_TO_MERGE_JSON}, and {META_JSON}")
|
||||||
f" and {META_JSON}")
|
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--model", default=MODEL,
|
"--model", default=MODEL,
|
||||||
help=f"the sentence embedding model (default \"{MODEL}\")")
|
help=f"the sentence embedding model (default \"{MODEL}\")")
|
||||||
@@ -159,7 +150,7 @@ def parse_song_id(item_id: str, path: Path) -> int:
|
|||||||
return int(item_id[len(prefix):])
|
return int(item_id[len(prefix):])
|
||||||
|
|
||||||
|
|
||||||
def load_run(run_dir: Path) -> tuple[str, Records]:
|
def load_run(run_dir: Path) -> Records:
|
||||||
"""Load and validate the keyword records of one tagging run.
|
"""Load and validate the keyword records of one tagging run.
|
||||||
|
|
||||||
Records carrying an "error" field are skipped. A "text"
|
Records carrying an "error" field are skipped. A "text"
|
||||||
@@ -169,9 +160,8 @@ def load_run(run_dir: Path) -> tuple[str, Records]:
|
|||||||
|
|
||||||
:param run_dir: The run's archive directory, containing
|
:param run_dir: The run's archive directory, containing
|
||||||
``output.jsonl``.
|
``output.jsonl``.
|
||||||
:return: The run label (the directory's basename) and its
|
:return: The run's valid records, each the song ID and the
|
||||||
valid records, each the song ID and the parsed keyword
|
parsed keyword mapping, in file order.
|
||||||
mapping, in file order.
|
|
||||||
:raises OSError: When ``output.jsonl`` cannot be read.
|
:raises OSError: When ``output.jsonl`` cannot be read.
|
||||||
:raises ValueError: When a line is not a well-formed output
|
:raises ValueError: When a line is not a well-formed output
|
||||||
record, or a "text" field is invalid per the rules above.
|
record, or a "text" field is invalid per the rules above.
|
||||||
@@ -205,33 +195,23 @@ def load_run(run_dir: Path) -> tuple[str, Records]:
|
|||||||
f"{path}: id {record['id']}: \"text\" does not"
|
f"{path}: id {record['id']}: \"text\" does not"
|
||||||
" parse to a JSON object")
|
" parse to a JSON object")
|
||||||
records.append((song_id, keywords))
|
records.append((song_id, keywords))
|
||||||
return run_dir.name, records
|
return records
|
||||||
|
|
||||||
|
|
||||||
def pool_keywords(runs: list[tuple[str, Records]],
|
def pool_keywords(runs: list[Records]) -> list[str]:
|
||||||
) -> tuple[list[str], Provenance]:
|
|
||||||
"""Pool the keywords of the given tagging runs.
|
"""Pool the keywords of the given tagging runs.
|
||||||
|
|
||||||
:param runs: The runs, each the run label and its valid
|
:param runs: The runs, each its valid records (song ID,
|
||||||
records (song ID, keyword mapping).
|
keyword mapping).
|
||||||
:return: The sorted, exact-string-deduplicated keyword list
|
:return: The sorted, exact-string-deduplicated keyword list.
|
||||||
and the provenance mapping from each keyword to its
|
|
||||||
occurrences, sorted by (run label, song ID).
|
|
||||||
"""
|
"""
|
||||||
provenance: Provenance = {}
|
pool: set[str] = set()
|
||||||
label: str
|
|
||||||
records: Records
|
records: Records
|
||||||
for label, records in runs:
|
for records in runs:
|
||||||
song_id: int
|
|
||||||
keywords: dict[str, Any]
|
keywords: dict[str, Any]
|
||||||
for song_id, keywords in records:
|
for _, keywords in records:
|
||||||
keyword: str
|
pool.update(keywords)
|
||||||
for keyword in keywords:
|
return sorted(pool)
|
||||||
provenance.setdefault(keyword, []).append(
|
|
||||||
(label, song_id))
|
|
||||||
for occurrences in provenance.values():
|
|
||||||
occurrences.sort()
|
|
||||||
return sorted(provenance.keys()), provenance
|
|
||||||
|
|
||||||
|
|
||||||
def write_pool(path: Path, keywords: list[str]) -> None:
|
def write_pool(path: Path, keywords: list[str]) -> None:
|
||||||
@@ -249,31 +229,6 @@ def write_pool(path: Path, keywords: list[str]) -> None:
|
|||||||
encoding="utf-8")
|
encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
def write_provenance(path: Path, provenance: Provenance) -> None:
|
|
||||||
"""Write the keyword provenance mapping.
|
|
||||||
|
|
||||||
Writes a CSV file with the header row
|
|
||||||
``Keyword,Run,Song``, one row per occurrence, long format.
|
|
||||||
Rows are sorted by keyword lexicographically, then by run
|
|
||||||
label, then by song ID.
|
|
||||||
|
|
||||||
:param path: The path of the provenance CSV file to write.
|
|
||||||
:param provenance: The provenance mapping from each keyword
|
|
||||||
to its occurrences (run label, song ID).
|
|
||||||
:return: None.
|
|
||||||
:raises OSError: When the file cannot be written.
|
|
||||||
"""
|
|
||||||
keyword: str
|
|
||||||
with open(path, "w", encoding="utf-8", newline="") as file:
|
|
||||||
writer: Any = csv.writer(file)
|
|
||||||
writer.writerow(["Keyword", "Run", "Song"])
|
|
||||||
for keyword in sorted(provenance.keys()):
|
|
||||||
label: str
|
|
||||||
song_id: int
|
|
||||||
for label, song_id in provenance[keyword]:
|
|
||||||
writer.writerow([keyword, label, song_id])
|
|
||||||
|
|
||||||
|
|
||||||
def encode_keywords(keywords: list[str], model_name: str,
|
def encode_keywords(keywords: list[str], model_name: str,
|
||||||
revision: str | None) -> Any:
|
revision: str | None) -> Any:
|
||||||
"""Encode the keywords into L2-normalized sentence embeddings.
|
"""Encode the keywords into L2-normalized sentence embeddings.
|
||||||
@@ -507,14 +462,14 @@ def write_keywords_to_merge(path: Path,
|
|||||||
|
|
||||||
|
|
||||||
def build_meta(
|
def build_meta(
|
||||||
run1: tuple[str, Records], run2: tuple[str, Records],
|
run1: Records, run2: Records,
|
||||||
args: argparse.Namespace, keyword_count: int,
|
args: argparse.Namespace, keyword_count: int,
|
||||||
extra_keywords: list[str],
|
extra_keywords: list[str],
|
||||||
versions: dict[str, str]) -> dict[str, Any]:
|
versions: dict[str, str]) -> dict[str, Any]:
|
||||||
"""Build the run metadata recorded into :data:`META_JSON`.
|
"""Build the run metadata recorded into :data:`META_JSON`.
|
||||||
|
|
||||||
:param run1: The first run's label and valid records.
|
:param run1: The first run's valid records.
|
||||||
:param run2: The second run's label and valid records.
|
:param run2: The second run's valid records.
|
||||||
:param args: The parsed command-line arguments.
|
:param args: The parsed command-line arguments.
|
||||||
:param keyword_count: The number of pooled keywords.
|
:param keyword_count: The number of pooled keywords.
|
||||||
:param extra_keywords: The extra a-priori keywords given via
|
:param extra_keywords: The extra a-priori keywords given via
|
||||||
@@ -526,7 +481,7 @@ def build_meta(
|
|||||||
return {
|
return {
|
||||||
"script_version": SCRIPT_VERSION,
|
"script_version": SCRIPT_VERSION,
|
||||||
"source_runs": [str(args.run_dir_1), str(args.run_dir_2)],
|
"source_runs": [str(args.run_dir_1), str(args.run_dir_2)],
|
||||||
"source_records": [len(run1[1]), len(run2[1])],
|
"source_records": [len(run1), len(run2)],
|
||||||
"embedding": {
|
"embedding": {
|
||||||
"model": args.model, "revision": args.revision},
|
"model": args.model, "revision": args.revision},
|
||||||
"clustering": {
|
"clustering": {
|
||||||
@@ -563,18 +518,17 @@ def write_meta(path: Path, meta: dict[str, Any]) -> None:
|
|||||||
def main(argv: list[str] | None = None) -> int:
|
def main(argv: list[str] | None = None) -> int:
|
||||||
"""Pool the two tagging runs' keywords and cluster them.
|
"""Pool the two tagging runs' keywords and cluster them.
|
||||||
|
|
||||||
Writes the six fixed-named artifacts under the output
|
Writes the five fixed-named artifacts under the output
|
||||||
directory, creating it (with parents) if it does not exist:
|
directory, creating it (with parents) if it does not exist:
|
||||||
the pooled keyword text file and the keyword provenance CSV
|
the pooled keyword text file; then the group membership CSV
|
||||||
file; then the group membership CSV file, holding the
|
file, holding the clustering result alone; the group name
|
||||||
clustering result alone; the group name keyword text file,
|
keyword text file, holding the same group names as a readable
|
||||||
holding the same group names as a readable list; the coding
|
list; the coding keyword set JSON file, holding the group
|
||||||
keyword set JSON file, holding the group names plus every
|
names plus every extra keyword given via ``--extra-keyword``;
|
||||||
extra keyword given via ``--extra-keyword``; and the run
|
and the run metadata JSON file, recording the command-line
|
||||||
metadata JSON file, recording the command-line choices and
|
choices and the environment. When the input is rejected, or
|
||||||
the environment. When the input is rejected, or an extra
|
an extra keyword duplicates a group name or another extra
|
||||||
keyword duplicates a group name or another extra keyword,
|
keyword, none of the five files is written.
|
||||||
none of the six files is written.
|
|
||||||
|
|
||||||
:param argv: The command-line arguments, or None for
|
:param argv: The command-line arguments, or None for
|
||||||
``sys.argv``.
|
``sys.argv``.
|
||||||
@@ -582,17 +536,15 @@ def main(argv: list[str] | None = None) -> int:
|
|||||||
"""
|
"""
|
||||||
started: float = time.monotonic()
|
started: float = time.monotonic()
|
||||||
args: argparse.Namespace = parse_args(argv)
|
args: argparse.Namespace = parse_args(argv)
|
||||||
run1: tuple[str, Records]
|
run1: Records
|
||||||
run2: tuple[str, Records]
|
run2: Records
|
||||||
try:
|
try:
|
||||||
run1 = load_run(args.run_dir_1)
|
run1 = load_run(args.run_dir_1)
|
||||||
run2 = load_run(args.run_dir_2)
|
run2 = load_run(args.run_dir_2)
|
||||||
except (OSError, ValueError) as error:
|
except (OSError, ValueError) as error:
|
||||||
print(f"error: {error}", file=sys.stderr)
|
print(f"error: {error}", file=sys.stderr)
|
||||||
return 1
|
return 1
|
||||||
keywords: list[str]
|
keywords: list[str] = pool_keywords([run1, run2])
|
||||||
provenance: Provenance
|
|
||||||
keywords, provenance = pool_keywords([run1, run2])
|
|
||||||
try:
|
try:
|
||||||
embeddings: Any = encode_keywords(
|
embeddings: Any = encode_keywords(
|
||||||
keywords, args.model, args.revision)
|
keywords, args.model, args.revision)
|
||||||
@@ -612,8 +564,6 @@ def main(argv: list[str] | None = None) -> int:
|
|||||||
args.output_dir.mkdir(parents=True, exist_ok=True)
|
args.output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
write_pool(
|
write_pool(
|
||||||
args.output_dir / SOURCE_KEYWORDS_TXT, keywords)
|
args.output_dir / SOURCE_KEYWORDS_TXT, keywords)
|
||||||
write_provenance(
|
|
||||||
args.output_dir / SOURCE_PROVENANCE_CSV, provenance)
|
|
||||||
write_groups(args.output_dir / RESULT_GROUPS_CSV, groups)
|
write_groups(args.output_dir / RESULT_GROUPS_CSV, groups)
|
||||||
write_keyword_names(
|
write_keyword_names(
|
||||||
args.output_dir / RESULT_KEYWORDS_TXT, groups)
|
args.output_dir / RESULT_KEYWORDS_TXT, groups)
|
||||||
|
|||||||
@@ -39,9 +39,6 @@ class TestClusterKeywords(unittest.TestCase):
|
|||||||
self.__source_keywords_txt: Path \
|
self.__source_keywords_txt: Path \
|
||||||
= self.__output_dir \
|
= self.__output_dir \
|
||||||
/ cluster_keywords.SOURCE_KEYWORDS_TXT
|
/ cluster_keywords.SOURCE_KEYWORDS_TXT
|
||||||
self.__source_provenance_csv: Path \
|
|
||||||
= self.__output_dir \
|
|
||||||
/ cluster_keywords.SOURCE_PROVENANCE_CSV
|
|
||||||
self.__result_keywords_txt: Path \
|
self.__result_keywords_txt: Path \
|
||||||
= self.__output_dir \
|
= self.__output_dir \
|
||||||
/ cluster_keywords.RESULT_KEYWORDS_TXT
|
/ cluster_keywords.RESULT_KEYWORDS_TXT
|
||||||
@@ -179,16 +176,6 @@ class TestClusterKeywords(unittest.TestCase):
|
|||||||
self.assertEqual(lines[-1], "")
|
self.assertEqual(lines[-1], "")
|
||||||
return lines[:-1]
|
return lines[:-1]
|
||||||
|
|
||||||
def __read_source_provenance(self) -> list[list[str]]:
|
|
||||||
"""Read the source provenance CSV file.
|
|
||||||
|
|
||||||
:return: All rows, including the header row, in file
|
|
||||||
order.
|
|
||||||
"""
|
|
||||||
with open(self.__source_provenance_csv, encoding="utf-8",
|
|
||||||
newline="") as file:
|
|
||||||
return list(csv.reader(file))
|
|
||||||
|
|
||||||
def __read_groups(self) -> list[list[str]]:
|
def __read_groups(self) -> list[list[str]]:
|
||||||
"""Read the group membership CSV file.
|
"""Read the group membership CSV file.
|
||||||
|
|
||||||
@@ -303,7 +290,6 @@ class TestClusterKeywords(unittest.TestCase):
|
|||||||
self.assertEqual(status, 1)
|
self.assertEqual(status, 1)
|
||||||
self.assertIn("duplicate key", stderr)
|
self.assertIn("duplicate key", stderr)
|
||||||
self.assertFalse(self.__source_keywords_txt.exists())
|
self.assertFalse(self.__source_keywords_txt.exists())
|
||||||
self.assertFalse(self.__source_provenance_csv.exists())
|
|
||||||
self.assertFalse(self.__result_groups_csv.exists())
|
self.assertFalse(self.__result_groups_csv.exists())
|
||||||
self.assertFalse(self.__result_keywords_txt.exists())
|
self.assertFalse(self.__result_keywords_txt.exists())
|
||||||
self.assertFalse(self.__keywords_to_merge_json.exists())
|
self.assertFalse(self.__keywords_to_merge_json.exists())
|
||||||
@@ -324,60 +310,11 @@ class TestClusterKeywords(unittest.TestCase):
|
|||||||
self.assertEqual(status, 1)
|
self.assertEqual(status, 1)
|
||||||
self.assertIn("song-1", stderr)
|
self.assertIn("song-1", stderr)
|
||||||
self.assertFalse(self.__source_keywords_txt.exists())
|
self.assertFalse(self.__source_keywords_txt.exists())
|
||||||
self.assertFalse(self.__source_provenance_csv.exists())
|
|
||||||
self.assertFalse(self.__result_groups_csv.exists())
|
self.assertFalse(self.__result_groups_csv.exists())
|
||||||
self.assertFalse(self.__result_keywords_txt.exists())
|
self.assertFalse(self.__result_keywords_txt.exists())
|
||||||
self.assertFalse(self.__keywords_to_merge_json.exists())
|
self.assertFalse(self.__keywords_to_merge_json.exists())
|
||||||
self.assertFalse(self.__meta_json.exists())
|
self.assertFalse(self.__meta_json.exists())
|
||||||
|
|
||||||
def test_provenance_content_and_ordering(self) -> None:
|
|
||||||
"""Test the provenance content and its ordering: rows
|
|
||||||
sorted by keyword lexicographically, then by run label,
|
|
||||||
then by song ID."""
|
|
||||||
self.__write_output(self.__run1, [
|
|
||||||
{"id": "song-2", "text": json.dumps({"shared": 1})},
|
|
||||||
{"id": "song-1", "text": json.dumps({"shared": 1})},
|
|
||||||
])
|
|
||||||
self.__write_output(self.__run2, [
|
|
||||||
{"id": "song-5",
|
|
||||||
"text": json.dumps({"shared": 1, "b-middle": 1})},
|
|
||||||
])
|
|
||||||
vectors: Vectors = {
|
|
||||||
**self.__two_cluster_vectors(),
|
|
||||||
"shared": (1.0, 0.0)}
|
|
||||||
status: int
|
|
||||||
status, _ = self.__run_cluster(vectors=vectors)
|
|
||||||
self.assertEqual(status, 0)
|
|
||||||
rows: list[list[str]] = self.__read_source_provenance()
|
|
||||||
self.assertEqual(rows[1:], [
|
|
||||||
["b-middle", "run2", "5"],
|
|
||||||
["shared", "run1", "1"],
|
|
||||||
["shared", "run1", "2"],
|
|
||||||
["shared", "run2", "5"],
|
|
||||||
])
|
|
||||||
|
|
||||||
def test_provenance_file_header_and_row_count(self) -> None:
|
|
||||||
"""Test that the provenance CSV file starts with the
|
|
||||||
``Keyword,Run,Song`` header row and has exactly one row
|
|
||||||
per keyword occurrence."""
|
|
||||||
self.__write_output(self.__run1, [
|
|
||||||
{"id": "song-2", "text": json.dumps({"shared": 1})},
|
|
||||||
{"id": "song-1", "text": json.dumps({"shared": 1})},
|
|
||||||
])
|
|
||||||
self.__write_output(self.__run2, [
|
|
||||||
{"id": "song-5",
|
|
||||||
"text": json.dumps({"shared": 1, "b-middle": 1})},
|
|
||||||
])
|
|
||||||
vectors: Vectors = {
|
|
||||||
**self.__two_cluster_vectors(),
|
|
||||||
"shared": (1.0, 0.0)}
|
|
||||||
status: int
|
|
||||||
status, _ = self.__run_cluster(vectors=vectors)
|
|
||||||
self.assertEqual(status, 0)
|
|
||||||
rows: list[list[str]] = self.__read_source_provenance()
|
|
||||||
self.assertEqual(rows[0], ["Keyword", "Run", "Song"])
|
|
||||||
self.assertEqual(len(rows), 1 + 4)
|
|
||||||
|
|
||||||
def test_groups_csv_header_and_ordering(self) -> None:
|
def test_groups_csv_header_and_ordering(self) -> None:
|
||||||
"""Test the header row and the group/keyword ordering of
|
"""Test the header row and the group/keyword ordering of
|
||||||
the group membership CSV file."""
|
the group membership CSV file."""
|
||||||
@@ -529,7 +466,6 @@ class TestClusterKeywords(unittest.TestCase):
|
|||||||
self.assertEqual(status, 1)
|
self.assertEqual(status, 1)
|
||||||
self.assertIn("zzz-extra", stderr)
|
self.assertIn("zzz-extra", stderr)
|
||||||
self.assertFalse(self.__source_keywords_txt.exists())
|
self.assertFalse(self.__source_keywords_txt.exists())
|
||||||
self.assertFalse(self.__source_provenance_csv.exists())
|
|
||||||
self.assertFalse(self.__result_groups_csv.exists())
|
self.assertFalse(self.__result_groups_csv.exists())
|
||||||
self.assertFalse(self.__result_keywords_txt.exists())
|
self.assertFalse(self.__result_keywords_txt.exists())
|
||||||
self.assertFalse(self.__keywords_to_merge_json.exists())
|
self.assertFalse(self.__keywords_to_merge_json.exists())
|
||||||
@@ -555,7 +491,6 @@ class TestClusterKeywords(unittest.TestCase):
|
|||||||
self.assertEqual(status, 1)
|
self.assertEqual(status, 1)
|
||||||
self.assertIn("a-center", stderr)
|
self.assertIn("a-center", stderr)
|
||||||
self.assertFalse(self.__source_keywords_txt.exists())
|
self.assertFalse(self.__source_keywords_txt.exists())
|
||||||
self.assertFalse(self.__source_provenance_csv.exists())
|
|
||||||
self.assertFalse(self.__result_groups_csv.exists())
|
self.assertFalse(self.__result_groups_csv.exists())
|
||||||
self.assertFalse(self.__result_keywords_txt.exists())
|
self.assertFalse(self.__result_keywords_txt.exists())
|
||||||
self.assertFalse(self.__keywords_to_merge_json.exists())
|
self.assertFalse(self.__keywords_to_merge_json.exists())
|
||||||
|
|||||||
Reference in New Issue
Block a user