Compare commits
131
Commits
9b4ee3a658
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
502d8b6b5a | ||
|
|
47bef84e7b | ||
|
|
9b6450fcbb | ||
|
|
309667e82b | ||
|
|
195605826f | ||
|
|
e569b3481c | ||
|
|
f0547208ed | ||
|
|
eef434acf7 | ||
|
|
7a625d50fc | ||
|
|
e1e2c86661 | ||
|
|
2fa55f29d7 | ||
|
|
676d7788e8 | ||
|
|
6d055c1ee5 | ||
|
|
bfbbe1875e | ||
|
|
a5fd86deb6 | ||
|
|
85d776da8b | ||
|
|
e3cca7e538 | ||
|
|
245b38c306 | ||
|
|
d470600eb3 | ||
|
|
e7e6787927 | ||
|
|
28b2c18604 | ||
|
|
21a39618d3 | ||
|
|
7f3b5def0f | ||
|
|
30c9943bad | ||
|
|
5a91f66402 | ||
|
|
0db25812f2 | ||
|
|
63f00e44c5 | ||
|
|
d55725332e | ||
|
|
3a9eee19d1 | ||
|
|
02dfc794c4 | ||
|
|
e1355bcb70 | ||
|
|
a489f077c9 | ||
|
|
15cdb861ff | ||
|
|
854bfe5c4b | ||
|
|
7917d23d6e | ||
|
|
c3d6c06910 | ||
|
|
683b15e073 | ||
|
|
190f95b632 | ||
|
|
54e42c072c | ||
|
|
77049eb026 | ||
|
|
f87e3642f6 | ||
|
|
bebf85814e | ||
|
|
ea9df52bc3 | ||
|
|
62adf7f937 | ||
|
|
35d8e08871 | ||
|
|
693dbccc9d | ||
|
|
2ea23d9bce | ||
|
|
87926d9622 | ||
|
|
aa8b84c8cd | ||
|
|
7739abd105 | ||
|
|
dfdfccc460 | ||
|
|
4ae6fa7b2e | ||
|
|
3aa5c7a432 | ||
|
|
bc44d60d18 | ||
|
|
a64c04a128 | ||
|
|
c3e48aac56 | ||
|
|
40c41f2734 | ||
|
|
58189bae3a | ||
|
|
745a8eed9b | ||
|
|
7f43cf238b | ||
|
|
95a05a0631 | ||
|
|
24a4ca6970 | ||
|
|
c51c9d71e1 | ||
|
|
7008e14972 | ||
|
|
668f1bb118 | ||
|
|
019f317a82 | ||
|
|
20c48f04b9 | ||
|
|
3bae24b930 | ||
|
|
1714257c89 | ||
|
|
949a0f9c8d | ||
|
|
f2a80475a5 | ||
|
|
7c2e286c9d | ||
|
|
5325a8f9a4 | ||
|
|
71b640085a | ||
|
|
dcf9e280d0 | ||
|
|
e4f40e7c0f | ||
|
|
4dc526c4b0 | ||
|
|
15f1843cde | ||
|
|
8cdeb3db2e | ||
|
|
4fdc7b48fd | ||
|
|
de8b8ee678 | ||
|
|
5cf2b8ee8c | ||
|
|
04d5095e44 | ||
|
|
ce4f9ad481 | ||
|
|
41dc3ea2ea | ||
|
|
c03add4868 | ||
|
|
d73de3d01c | ||
|
|
14a3cbe121 | ||
|
|
4313d58471 | ||
|
|
db9fbd71d4 | ||
|
|
1802dac07a | ||
|
|
9b2af3bb50 | ||
|
|
99bbd566cf | ||
|
|
3f793afd34 | ||
|
|
05ba3942a2 | ||
|
|
d46c50a2db | ||
|
|
964006ec6d | ||
|
|
0c82da7a1d | ||
|
|
93a4f808bf | ||
|
|
7a58f37d9f | ||
|
|
b727284ef3 | ||
|
|
396848e898 | ||
|
|
55926b8dc4 | ||
|
|
ee6a1b88c5 | ||
|
|
1a7519becd | ||
|
|
a534b1b70b | ||
|
|
8fdd21f8a2 | ||
|
|
c561330bd0 | ||
|
|
64c62eb394 | ||
|
|
0d3cd69a0f | ||
|
|
81975b72e4 | ||
|
|
1b21471845 | ||
|
|
15f207ffc6 | ||
|
|
93aa06703b | ||
|
|
8fdd25f569 | ||
|
|
00b5c208ec | ||
|
|
c81bfcf51d | ||
|
|
942a74c86c | ||
|
|
1f95f0231d | ||
|
|
ae0e9d0a08 | ||
|
|
6a91bc73eb | ||
|
|
e15c988bf4 | ||
|
|
f36a30411d | ||
|
|
5d5e4fac50 | ||
|
|
e9027472f4 | ||
|
|
e0eb343ce0 | ||
|
|
8d223177cd | ||
|
|
f9ab79f0c2 | ||
|
|
b68393ab01 | ||
|
|
12ace6f45a | ||
|
|
9a42b463a3 |
@@ -1,6 +1,6 @@
|
||||
# 流行音樂中「女性力量」語彙的挪用與污染——以 Billboard Year-End Hot 100(2016-2025)為例的內容分析
|
||||
|
||||
這是研究論文《流行音樂中「女性力量」語彙的挪用與污染——以 Billboard Year-End Hot 100(2016-2025)為例的內容分析》的專案資料,包括論文本身、文件、紀錄、資料、工具程式、AI提示詞,等等。
|
||||
這是研究論文《流行音樂中「女性力量」語彙的挪用與污染——以 Billboard Year-End Hot 100(2016-2025)為例的內容分析》的專案資料,包括論文本身、文件、紀錄、資料、工具程式、LLM提示詞,等等。
|
||||
|
||||
## 論文和摘要
|
||||
|
||||
@@ -10,9 +10,17 @@
|
||||
|
||||
研究方法、記錄等文件,請參閱 docs/ 資料夾。
|
||||
|
||||
## LLM提示
|
||||
|
||||
LLM提示請參閱 prompts/ 資料夾。
|
||||
|
||||
## 輔助工具程式
|
||||
|
||||
輔助工具程式,請參閱 tools/ 資料夾。
|
||||
輔助工具程式請參閱 tools/ 資料夾。
|
||||
|
||||
## 研究資料
|
||||
|
||||
研究資料請參閱 data/ 資料夾。
|
||||
|
||||
## 授權
|
||||
|
||||
|
||||
|
Can't render this file because it is too large.
|
+61
-70
@@ -1,78 +1,69 @@
|
||||
# Project Conventions
|
||||
# 常設工作規範
|
||||
|
||||
The standing working rules of this project. Formerly the
|
||||
project `CLAUDE.md`; moved here so that Claude Code subagents
|
||||
do not inherit it into their context (blind-reading agents
|
||||
must not see it). A main session working on this project
|
||||
reads this file before touching the pipeline, the data, or
|
||||
the documents.
|
||||
本專案的常設工作規範。原專案 `CLAUDE.md`;移置於此,
|
||||
使 Claude Code subagent 不將其繼承入 context(盲判型
|
||||
agent 不得見之)。凡於本專案工作的主會話,動手管線、
|
||||
資料或文件之前,先讀本檔。
|
||||
|
||||
## Analysis pipeline
|
||||
## 分析管線
|
||||
|
||||
The pipeline has run to completion; these conventions govern
|
||||
any rerun or extension.
|
||||
管線已跑畢;本規範適用於任何重跑或擴充。
|
||||
|
||||
- LLM analysis runs via Python scripts calling the Anthropic
|
||||
Messages API, Batch API where possible. Steps 1 and 3 run
|
||||
on `claude-sonnet-4-6` with `temperature=0` and thinking
|
||||
disabled; steps 4 and 5 run on `claude-fable-5`, which
|
||||
accepts neither parameter -- step 4 absorbs its sampling
|
||||
variance by the majority vote, step 5 by consolidating the
|
||||
three readings.
|
||||
- Prompt definition files live in
|
||||
`prompts/<step><substep>-<task>.md` (e.g. 1-tag.md,
|
||||
5a-read.md; substeps are lettered, matching the step
|
||||
numbering of the paper; no version suffix -- versions live
|
||||
in git history) and are passed verbatim as the system
|
||||
prompt. The number names a step of the research
|
||||
procedure, not the file: the deterministic vocabulary step
|
||||
(step 2) has no definition file yet holds its own number.
|
||||
- Itemwise LLM judgments (per-song coding in step 3,
|
||||
per-keyword group selection in step 4) run the same
|
||||
definition file three times, independently, over the same
|
||||
input; a deterministic tally then assigns an item (a
|
||||
(song, keyword) or (group, keyword) pair) when at least
|
||||
two of the three runs assign it ("3 runs + majority
|
||||
vote"). Free-generation steps run
|
||||
twice and both outputs are pooled. The step-5 qualitative
|
||||
readings are neither: three independent readings per song,
|
||||
consolidated per song and synthesized across songs by
|
||||
their own definition files -- a qualitative protocol, not
|
||||
a vote (see docs/methodology.md). The vocabulary is
|
||||
built by a deterministic subcommand (embedding +
|
||||
clustering), not by an LLM. If a validation outcome is
|
||||
unexpected, revise the definition file and repeat that
|
||||
cycle; never patch results by hand.
|
||||
- Each run of a step is archived self-contained under the
|
||||
destination directory given explicitly on the `run-llm`
|
||||
command line (by convention `runs/<step>/run<N>/`):
|
||||
prompt snapshot, raw output, and `meta.json` (model ID,
|
||||
parameters, timestamps, batch ID). The runs of a step are
|
||||
that many separate invocations of `run-llm`. Replacing
|
||||
an existing run archive requires an explicit flag;
|
||||
superseded runs live in git history. Deterministic steps
|
||||
archive under `runs/<step>/` with no `run<N>` level.
|
||||
- Token usage and cost of every `run-llm` execution are
|
||||
recorded in `docs/run-costs.md` in the same commit as the
|
||||
run archive.
|
||||
- Scripts read the API key from the `ANTHROPIC_API_KEY`
|
||||
environment variable (`.env`, gitignored).
|
||||
- LLM 分析以 Python 腳本呼叫 Anthropic Messages API
|
||||
執行,能用 Batch API 處即用之。步驟 1 與步驟 3 以
|
||||
`claude-sonnet-4-6` 執行,`temperature=0`、thinking
|
||||
停用;步驟 4 與步驟 5 以 `claude-fable-5` 執行,該
|
||||
模型兩個參數皆不受理——步驟 4 的取樣變異由多數決
|
||||
吸收,步驟 5 由整合三份閱讀吸收。
|
||||
- 定義檔置於 `prompts/<步><次步>-<task>.md`(如
|
||||
1-tag.md、5a-read.md;次步以字母標示,與論文正文的
|
||||
步驟編號一致;不帶版本號——版本即 git 歷史),逐字
|
||||
作為 system prompt。
|
||||
- 逐項的 LLM 判斷(步驟 3 的逐首編碼、步驟 4 的逐碼
|
||||
入群判斷)以同一份定義檔、同一份輸入獨立執行三次;
|
||||
再由確定性計票將三次執行中至少兩次指派的項目(一個
|
||||
(歌,關鍵字)或(群,關鍵字)配對)收入定案
|
||||
(「三次執行+多數決」)。自由生成步驟執行兩次,
|
||||
兩份輸出進池。步驟 5 的質性閱讀兩者皆非:逐首三次
|
||||
獨立閱讀,由各自的定義檔逐首整合、跨首統整——質性
|
||||
協定,不是投票(見 docs/methodology.md)。詞彙表由
|
||||
確定性子命令建構(嵌入+分群),不經 LLM。驗證結果
|
||||
不符預期時,修訂定義檔並重複該循環;絕不手改結果。
|
||||
- 一步的每次執行皆自我完備歸檔於 `run-llm` 命令列上
|
||||
明示指定的目的目錄下(慣例為
|
||||
`data/runs/<步驟>/run<N>/`):定義檔快照、原始輸出與
|
||||
`meta.json`(model ID、參數、時間戳、batch ID)。
|
||||
一步的 N 次執行即 N 次各自的 `run-llm` 呼叫。覆蓋
|
||||
既有執行歸檔須明示旗標;被取代的執行留在 git 歷史。
|
||||
確定性步驟歸檔於 `data/runs/<步驟>/`,不分 `run<N>` 層。
|
||||
- 每次 `run-llm` 執行的 token 用量與費用記入
|
||||
`docs/run-costs.md`,與執行歸檔同一 commit。
|
||||
- 腳本自環境變數 `ANTHROPIC_API_KEY` 讀取 API key
|
||||
(`.env`,gitignored)。
|
||||
|
||||
## Data rules
|
||||
## 資料規則
|
||||
|
||||
- `data/source/` holds the immutable hand-placed raw files;
|
||||
`data/captures/` is written only by the fetch commands and the
|
||||
private import script; `data/manual/` is written only by the
|
||||
user's own hand; `data/derived/` is written only by the
|
||||
`build-db` subcommand.
|
||||
- Full lyrics are copyrighted: they stay in `data/captures/lyrics/`
|
||||
(gitignored) and must never be committed or reproduced in
|
||||
full anywhere in the repo.
|
||||
- `data/source/` 存手放後不動的原始檔;
|
||||
`data/captures/` 只由 fetch 命令與私人匯入腳本
|
||||
寫入;`data/manual/` 只由研究者親手寫入;
|
||||
`data/derived/` 只由 `build-db` 子命令寫入;
|
||||
`data/runs/` 存 LLM 執行的原始歸檔,只由執行程序
|
||||
寫入;`data/results/` 存論文引用的定案表,只由
|
||||
計票程序寫入。
|
||||
- 歌詞全文有版權:一律置於 `data/captures/lyrics/`
|
||||
(gitignored),絕不 commit,亦絕不於 repo 任何處
|
||||
全文重現。
|
||||
|
||||
## Documents
|
||||
## 文件
|
||||
|
||||
- `results/` holds the final tallied tables (what the paper
|
||||
cites); `runs/` holds raw audit records. The paper cites
|
||||
`results/` only.
|
||||
- Any change to a definition file or the plan is recorded in
|
||||
`docs/decision-log.md` with date and reason.
|
||||
- `data/results/` 存計票後的定案表(論文所引);
|
||||
`data/runs/` 存原始稽核紀錄。論文只引 `data/results/`。
|
||||
- **Commit 判準**:凡能由「committed 的輸入+committed
|
||||
的程式」決定性再生者不 commit;凡不能者一律以文字
|
||||
格式 commit,格式跟著上文「資料規則」一節所定的
|
||||
層次走。例外:`data/results/` 定案表與
|
||||
`data/derived/` 人讀報表雖可再生仍 commit——理由是
|
||||
引用穩定性、審稿人零門檻、撰稿期數字變動可 diff;
|
||||
兩者皆與工作儲存同一動作產出,稽核鏈無中間空缺。
|
||||
- 凡定義檔或研究規劃之更動,皆記入
|
||||
`docs/decision-log.md`,註明日期與原因。
|
||||
|
||||
+17
-18
@@ -99,8 +99,8 @@
|
||||
|
||||
## 2026-08-02
|
||||
|
||||
- **`import-lyrics` 不設為子命令,pilot 歌詞改以 `excludes/`
|
||||
私人腳本匯入**。理由:pilot 捕捉檔不隨論文發布,子命令
|
||||
- **`import-lyrics` 不設為子命令,pilot 歌詞改以私人腳本
|
||||
匯入**。理由:pilot 捕捉檔不隨論文發布,子命令
|
||||
形式會在發布的 CLI 裡留下讀者無法執行的死命令——要交待的
|
||||
是「沿用 pilot 捕捉」的事實(記於 lyrics-provenance.csv 與
|
||||
論文方法節),不是工具本身;工具移出專案,發布管線即
|
||||
@@ -178,8 +178,8 @@
|
||||
來源公開可稽核,快照乾淨、overrides 縮小;論文方法節揭露
|
||||
「捕捉前經研究者查證補完」。性別以公開自我認同為準,
|
||||
推測不確定且無佐證者列疑慮清單;查證屬資料策展而非
|
||||
分析,以 Claude Code 輔助、不走分析 API;QS 批次存
|
||||
`excludes/`(私人工作檔,不隨論文發布)。
|
||||
分析,以 Claude Code 輔助、不走分析 API;QS 批次以
|
||||
私人工作檔留存,不隨論文發布。
|
||||
|
||||
## 2026-08-04
|
||||
|
||||
@@ -227,7 +227,7 @@
|
||||
- **Markdown 檔名一律以 dash 連接**(decision-log.md、
|
||||
research-plan.md、project-structure.md 等;與 data/ 層
|
||||
CSV 檔名慣例一致),定義檔命名慣例同步改為
|
||||
`prompts/<task>-v<N>.md`(如 screen-v1.md),`excludes/`
|
||||
`prompts/<task>-v<N>.md`(如 screen-v1.md),版本庫外的
|
||||
私人工作檔一併改名;`conference_abstract.md` 改名並搬入
|
||||
`paper/`——它是本次年會實際送出的摘要,與全文同屬投稿
|
||||
血脈,不是 `docs/` 的內部工作文件。全 repo 指涉同步更新。
|
||||
@@ -583,8 +583,8 @@
|
||||
Billboard 署名 Pinkfong 為品牌(Wikidata Q55735607,型態
|
||||
brand),非演唱者;其唯一上榜曲 Baby Shark(2019#75)
|
||||
的實際演唱者為 Hope Segoine(KTVB 2019-03-06 報導、
|
||||
Songfacts、經紀簡介,出處詳
|
||||
`excludes/quickstatements/hope-segoine.md`)。藉既有的
|
||||
Songfacts、經紀簡介,出處詳私人留存的
|
||||
QuickStatements 草稿)。藉既有的
|
||||
署名正規化機制(`ArtistImporter.CANONICAL_ARTIST_NAMES`)
|
||||
指認:歌曲的署名字串維持榜單所印的 "Pinkfong",解析出
|
||||
的演出者實體為 Hope Segoine。品牌無性別可言,人有;
|
||||
@@ -637,10 +637,9 @@
|
||||
(sentence-transformers/all-mpnet-base-v2, revision
|
||||
e8c3b32e)對 `runs/02-cluster/source-keywords.txt` 的
|
||||
5,999 個關鍵字重新計算。組內一致性定義為組內成員兩兩
|
||||
餘弦相似度之平均。k=50 直接取自留存的分群結果
|
||||
(`excludes/k50/02-cluster/groups.csv`);k=30 以同參數
|
||||
重跑(確定性演算法),重算所得最大組恰為 491 詞,與
|
||||
當時粗算紀錄吻合,佐證重算與原實驗一致。數表:
|
||||
餘弦相似度之平均。k=50 直接取自私人留存的分群結果;
|
||||
k=30 以同參數重跑(確定性演算法),重算所得最大組恰為
|
||||
491 詞,與當時粗算紀錄吻合,佐證重算與原實驗一致。數表:
|
||||
|
||||
| | k=30 | k=50 | k=100 |
|
||||
|---|---|---|---|
|
||||
@@ -703,8 +702,8 @@
|
||||
力量語意即入選,women-power 群 11 碼,並產詞彙表外
|
||||
幻覺碼一筆),claude-fable-5 讀為詞彙化概念(要求女性
|
||||
標記,women-power 群 2 碼),與草稿盲選及深度閱讀輔助
|
||||
判讀同讀法;sonnet 對照執行歸檔備份於
|
||||
`excludes/experiments/`,支出留帳。claude-fable-5 不
|
||||
判讀同讀法;sonnet 對照執行歸檔另行私人備份,不入
|
||||
版本庫,支出留帳。claude-fable-5 不
|
||||
受理 temperature 與 thinking 參數(均不送出),無法釘
|
||||
temperature=0;執行間變異實測存在(run1/run2 於陽剛、
|
||||
脆弱邊緣碼分歧),由三票多數決吸收;詞彙表外輸出項
|
||||
@@ -751,9 +750,9 @@
|
||||
編碼、票數),納入重建摘要與清空範圍,供群層次查詢;
|
||||
重建命令自此帶 `--groups results/groups.csv`。
|
||||
|
||||
- **Fable 5 輔助判讀實驗總錄(非正式編碼;結果檔於
|
||||
gitignored 之 `excludes/`,本條為其版本庫內的程序
|
||||
錨點)**:2026-08-08 起以 Claude Code subagent(模型
|
||||
- **Fable 5 輔助判讀實驗總錄(非正式編碼;結果檔私人
|
||||
留存、不入版本庫,本條為其版本庫內的程序錨點)**:
|
||||
2026-08-08 起以 Claude Code subagent(模型
|
||||
Fable 5)進行七項深度閱讀實驗。共同程序:每首歌一個
|
||||
獨立會話、提示最少化且逐字統一、判讀者互不知情、不經
|
||||
三票制;定位為研究者的輔助判讀(初篩),不改動任何
|
||||
@@ -783,7 +782,7 @@
|
||||
`male-mixed-wp-fe-feminist-reading.md` 與同名 .ods)。
|
||||
此系列即論文方法節「最後編碼的結果,再由研究者與 LLM
|
||||
(Claude Code)輔助判讀」之所指;結果檔含大量歌詞引文,
|
||||
依著作權紀律留置 `excludes/`,不入版本庫。
|
||||
依著作權紀律私人留置,不入版本庫。
|
||||
|
||||
- **步驟 5:女性主義問題之質性深讀(設計定案)**:論文
|
||||
題目「挪用與污染」需要框架層的系統性證據;47 首
|
||||
@@ -899,7 +898,7 @@
|
||||
performer_gender 手工修正)、四份 results 定案表、時程
|
||||
展延一至二日與剩餘撰寫工作。仍然有效的原則(先導不
|
||||
比較、提示只定格式、歌手背景防火牆、commit 判準、
|
||||
版權規則、positionality)保留。
|
||||
版權規則)保留。
|
||||
|
||||
## 2026-08-17
|
||||
|
||||
|
||||
+101
-202
@@ -1,12 +1,12 @@
|
||||
# 方法細節
|
||||
|
||||
(全文方法節底稿。演算法在執行前寫定;任何修訂記入
|
||||
`decision-log.md`。定義檔全文見 `prompts/`,執行紀錄見
|
||||
`runs/`。)
|
||||
`decision-log.md`。)
|
||||
|
||||
## 自然編碼管線總覽
|
||||
|
||||
五個步驟:步驟 1 自由標註(兩次執行進池)→ 步驟 2 詞彙表
|
||||
五個步驟:步驟 1 自由標註(`claude-sonnet-4-6`,
|
||||
temperature=0、thinking 關閉,兩次執行進池)→ 步驟 2 詞彙表
|
||||
建構(詞向量分群,確定性)→ 步驟 3 全量編碼(三次執行+
|
||||
多數決)→ 步驟 4 語意編碼群(三次執行+多數決)→ 步驟 5
|
||||
女性主義問題之質性深讀(三次閱讀+逐首整合+樣態統整)。
|
||||
@@ -14,13 +14,6 @@
|
||||
不接觸歌詞,步驟 2 亦不呼叫 LLM。設計原則見
|
||||
`research-plan.md`;本檔記載可重現的演算法細節。
|
||||
|
||||
編號的所指為**研究程序的工序**,不是定義檔:步驟 1、
|
||||
步驟 3、步驟 4 與步驟 5 有定義檔(`prompts/`;步驟 5 依
|
||||
子工序有三份),步驟 2 沒有——它是單一確定性計算,由
|
||||
`cluster-keywords` 一個子命令完成。有無定義檔的區別即
|
||||
「該步是否為 LLM 判斷」,由 `prompts/` 是否存在同號檔案
|
||||
直接可見。
|
||||
|
||||
## 步驟 2 詞彙表建構——詞向量分群
|
||||
|
||||
詞彙表由確定性程序產生,不經 LLM。完整分割(每個關鍵字
|
||||
@@ -30,8 +23,7 @@
|
||||
### 進池
|
||||
|
||||
兩次標註執行的全部關鍵字取聯集、逐字串精確去重、字典序
|
||||
排列。失敗與拒答的記錄跳過(其歌曲不貢獻關鍵字);解析
|
||||
時偵測重複鍵,違規即失敗。
|
||||
排列。失敗與拒答的記錄跳過(其歌曲不貢獻關鍵字)。
|
||||
|
||||
### 分群
|
||||
|
||||
@@ -39,14 +31,10 @@
|
||||
(釘定 revision),關鍵字的連字號先還原為空格再編碼,
|
||||
輸出 768 維向量並 L2 正規化。
|
||||
- **分群**:階層式聚合分群(Ward linkage),k=100。
|
||||
向量既已正規化,歐氏距離與餘弦相似度單調對應;三種
|
||||
linkage 實測比較,Ward 於各個 k 的組內一致性均最高
|
||||
(average 與 complete 皆產生吞噬半數語料的巨大異質
|
||||
組)。
|
||||
向量既已正規化,歐氏距離與餘弦相似度單調對應。
|
||||
- **組數的取捨**:k 太小則壓縮比過高,樹上層被迫併入
|
||||
不相干的詞,組雖大而無主題(k=30 最大組 491 詞、
|
||||
組內一致性 0.41,成員橫跨籃球、海灘、外星人綁架);
|
||||
k 太大則人工難以通覽。定於 100,理由是實測顯示雜物櫃
|
||||
不相干的詞,組雖大而無主題;k 太大則人工難以通覽。
|
||||
定於 100,理由是實測顯示雜物櫃
|
||||
組於此始裂解為有主題的組,且編碼實測未見碼數過多的
|
||||
副作用——全量三次執行下 101 個碼全數用到;模型另行
|
||||
造出的碼共 13 筆,佔 44,149 筆標籤指派的 0.03%。
|
||||
@@ -55,39 +43,23 @@
|
||||
自己產出過的關鍵字,非任何人事後撰寫。已知限制:
|
||||
組越大越異質時,medoid 只是折衷詞,可能代表不了組內
|
||||
內容(實測 `mutual-individuality` 組內一致 0.70 而
|
||||
編碼從未使用);此類碼於結果中呈現為零使用,據實
|
||||
報告,不事後改名。
|
||||
- **取捨紀錄**:曾以 LLM 單發收斂(merge/cap 兩步)
|
||||
實作本步,四種模型六次執行全部無法維持完整分割,
|
||||
已棄用(詳見 `decision-log.md` 2026-08-05;棄用的
|
||||
定義檔止於 git 歷史,見 `git log -- prompts/`)。
|
||||
- **產物**:五份,前綴分別標示來源與結果。
|
||||
`source-keywords.txt`(進池後的關鍵字,一行一個)
|
||||
記錄進來的是什麼;`result-keywords.txt`(組名,一行
|
||||
一個)與 `groups.csv`(欄位 Group、Keyword,一列一個
|
||||
成員)記錄算出來的分割;`keywords-to-merge.json`
|
||||
(`{"keywords": [...]}`)是實際交給模型的碼,即組名
|
||||
加上先驗主題詞——五份中只有這一份含研究者的介入。
|
||||
`meta.json` 記錄執行本身:進池的兩份執行歸檔與其有效
|
||||
筆數、嵌入模型與釘定 revision、分群參數與組數、外加
|
||||
的先驗詞、關鍵字總數,以及產生數字的套件版本。凡命令
|
||||
列上的選擇與環境事實皆在此,不記時間戳與輸入雜湊
|
||||
——前者使同環境重跑逐位元組可再生,後者只會重述 git
|
||||
已保證的事。
|
||||
編碼從未使用);此類碼於結果中呈現為零使用。
|
||||
- **取捨紀錄**:詞彙表分群曾比較的替代法與棄用理由,見
|
||||
`decision-log.md` 2026-08-05 條。
|
||||
- **產物**:分群輸出中,只有實際交給模型的碼表——組名
|
||||
加上先驗主題詞——含研究者的介入;其餘皆為分群過程本身
|
||||
的機械紀錄。
|
||||
- **可重現性**:同一輸入、同一釘定模型、同一參數逐次
|
||||
重現。不同 CPU/BLAS 實作的浮點尾數差異可能使邊界
|
||||
詞的歸屬翻動,屬已揭露的限制;論文所用碼表逐字
|
||||
commit,引用單位為該份定案檔案。
|
||||
詞的歸屬翻動,屬已揭露的限制。
|
||||
|
||||
### women-power 的注入
|
||||
|
||||
定案詞彙表為 100 個分群組名再加上 `women-power` 一詞,
|
||||
共 101 個碼。`women-power` 是研究者任意決定的先驗主題(即本
|
||||
論文的主題本身),不由資料產生,屬揭露的儀器介入。該詞
|
||||
於執行時以 `--extra-keyword` 明示加入,不寫死在程式裏
|
||||
——研究者的介入因此每次都出現在重現命令上,而非無聲
|
||||
發生;分群結果的兩份產物不含它,只有交給模型的碼表含
|
||||
它。
|
||||
的加入於每次執行皆為明示、可稽核的介入,而非無聲內建於
|
||||
判斷邏輯;分群結果本身不含它,只有交給模型的碼表含它。
|
||||
|
||||
注入而非另設篩選軌的理由:讓研究者的主題詞與模型自己
|
||||
收斂出的類別(分群已自行長出 `female-empowerment` 等組)
|
||||
@@ -97,24 +69,17 @@
|
||||
|
||||
## 步驟 3 編碼的三次執行與多數決
|
||||
|
||||
- **模型**:`claude-sonnet-4-6`,temperature=0、thinking 關閉。
|
||||
- **三次執行**:同一份定義檔、同一份輸入檔,獨立執行
|
||||
三次,三份歸檔並列(`runs/3a-code/run1`、`run2`、
|
||||
`run3`),彼此無先後主從之別。
|
||||
三次,三份歸檔並列,彼此無先後主從之別。
|
||||
- **多數決**:一首歌的一個標籤,三次執行中至少兩次標出
|
||||
即收入定案編碼。三票不平手,裁決規則因此無例外條款,
|
||||
計票由確定性子命令完成(見交接契約)。
|
||||
- **第三票取全量**:只對前兩次分歧的標籤補問第三票,
|
||||
計票結果相同;仍採全量執行——全部歌曲、全部關鍵字
|
||||
——使三票在同一條件下取得。
|
||||
計票由確定性程序完成(見交接契約)。
|
||||
- **第三票取全量**:第三次執行同為全量——全部歌曲、
|
||||
全部關鍵字——使三票在同一條件下取得。
|
||||
- **對邊緣標籤的作用**:兩次執行只分得出「兩次皆標」與
|
||||
「僅一次標」;三次執行還分得出 3-0 與 2-1,故「定案
|
||||
編碼中有多少比例僅以一票之差成立」成為可報告的量。
|
||||
至於判定本身,兩次執行相左的標籤在何種協定下都由第三
|
||||
個判斷定奪,票數不使不確定性消失:某標籤於單次執行被
|
||||
標出的傾向若恰為一半,任何票數皆為擲幣。三票之效在
|
||||
傾向偏離一半處——多數決將判定推向該傾向本身(單次
|
||||
0.7 者為 0.78,0.9 者為 0.97),程序重跑的一致性因而
|
||||
高於單次執行,唯獨恰半處無從改善。
|
||||
|
||||
## 步驟 4 語意編碼群
|
||||
|
||||
@@ -125,103 +90,79 @@
|
||||
由 LLM 依編碼名的字面語意判斷。
|
||||
- **任務**:每筆輸入為一個群名加 101 個編碼的字母序
|
||||
清單,輸出為入選編碼的單層 JSON 陣列;定義檔
|
||||
`prompts/4-group.md` 只定格式,不含任何群的語意定義。
|
||||
- **模型**:`claude-fable-5`(步驟 1、3 為
|
||||
`claude-sonnet-4-6`)。該模型不受理 `temperature` 與
|
||||
`thinking` 參數,兩者均不送出;取樣變異由多數決吸收。
|
||||
只定格式,不含任何群的語意定義。
|
||||
- **模型**:`claude-fable-5`;取樣變異由多數決吸收;
|
||||
模型裁定的理由與對照實驗見決策日誌。
|
||||
- **三次執行**:同一份定義檔、同一份輸入檔,獨立執行
|
||||
三次,歸檔並列(`runs/4-group/run1`、`run2`、`run3`)。
|
||||
三次,歸檔並列。
|
||||
- **多數決**:一個(群,編碼)配對,三次執行中至少兩次
|
||||
入選即屬該群;不在 101 碼詞彙表內的輸出項無效,每筆
|
||||
丟棄印於標準錯誤。計票由確定性子命令 `tally-groups`
|
||||
完成:`tally-groups <執行歸檔 1> <執行歸檔 2> <執行歸檔
|
||||
3> <合法碼清單> <輸出 CSV>`,合法碼清單之產法同步驟 3。
|
||||
定案分群寫入 `results/groups.csv`,欄位 `Group`、
|
||||
`Keyword`、`Votes`,列序先依群名、再依編碼,一律以
|
||||
Unicode 碼位比較,換行為 CRLF。
|
||||
- **工作儲存**:`build-db --groups <定案分群 CSV>` 將定案
|
||||
分群逐欄照存入 `groups` 資料表(群、編碼、票數),供
|
||||
群層次查詢。
|
||||
入選即屬該群;不在 101 碼詞彙表內的輸出項無效,逐筆
|
||||
記錄後丟棄。計票由確定性計票程序完成,合法碼
|
||||
清單之產法同步驟 3,結果為定案分群表。
|
||||
|
||||
## 步驟 5 女性主義問題之質性深讀
|
||||
|
||||
本步驟之 5a 至 5c 為**質性閱讀,非編碼**:輸出為自由
|
||||
文字的問題閱讀報告,無可逐項機械比對的單位,故不適用
|
||||
三票多數決與仲裁;
|
||||
三次獨立閱讀為分析者三角檢核,逐首整合為整合而非裁決,
|
||||
跨首統整之產出為草稿,終審與詮釋由研究者為之。論文引用
|
||||
本步驟之 5a 至 5c 為**質性閱讀,非編碼**:輸出為自由
|
||||
文字的問題閱讀報告,無可逐項機械比對的單位,故不適用
|
||||
三票多數決與仲裁;
|
||||
三次獨立閱讀為分析者三角檢核,逐首整合為整合而非裁決,
|
||||
跨首統整之產出為草稿,終審與詮釋由研究者為之。論文引用
|
||||
本步驟時不作次數宣稱。
|
||||
|
||||
- **對象**:定案編碼含 `women-power` 或
|
||||
- **對象**:定案編碼含 `women-power` 或
|
||||
`female-empowerment` 的 145 首歌。
|
||||
- **5a 逐首閱讀**:每筆輸入為一首歌的完整歌詞逐字全文,
|
||||
不含歌名與演唱者(盲讀);定義檔
|
||||
`prompts/5a-read.md`。同一份定義檔、同一份輸入檔,
|
||||
獨立執行三次,歸檔並列(`runs/5a-read/run1`、
|
||||
`run2`、`run3`)。
|
||||
- **5b 逐首整合**:每筆輸入為該首歌的三份閱讀報告
|
||||
(不含歌詞);以問題機制為單位保守合併,標收斂註記
|
||||
((3/3)、(2/3)),主清單僅列兩讀以上提出者,單讀發現
|
||||
以一行存目;定義檔 `prompts/5b-consolidate.md`,
|
||||
執行一次,歸檔 `runs/5b-consolidate/run1`。
|
||||
- **5c 樣態統整**:單筆輸入為一批整合報告;歸納問題
|
||||
**樣態**——問題呈現與運作的重複形態,非問題分類,
|
||||
代表引句僅取自主清單;定義檔
|
||||
`prompts/5c-synthesize.md`。四種輸入範圍各執行
|
||||
一次:全 145 首之基底統整(歸檔
|
||||
`runs/5c-synthesize/run1`),及依 `performer_gender`
|
||||
(演唱聲音之性別)切分之三個發話脈絡統整——男聲
|
||||
(male,歸檔 `run2`)、女聲(female,歸檔 `run3`)、
|
||||
混合(mixed,歸檔 `run4`);基底看橫貫各脈絡之樣態,
|
||||
分組看各權力脈絡下之樣態。genderfluid 與 non-binary
|
||||
共 3 首不設群——樣本數不支持歸納——僅入基底統整,
|
||||
由研究者以個案閱讀。同一輸入不重複執行:自由歸納之
|
||||
產出無機械合併可言,其變異由 5a 三讀、5b 整合與
|
||||
草稿地位承接,四份草稿互為對照,由研究者終審裁決。
|
||||
- **5d 樣態標註**:以 (歌, 樣態) 對為可逐項機械比對之
|
||||
單位,回歸「三次執行+多數決」協定。對象為「有問題」
|
||||
的歌——三讀中至多一個「無」(多數決精神;恰兩「無」
|
||||
者其整合報告主清單必為空,與 5b 主清單規則自洽),
|
||||
計 111 首(男聲 12、女聲 69、混合 29、genderfluid 1)。
|
||||
樣態表為三份分組統整草稿原文:男聲 13 條(M1–M13)、
|
||||
女聲 14 條(F1–F14)、混合 16 條(X1–X16);基底 15 條
|
||||
- **5a 逐首閱讀**:每筆輸入為一首歌的完整歌詞逐字全文,
|
||||
不含歌名與演唱者(盲讀)。同一份定義檔、同一份輸入檔,
|
||||
獨立執行三次,歸檔並列。
|
||||
- **5b 逐首整合**:每筆輸入為該首歌的三份閱讀報告
|
||||
(不含歌詞);以問題機制為單位保守合併,標收斂註記
|
||||
((3/3)、(2/3)),主清單僅列兩讀以上提出者,單讀發現
|
||||
以一行存目;定義檔,執行一次。
|
||||
- **5c 樣態統整**:單筆輸入為一批整合報告;歸納問題
|
||||
**樣態**——問題呈現與運作的重複形態,非問題分類,
|
||||
代表引句僅取自主清單。四種輸入範圍各執行
|
||||
一次:全 145 首之基底統整,及依演唱聲音之性別切分
|
||||
之三個發話脈絡統整——男聲(male)、女聲(female)、
|
||||
混合(mixed);基底看橫貫各
|
||||
脈絡之樣態,分組看各權力脈絡下之樣態。genderfluid
|
||||
與 non-binary 共 3 首不設群——樣本數不支持歸納——
|
||||
僅入基底統整。同一輸入不重複
|
||||
執行:自由歸納之產出無機械合併可言,其變異由 5a
|
||||
三讀、5b 整合與草稿地位承接,四份草稿互為對照,由
|
||||
研究者終審裁決。
|
||||
- **5d 樣態標註**:以(歌,樣態)對為可逐項機械比對之
|
||||
單位,回歸「三次執行+多數決」協定。對象為「有問題」
|
||||
的歌——三讀中至多一個「無」(多數決精神;恰兩「無」
|
||||
者其整合報告主清單必為空,與 5b 主清單規則自洽),
|
||||
計 111 首(男聲 12、女聲 69、混合 29、genderfluid 1)。
|
||||
樣態表為三份分組統整草稿原文:男聲 13 條(M1–M13)、
|
||||
女聲 14 條(F1–F14)、混合 16 條(X1–X16);基底 15 條
|
||||
不入矩陣——全體歸納之一條樣態可能疊合不同方向的
|
||||
權力關係(男對女、女對男),實為多個樣態共用一名。
|
||||
檢驗範圍:男聲樣態不檢驗純女聲歌、女聲樣態不檢驗
|
||||
純男聲歌(發話位置範疇錯置,檢查無意義);混合樣態
|
||||
檢驗全部(其男女聲部無系統化切分方式,無意義之標註
|
||||
容忍之);genderfluid 歌三套全查(無自身透鏡,發話
|
||||
位置無法先驗決定),其歸屬引用維持個案地位。每筆
|
||||
輸入為一首歌之整合報告(「僅單獨提及」行於組裝時
|
||||
剝除,標註僅依主清單)與該首適用之樣態表;定義檔
|
||||
`prompts/5d-annotate.md`,獨立執行三次
|
||||
(`runs/5d-annotate/run1`~`run3`),(歌, 樣態) 對
|
||||
權力關係(男對女、女對男),實為多個樣態共用一名。
|
||||
檢驗範圍:男聲樣態不檢驗純女聲歌、女聲樣態不檢驗
|
||||
純男聲歌(發話位置範疇錯置,檢查無意義);混合樣態
|
||||
檢驗全部(其男女聲部無系統化切分方式,無意義之標註
|
||||
容忍之);genderfluid 歌三套全查(無自身透鏡,發話
|
||||
位置無法先驗決定),其歸屬引用維持個案地位。每筆
|
||||
輸入為一首歌之整合報告(「僅單獨提及」行於組裝時
|
||||
剝除,標註僅依主清單)與該首適用之樣態表;定義檔
|
||||
獨立執行三次,(歌,樣態)對
|
||||
得兩票以上者定案。
|
||||
- **模型**:`claude-fable-5`(與先導深讀同儀器;
|
||||
`temperature` 與 `thinking` 參數不適用,均不送出)。
|
||||
- **輸入組裝**:確定性行內腳本。5a:145 首依歌曲 ID
|
||||
升序,`content` 為歌詞逐字全文;5b:每筆
|
||||
`{"reports": [run1 輸出, run2 輸出, run3 輸出]}`;
|
||||
5c:單筆以 `song-<ID>` 為鍵、整合報告為值之 JSON
|
||||
物件,鍵集合為該次統整之範圍(基底為全 145 首,
|
||||
分組依工作庫 `performer_gender` 切分);5d:每筆
|
||||
`{"report": 主清單, "patterns": [{"id", "name",
|
||||
"description"}]}`,樣態條目自分組統整草稿機械切出,
|
||||
代表引句不隨附——引句出自特定歌曲,判該曲時形同
|
||||
預答。各輸入檔之 SHA-256 記入該步 meta。
|
||||
- **模型**:`claude-fable-5`。
|
||||
- **輸入組裝**:各步輸入檔由確定性程序自上游產物組裝。
|
||||
5d 之樣態條目自分組統整草稿機械切出,代表引句不
|
||||
隨附——引句出自特定歌曲,判該曲時形同預答。
|
||||
|
||||
## 女性力量候選集
|
||||
|
||||
候選集為兩類歌曲的合集:定案編碼含 `women-power` 者,
|
||||
以及定案編碼含研究者指認之女性力量概念域分群組者。
|
||||
指認於詞彙表定案後、黃金標準編碼開始前完成,指認清單
|
||||
與理由記入決策日誌。
|
||||
候選集為定案編碼含 `women-power` 或 `female-empowerment`
|
||||
(步驟 4 女性力量群的兩個編碼)之歌曲聯集:wp 66 首、
|
||||
fe 144 首,聯集 145 首。
|
||||
|
||||
## 軌跡對映(診斷用)
|
||||
|
||||
沿收斂軌跡的機械對映:原始關鍵字 →(兩份標註執行歸檔的
|
||||
`output.jsonl`)歌曲、原始關鍵字 →(分群)組,純程式查表,
|
||||
原始輸出)歌曲、原始關鍵字 →(分群)組,純機械查表,
|
||||
決定性。以其結果與步驟 3 直接編碼的差異率作為「收斂軌跡
|
||||
扭曲」的診斷量,不作主結果。
|
||||
|
||||
@@ -230,78 +171,36 @@
|
||||
每一步的輸出如何變成下一步的輸入,皆為確定性程序,規則
|
||||
明定如下:
|
||||
|
||||
- **歌詞輸入檔(步驟 1)**:`export-llm-input` 自工作
|
||||
儲存產出,每筆 `{"id": "song-<ID>", "content": <歌詞>}`,
|
||||
依歌曲 ID 升序。步驟 3 的輸入由同一子命令、同一工作
|
||||
儲存產出(見下),兩步的語料同一性由此成立;各步
|
||||
輸入檔的 SHA-256 記入該步 meta。
|
||||
- **步驟 1 → 2**:`cluster-keywords` 讀兩份執行歸檔的
|
||||
`output.jsonl`(一律以換行字元 `\n` 切行——歌詞含
|
||||
U+0085 等控制字元時,`str.splitlines()` 類的通用切行
|
||||
會截斷 JSON 字串,實測踩中),進池後直接分群,一次
|
||||
產出上列五份檔案。
|
||||
- **步驟 2 → 3 輸入檔**:`export-llm-input --extras
|
||||
<定案碼表>` 自工作儲存產出步驟 3 的輸入,每筆
|
||||
`{"id": "song-<ID>", "content": <字串>}`,`content` 為
|
||||
固定鍵序序列化的 `{"lyrics": …, "keywords": [...]}`,
|
||||
依歌曲 ID 升序。碼表以參數傳入而非填進定義檔——定義
|
||||
- **歌詞輸入檔(步驟 1)**:由確定性的匯出程序自工作
|
||||
儲存產出,一筆一首歌,依歌曲 ID 升序。步驟 3 的輸入
|
||||
由同一匯出程序、同一工作儲存產出(見下),兩步的語料
|
||||
同一性由此成立。
|
||||
- **步驟 1 → 2**:確定性的分群程序讀兩份執行歸檔的
|
||||
執行紀錄,進池後直接分群,產出詞彙表與交給
|
||||
模型的碼表。
|
||||
- **步驟 2 → 3 輸入檔**:同一匯出程序自工作儲存產出
|
||||
步驟 3 的輸入,一筆一首歌,兼含歌詞與定案碼表,依
|
||||
歌曲 ID 升序。碼表以參數傳入而非填進定義檔——定義
|
||||
檔只規定任務形狀,換詞彙表、換演算法都不必改它。
|
||||
- **步驟 3 定案**:`tally-codings <執行歸檔 1> <執行歸檔
|
||||
2> <執行歸檔 3> <輸出 CSV> --corrections <更正表>
|
||||
--valid-keywords <合法碼清單>` 讀三份執行歸檔的
|
||||
`output.jsonl`,依序套用更正表、驗證所有標籤皆在合法碼
|
||||
- **步驟 3 定案**:確定性的計票程序讀三份執行歸檔的
|
||||
執行紀錄,依序套用更正表、驗證所有標籤皆在合法碼
|
||||
清單之內、計票。
|
||||
- **更正表**:`data/manual/coding-corrections.csv`,研究者
|
||||
逐列校定的人工著作,欄位 `Song ID`、`Run`、`Type`、
|
||||
`To Be Replaced`、`Correct Term`。`Type` 為 `keyword`
|
||||
或 `evidence`,分別更正標籤與引述;`Correct Term` 為
|
||||
替代字串,或 `**REMOVE**` 表示刪去該筆標籤指派(`keyword`)
|
||||
或該句引述(`evidence`)。一筆 `evidence` 更正套用於該
|
||||
首歌該次執行的所有出現處。兩個文字欄以歌詞慣例「 / 」
|
||||
表示換行(與載入後的執行紀錄同一表示法,逐字比對、不再
|
||||
轉換),故一列一行,純文字工具可逐列處理。表中任一列若
|
||||
在資料中找不到對應者,即中止;校定的判準記於
|
||||
`decision-log.md`。
|
||||
- **合法碼清單**:純文字、一行一個碼,自詞彙表產出:
|
||||
`{ cat runs/2-cluster/result-keywords.txt; echo
|
||||
women-power; } | sort`。
|
||||
- **定案表**:`results/codings.csv`,欄位 `Song`、
|
||||
`Artist Credit`、`Keyword`、`Quote`,一列一個標籤。歌名
|
||||
與演出者名銜逐首查工作儲存取得,故本子命令須在
|
||||
`build-db` 之後執行。`Quote` 為該標籤在計票中各份執行
|
||||
所引的歌詞行:各份的引述串接後逐字去重,按 Unicode
|
||||
碼位排序,以單一 `|` 相接(三份執行彼此無先後主從之
|
||||
別,引述之序取決於引述本身);引述內的換行於執行紀錄
|
||||
載入時一次換成歌詞慣例「 / 」,此後更正表、定案表與
|
||||
工作儲存全鏈路同一表示法,不再還原。「 / 」的無歧義性
|
||||
是語料事實而非結構保證:全 883 首歌詞經窮舉查核不含
|
||||
「 / 」;換語料須重查。列序依印出的前三欄依序排:
|
||||
歌名、演出者名銜、
|
||||
標籤,一律以 Unicode 碼位比較,換行為 CRLF(同專案
|
||||
其他 CSV)。
|
||||
- **序列化通則**:所有中間檔為 UTF-8,欄序、鍵序與元素
|
||||
序皆依上列規則明定,無時間戳、無隨機成分;JSON 解析
|
||||
一律偵測重複鍵,違規即失敗。人讀為主的產物採純文字或
|
||||
CSV(CSV 依 RFC 4180,標題列字首大寫),機器交接檔採
|
||||
JSON。給定相同的 LLM 執行輸出,全部交接產物逐位元組
|
||||
可再生。
|
||||
- **更正表**:研究者逐列校定的人工著作,逐筆更正標籤
|
||||
或引述——以替代字串取代,或刪去該筆標籤指派或該句
|
||||
引述。一筆引述更正套用於該首歌該次執行的所有出現處。
|
||||
表中任一列若在資料中找不到對應者,即中止;校定的
|
||||
判準記於 `decision-log.md`。
|
||||
- **合法碼清單**:自詞彙表產出:分群組名加上
|
||||
`women-power`。
|
||||
- **定案編碼表**:歌名與演出者名銜逐首查工作儲存取得。
|
||||
每個定案標籤隨附其在計票中各份執行所引的
|
||||
歌詞行,供逐碼查核(三份執行彼此無先後主從之別)。
|
||||
|
||||
## 執行與稽核
|
||||
|
||||
- LLM 步驟以 `run-llm <定義檔> <輸入檔> <歸檔目錄>`
|
||||
執行;一步的 N 次執行=重現命令清單上的 N 行命令,
|
||||
各自歸檔(`runs/<步驟>/run1`、`run2`,三票制步驟另有
|
||||
`run3`)。
|
||||
- 確定性步驟(進池、分群、計票、對映)為子命令,其
|
||||
輸入輸出檔同隨 `runs/` 歸檔;因無執行變異,歸檔目錄
|
||||
下不分 `run<N>` 層。
|
||||
- LLM 步驟以批次執行程序執行,一份定義檔配一份
|
||||
輸入檔;一步的 N 次執行為 N 次各自獨立的呼叫,各自
|
||||
歸檔自我完備。
|
||||
- Batch API 的每筆請求自含全部脈絡且互不可見(平台
|
||||
契約),歌與歌之間的獨立性由此成立;各次執行的獨立
|
||||
性由「一次呼叫、一個批次、一份歸檔」的執行結構自明。
|
||||
- 每次 `run-llm` 執行的 token 用量與費用記入
|
||||
`run-costs.md`,被取代的執行一併保留供總支出核算。
|
||||
|
||||
## 映射分析方法
|
||||
|
||||
(依 2026-07-30 決策,於看到結果前寫定;待黃金標準
|
||||
編碼展開前補入。)
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# LLM 輸出的契約查核
|
||||
|
||||
(2026-08-06 量測。對象為步驟 3 的三份執行歸檔
|
||||
`runs/3a-code/run1`–`run3`,共 883 首歌、44,149 筆標籤
|
||||
`data/runs/3a-code/run1`–`run3`,共 883 首歌、44,149 筆標籤
|
||||
指派、44,146 句引述。歌詞以模型實際看到的那一份為準,即
|
||||
`tools/instance/llm-input-code.jsonl`。)
|
||||
|
||||
@@ -46,9 +46,8 @@ song-750 run3: -use=[] -abuse=["I got a thing for the hard
|
||||
liquor on ice"]
|
||||
```
|
||||
|
||||
模型寫錯後綴、已輸出的 token 收不回,遂以空陣列收束該鍵,
|
||||
再於正確的鍵補上引述。song-750 的 run1 與 run2 則整筆使用
|
||||
錯拼的鍵並附上引述,故該首的三票分裂於兩種拼寫之間。
|
||||
song-750 的 run1 與 run2 則整筆使用錯拼的鍵並附上引述,
|
||||
故該首的三票分裂於兩種拼寫之間。
|
||||
|
||||
## 引述的存在
|
||||
|
||||
|
||||
+18
-40
@@ -2,15 +2,14 @@
|
||||
|
||||
(2026-08-17 認清並記錄。本檔記述正式研究之前的先導研究:
|
||||
它做了什麼、研究者檢視到哪一層、哪些成果被沿用、哪些被
|
||||
棄用,以及它如何促成正式研究的設計。散見於
|
||||
`decision-log.md` 的相關條目在文中逐一指出。)
|
||||
棄用,以及它如何促成正式研究的設計。)
|
||||
|
||||
## 一、先導研究是什麼
|
||||
|
||||
正式研究之前,研究者曾以 Claude Code 對 Billboard Year-End
|
||||
Hot 100(2018–2025、684 首)的歌詞做過一輪探索性分析,檢視
|
||||
「女性力量」語彙的使用狀況,並附帶分析 pussy 一詞作為女性
|
||||
代稱的修辭。2026 年 8 月投出的研討會摘要即根據該輪分析撰寫。
|
||||
代稱的修辭。2026 年 4 月投出的研討會摘要即根據該輪分析撰寫。
|
||||
|
||||
## 二、它實際的執行方式
|
||||
|
||||
@@ -35,27 +34,24 @@ frame-aware 提示修正實驗,均由當時協作的 Claude Code 設計並
|
||||
## 三、被沿用的成果
|
||||
|
||||
- **歌詞捕捉檔**:先導研究蒐集的 lyrics.json(684 首,
|
||||
2018–2025)以 `excludes/` 的私人腳本匯入歌詞快取,只取
|
||||
2018–2025)以私人腳本(不入版本庫)匯入歌詞快取,只取
|
||||
識別欄位與歌詞本文,先導的分析欄位一概不匯入;出處記於
|
||||
`data/captures/lyrics-provenance.csv`,method 欄標
|
||||
`pilot-import`(見 `decision-log.md` 2026-07-31、08-02 條)。
|
||||
`data/captures/lyrics-provenance.csv`
|
||||
(見 `decision-log.md` 2026-07-31、08-02 條)。
|
||||
- **假說方向**:女性力量語彙的挪用與污染,成為正式研究的
|
||||
研究問題。
|
||||
- **粒度選擇**:正式研究鎖定 thematic keywords 這一粒度,
|
||||
係繼承先導研究三種粒度的比較結果(keywords 過碎、themes
|
||||
過早抽象),且於執行前鎖定以防事後擇優。
|
||||
過早抽象)。
|
||||
- **`women-power` 一詞的來歷**:考據先導研究的 local agent
|
||||
存檔可知,其第一步指令含數十個範例 thematic keywords,
|
||||
其中即有 women-power——為當時協作的 Claude Code 依研究者
|
||||
長期表達的關注主動加入(研究者端播種,非明示指定);先導
|
||||
的標籤 `women-power-and-empowerment` 則是第三步強制合併
|
||||
兩個關鍵字的管線人工產物。正式研究的探針因此指向被播種的
|
||||
本詞 `women-power`,不用合併假影(見 `decision-log.md`
|
||||
2026-08-04 條)。**附帶認清:先導第一步並非零語意提示,
|
||||
此即正式研究「提示詞只定格式、不定語意」設計所矯正者。**
|
||||
兩個關鍵字的管線人工產物(見 `decision-log.md` 2026-08-04
|
||||
條)。
|
||||
- **簿記容量的教訓**:先導研究九百餘詞可以在單一回應內完成
|
||||
分組,正式研究的 5,999 個關鍵字則四種模型六次執行全部未
|
||||
通過完整分割驗證,遂改用詞向量嵌入+確定性分群(見
|
||||
分組;此法未沿用,詞彙表改由確定性程序產生(見
|
||||
`decision-log.md` 2026-08-06 條)。
|
||||
|
||||
## 四、被棄用的成果
|
||||
@@ -63,35 +59,20 @@ frame-aware 提示修正實驗,均由當時協作的 Claude Code 設計並
|
||||
以下先導研究的產物未進入正式研究,論文亦未引用:
|
||||
|
||||
- **genuine/peripheral/fake 三分類與「44% 假女性力量」**:
|
||||
構念與分母皆與正式研究不同(正式研究為 145 首女性力量群
|
||||
歌曲中 111 首有性別問題),兩者不可對讀。
|
||||
- **五種「假女性力量」類型(A–E)**:正式研究改由儀器分三個
|
||||
發話脈絡各自歸納,得 43 條問題樣態(`results/patterns.csv`)。
|
||||
- **pussy 一詞的修辭分類**:正式研究範圍收斂至語彙的挪用與
|
||||
污染,未納入。
|
||||
構念與分母均未經稽核回溯,未沿用。
|
||||
- **五種「假女性力量」類型(A–E)**:未沿用。
|
||||
- **pussy 一詞的修辭分類**:未沿用。
|
||||
- **frame-aware(框架感知)提示修正法**:先導研究以框架判準
|
||||
寫進提示以提高準確率;正式研究刻意不定義編碼,因為研究
|
||||
對象正是 LLM 未受引導的自然編碼——把框架寫進提示,即無法
|
||||
再以其輸出為批判對象。兩者的設計方向相反,故未沿用。
|
||||
- **三個獨立 LLM subagent+人工仲裁的流程**:正式研究改以
|
||||
Anthropic API 逐首獨立呼叫,原因是實測證實 subagent 會繼承
|
||||
CLAUDE.md 與環境資訊,context 無法僅憑定義檔重現(見
|
||||
`decision-log.md` 2026-07-30 條)。
|
||||
寫進提示以提高準確率;未沿用。
|
||||
- **三個獨立 LLM subagent+人工仲裁的流程**:未沿用,原因見
|
||||
`decision-log.md` 2026-07-30 條。
|
||||
|
||||
## 五、它如何促成正式研究的設計
|
||||
## 五、它促成了可稽核的設計
|
||||
|
||||
先導研究的根本限制不在結論對錯,而在**不可稽核**:沒有定義檔
|
||||
快照、沒有原始輸出歸檔、沒有參數紀錄,因此任何一個數字都無法
|
||||
回溯到產生它的那一次執行。正式研究的幾項設計正是針對這一點:
|
||||
|
||||
- 所有 LLM 步驟改以 API script 執行,每次執行自我完備歸檔於
|
||||
`runs/`(定義檔快照、原始輸出、model ID 與參數、批次 ID、
|
||||
token 用量);
|
||||
- 可逐項機械比對的判斷採三次執行+多數決,並量測其穩定性
|
||||
(見 `reliability.md`);
|
||||
- 定義檔只規定任務形狀,不給主題定義、判準或範例;
|
||||
- 每一筆編碼必附歌詞引述,供逐筆查核;
|
||||
- 輸出的契約遵從另行查核(見 `output-validation.md`)。
|
||||
回溯到產生它的那一次執行。正式研究的可稽核設計,正是針對這一點
|
||||
而立。
|
||||
|
||||
換言之,正式研究對 LLM 輸出所採取的「不信任、須查核」立場,
|
||||
其第一個案例就是先導研究本身。
|
||||
@@ -102,6 +83,3 @@ frame-aware 提示修正實驗,均由當時協作的 Claude Code 設計並
|
||||
的前身),樹中不另存副本——摘要只是產物,不足以記述過程,
|
||||
故另立本檔。
|
||||
- 先導研究的歌詞捕捉檔沿用事實:`data/captures/lyrics-provenance.csv`。
|
||||
- 相關決策條目:`decision-log.md` 2026-07-30(不與先導比較、
|
||||
不用 subagent)、07-31 與 08-02(歌詞沿用與私人匯入腳本)、
|
||||
08-04(women-power 來歷考據)、08-06(詞彙表改用詞向量分群)。
|
||||
|
||||
@@ -1,139 +0,0 @@
|
||||
# 專案目錄結構
|
||||
|
||||
(2026-07-30 討論定案;2026-07-31 更新為 tools/ 子專案與
|
||||
SQLite 工作儲存架構;2026-08-17 依完成後的現況更新)
|
||||
|
||||
```
|
||||
pop-fem-audit/
|
||||
├── README.md # 專案說明
|
||||
├── .gitignore # captures/lyrics/、.env、scratch
|
||||
├── data/ # 依生命週期分層(文字格式)
|
||||
│ ├── source/ # 源頭:手放後不動
|
||||
│ │ └── yearend_hot100_2016_2025.csv # 原始榜單
|
||||
│ ├── captures/ # 外部捕捉:只由 fetch 命令與
|
||||
│ │ │ # 私人匯入腳本寫入
|
||||
│ │ ├── artists-wikidata.csv # Wikidata 快照
|
||||
│ │ ├── lyrics-provenance.csv # 歌詞出處
|
||||
│ │ └── lyrics/ # 歌詞 .txt 快取
|
||||
│ │ # (gitignored,版權)
|
||||
│ ├── manual/ # 人工著作:只由研究者手寫
|
||||
│ │ ├── coding-corrections.csv # 編碼與引述的校對表
|
||||
│ │ └── performer-gender-corrections.csv # 演唱聲音性別的
|
||||
│ │ # 手工修正
|
||||
│ └── derived/ # 衍生:只由 build-db 寫入
|
||||
│ ├── songs.csv # 歌曲報表(人讀;進 git)
|
||||
│ └── artists.csv # 歌手報表(人讀;進 git)
|
||||
├── prompts/ # LLM 定義檔(逐字作為 system prompt)
|
||||
│ └── <步><次步>-<task>.md # 1-tag.md、3a-code.md、
|
||||
│ # 4-group.md、5a-read.md、
|
||||
│ # 5b-consolidate.md、
|
||||
│ # 5c-synthesize.md、
|
||||
│ # 5d-annotate.md
|
||||
│ # (次步以字母標示,與論文正文
|
||||
│ # 的步驟編號一致;步內僅一個
|
||||
│ # 執行時省略次步)
|
||||
│ # 不帶版本號,版本即 git 歷史
|
||||
│ # (編號的所指是工序:確定性
|
||||
│ # 的步驟 2 無定義檔仍佔一號)
|
||||
├── tools/ # 輔助工具子專案(src-layout)
|
||||
│ ├── pyproject.toml # 發行名 pop-fem-audit-tools;
|
||||
│ │ # pip install -e tools/ 安裝
|
||||
│ ├── README.rst LICENSE MANIFEST.in .env.example .gitignore
|
||||
│ ├── docs/ # Sphinx API 文件
|
||||
│ ├── instance/ # SQLite 工作儲存(generated、
|
||||
│ │ # gitignored;含歌詞全文)
|
||||
│ ├── src/pop_fem_audit_tools/
|
||||
│ │ ├── __main__.py # 套件 CLI 進入點(分派子命令)
|
||||
│ │ ├── commands/ # CLI 子命令模組(登記於 __init__)
|
||||
│ │ │ ├── build_db.py # build the SQLite working store
|
||||
│ │ │ │ # from the inputs
|
||||
│ │ │ ├── export_llm_input.py # export the LLM input JSONL
|
||||
│ │ │ │ # (lyrics only) from the
|
||||
│ │ │ │ # working store
|
||||
│ │ │ ├── fetch_artists.py # fetch artist metadata from
|
||||
│ │ │ │ # Wikidata into the snapshot CSV
|
||||
│ │ │ ├── fetch_lyrics.py # fetch missing lyrics from the
|
||||
│ │ │ │ # public APIs into the lyrics dir
|
||||
│ │ │ ├── cluster_keywords.py # pool the tagging runs'
|
||||
│ │ │ │ # keywords and cluster them
|
||||
│ │ │ │ # into the codes (step 2)
|
||||
│ │ │ ├── tally_codings.py # settle step 3 by majority
|
||||
│ │ │ ├── tally_groups.py # settle step 4 by majority
|
||||
│ │ │ ├── tally_annotations.py # settle step 5d by majority
|
||||
│ │ │ └── run_llm.py # API 執行器:一份定義檔+一份輸入
|
||||
│ │ │ # →歸檔至指定目錄(Batch API);
|
||||
│ │ │ # 多次執行的計票由獨立子命令承擔
|
||||
│ │ ├── config.py # pydantic-settings 設定(.env)
|
||||
│ │ ├── database.py # SQLAlchemy engine / session / Base
|
||||
│ │ ├── models.py # SQLAlchemy ORM 資料模型
|
||||
│ │ └── utils.py # 共用工具(format_duration)
|
||||
│ └── tests/ # 單元測試(unittest)
|
||||
├── runs/ # 現行執行的完整稽核紀錄(進 git;
|
||||
│ │ # 重跑同一 run 須明示 --replace)
|
||||
│ ├── <步驟名>/ # 一步一個目錄(1-tag、3a-code、
|
||||
│ │ # 4-group、5a-read、
|
||||
│ │ # 5b-consolidate、
|
||||
│ │ # 5c-synthesize、5d-annotate)
|
||||
│ │ └── run<N>/ # LLM 步驟:每個 run 一份自我
|
||||
│ │ ├── prompt.md # 完備歸檔(定義檔快照)
|
||||
│ │ ├── output.jsonl # 該次執行原始輸出
|
||||
│ │ └── meta.json # model ID、temperature、時間戳、
|
||||
│ │ # batch ID、token 用量
|
||||
│ └── 2-cluster/ # 確定性步驟:無執行變異,
|
||||
│ # 不分 run<N> 層
|
||||
├── results/ # 論文引用的定案表 CSV(計票子命令
|
||||
│ │ # 產出;「可再生仍 commit」的例外)
|
||||
│ ├── codings.csv # 步驟 3 定案編碼
|
||||
│ ├── groups.csv # 步驟 4 定案編碼群
|
||||
│ ├── patterns.csv # 步驟 5c 定案樣態表
|
||||
│ ├── annotations.csv # 步驟 5d 定案歌×樣態
|
||||
│ └── pattern-matrix.csv # 前四者的人讀寬表
|
||||
├── docs/
|
||||
│ ├── conventions.md # 常設工作規範(原 CLAUDE.md;
|
||||
│ │ # 移入 docs/ 使 subagent 不繼承)
|
||||
│ ├── research-plan.md # 研究步驟規劃(本檔之姊妹篇)
|
||||
│ ├── project-structure.md # 本檔
|
||||
│ ├── output-validation.md # LLM 輸出的契約查核紀錄
|
||||
│ ├── decision-log.md # 決策日誌:每次改定義檔的原因
|
||||
│ ├── run-costs.md # 每次執行的 token 用量與費用
|
||||
│ ├── reliability.md # 信度:量測方式與結果
|
||||
│ ├── pilot-study.md # 先導研究的來歷與地位
|
||||
│ └── methodology.md # 方法細節(全文方法節底稿;
|
||||
│ # 映射分析方法須在看結果前寫定)
|
||||
└── paper/
|
||||
├── abstract.md # 摘要
|
||||
└── 流行音樂中「女性力量」….odt # 全文
|
||||
```
|
||||
|
||||
## 設計理由
|
||||
|
||||
- **`runs/` 自我完備**:每個執行目錄含定義檔快照 + 原始輸出 +
|
||||
meta,讀者不需 git 考古即可稽核任一筆結果。
|
||||
- **`runs/`(原始稽核資料)與 `results/`(最終表)分離**:
|
||||
論文只引 `results/`,其來源可回溯至 `runs/`。
|
||||
- **`prompts/` 檔名不帶版本號**:版本即 git 歷史,失敗的
|
||||
版本不保留;論文引用的單位是 `runs/` 內隨執行保存的定義檔
|
||||
快照(每個執行目錄自我完備),不需檔名可指的版本名。
|
||||
- **工作儲存的資料表**:`songs`(含 `performer_gender`=演唱
|
||||
聲音的性別)、`chart_entries`、`artists`、`song_artists`、
|
||||
`codings`(定案編碼:一歌一標籤一列,`quotes` 存該標籤所據的
|
||||
歌詞引述,多句以 `|` 相接)、`groups`(語意編碼群)、
|
||||
`patterns`(深讀樣態)、`annotations`(歌×樣態定案矩陣)。
|
||||
各定案表經 `build-db` 的 `--codings`、`--groups`、
|
||||
`--patterns`、`--annotations` 匯入,性別修正經
|
||||
`--gender-corrections` 套用,與其餘資料同一交易,儲存不會
|
||||
半建;詳見 `research-plan.md`「資料儲存與模型」。
|
||||
- **Commit 判準**:能由「committed 輸入+程式」決定性再生者不
|
||||
commit(SQLite 工作儲存、LLM 輸入檔);源頭、捕捉、人工著作
|
||||
一律以文字 commit。「可再生仍 commit」的例外有二:
|
||||
`results/` 報表(引用穩定性、審稿人零門檻、撰稿期可 diff)
|
||||
與 `data/derived/` 人讀報表(與工作儲存同一動作產出,稽核
|
||||
鏈無中間空缺)。詳見 `research-plan.md`「資料儲存與模型」。
|
||||
- **設定**經 pydantic-settings 統一:`.env`(gitignored,範本
|
||||
`tools/.env.example`)供應 `SQLALCHEMY_DATABASE_URL` 與
|
||||
`ANTHROPIC_API_KEY`,絕不寫入 repo。
|
||||
- **不設 CLAUDE.md**:實測證實 Claude Code subagent 會繼承專案
|
||||
CLAUDE.md 全文(原本因此只放極簡工作規則),2026-08-18 進一步
|
||||
將其移為 `docs/conventions.md`——docs/ 不會自動注入 subagent
|
||||
的 context,盲判型 agent 便不會看到工作規範;主會話動手前
|
||||
自行閱讀之。
|
||||
+4
-15
@@ -1,8 +1,7 @@
|
||||
# 信度:量測方式與結果
|
||||
|
||||
(2026-08-16 量測。對象為步驟 3a 的三份執行歸檔
|
||||
`runs/3a-code/run1`–`run3`,與步驟 5d 的三份執行歸檔
|
||||
`runs/5d-annotate/run1`–`run3`;數字由原始輸出直接計算。)
|
||||
(2026-08-16 量測。對象為步驟 3a 與步驟 5d 各三次執行的
|
||||
原始輸出;數字由原始輸出直接計算。)
|
||||
|
||||
## 一、本研究的信度是什麼
|
||||
|
||||
@@ -34,11 +33,7 @@ Krippendorff 依產生資料的設計,把信度分成三型
|
||||
**精密度(precision)**,而**準確度(accuracy)** 未經量測。
|
||||
若儀器是確定性的,重複性無須量測;但 LLM 不是——同樣的提示、
|
||||
同樣的輸入,三次會給出不同結果,因此重複性是必須報告的儀器
|
||||
規格,而非慣例儀式。
|
||||
|
||||
這一點與 LLM 標註的近期文獻一致:同一模型重複取樣量測的是
|
||||
自我一致性(self-consistency),而重複執行後取多數決可提升
|
||||
標註穩定度(如 Prompt Stability Scoring,arXiv:2407.02039)。
|
||||
規格。
|
||||
|
||||
## 二、三個常用指標
|
||||
|
||||
@@ -106,7 +101,7 @@ Jaccard(只看「有標到」的格子,忽略雙方都沒標的)約 0.89–0.91,
|
||||
比百分比一致率低而更誠實——因為 82% 的格子是雙方都判「無」,
|
||||
那些一致並不費力。
|
||||
|
||||
## 四、隨機性的規模與三票制的作用
|
||||
## 四、隨機性的規模
|
||||
|
||||
信度數字回答的實際問題是:**這台儀器的隨機性,大到會不會
|
||||
改變結論?**
|
||||
@@ -116,12 +111,6 @@ Jaccard(只看「有標到」的格子,忽略雙方都沒標的)約 0.89–0.91,
|
||||
一半會成為誤收、另一半會成為漏收。三票多數決把兩票以上者
|
||||
收入、一票者剔除,處理的正是這一批。
|
||||
|
||||
三票制並非消除隨機性,而是**把隨機性往案例原本的傾向推**。
|
||||
設某個邊緣案例的符合程度為 p,單次執行以機率 p 標出,三次
|
||||
多數決則以 p³+3p²(1−p) 標出:p=0.9 者由 0.9 提高到 0.972,
|
||||
p=0.1 者由 0.1 壓低到 0.028,而 p=0.5 者仍是 0.5——真正
|
||||
模稜兩可的案例,任何票制都救不了。
|
||||
|
||||
以此規模判斷,結論層的三條帶狀結構(女性力量與陽剛群共現、
|
||||
與脆弱群互斥、與厭女群獨立)不可能由這個量級的雜訊翻轉。
|
||||
反過來說,若一致率只有 0.6,同一組結論就不能採信——信度
|
||||
|
||||
+20
-85
@@ -11,81 +11,33 @@
|
||||
**不與先導研究做比較**;
|
||||
全文數字一律以正式研究結果為準。先導研究僅作為假說的
|
||||
內部來源,記於決策日誌,不進入論文敘事。
|
||||
- **執行原則**:主會話只做討論;所有分析由 deterministic
|
||||
script 執行。LLM 步驟以 Python script 呼叫 Anthropic
|
||||
Messages API(個人 Console 帳號、Batch API 五折),定義
|
||||
檔逐字作為 system prompt。可逐項機械比對的判斷(步驟 3
|
||||
編碼、步驟 4 選群、步驟 5d 樣態標註)採「同一定義檔
|
||||
獨立執行三次+多數決」,計票由確定性子命令完成。自由
|
||||
生成(步驟 1 自由標註)兩次執行全數進池。步驟 5a 至
|
||||
5c 為質性閱讀協定:三次獨立閱讀為分析者三角檢核、
|
||||
逐首整合、樣態統整產出草稿,不適用投票與仲裁。詞彙表
|
||||
不經 LLM,由詞向量嵌入+確定性分群產生。驗證結果不符
|
||||
預期則修訂定義檔重跑該循環,絕不手改結果。
|
||||
- **執行原則**:主會話只做討論,分析一律由確定性程序
|
||||
執行。可逐項機械比對的判斷(步驟 3 編碼、步驟 4 選群、
|
||||
步驟 5d 樣態標註)採「三次獨立執行+多數決」定案,
|
||||
自由生成(步驟 1 自由標註)採兩次執行全數進池,步驟
|
||||
5a 至 5c 定為質性閱讀協定;程序細節見 `methodology.md`,
|
||||
工作規範見 `conventions.md`。
|
||||
- **提示詞只定格式、不定語意**:研究對象是通用 LLM 以其
|
||||
網路語料知識背景所做的自然編碼與閱讀,其結果本身是
|
||||
批判對象。定義檔只規定任務形狀(輸入、數量範圍、輸出
|
||||
格式),不給任何主題的定義、判準或範例。LLM 判斷一律
|
||||
要求逐項引述歌詞原句,作為檢視偏差的依據。
|
||||
- **模型**:步驟 1、3 用 `claude-sonnet-4-6`(temperature=0、
|
||||
thinking 關閉);步驟 4、5 用 `claude-fable-5`(兩參數
|
||||
不適用,均不送出;取樣變異由多數決或整合吸收)——實測
|
||||
發現 sonnet 將「Women Power」拆讀為 women+power 的組合
|
||||
語意,fable-5 讀為詞彙化概念,語意層任務因此換用
|
||||
fable-5(經過見決策日誌)。
|
||||
- **模型**:步驟 1、3 用 `claude-sonnet-4-6`;步驟 4、5
|
||||
用 `claude-fable-5`——實測發現 sonnet 將「Women Power」
|
||||
拆讀為 women+power 的組合語意,fable-5 讀為詞彙化
|
||||
概念,語意層任務因此換用 fable-5(經過見決策日誌)。
|
||||
- **信度與效度**:三次獨立執行量測穩定性(intra-rater
|
||||
reliability,可報告兩兩一致率);封閉母體全量檢查取代
|
||||
抽樣防衛。可重現性定義為「程序透明+可稽核」:公開定義
|
||||
檔、記錄 model ID 與執行時間、保存全部原始輸出。
|
||||
|
||||
## 資料儲存與模型
|
||||
## 範圍性定案
|
||||
|
||||
- **Commit 判準**:凡能由「committed 的輸入+committed 的
|
||||
程式」決定性再生者,不 commit;凡不能者——源頭資料、
|
||||
外部世界的捕捉(Wikidata 快照、LLM 原始輸出)、人工
|
||||
著作——一律以文字格式 commit。格式跟著層次走。
|
||||
- **分層**:
|
||||
- 源頭:原始榜單 CSV(進 git)。
|
||||
- 捕捉:Wikidata 快照 CSV、`runs/` JSONL(皆進 git);
|
||||
歌詞 `.txt` 快取(版權因素 gitignored,為已知的稽核
|
||||
缺口)。
|
||||
- 人工:`data/manual/`(僅研究者親手寫入:編碼修正、
|
||||
演唱者性別修正)。
|
||||
- 工作儲存:SQLite 單檔(`tools/instance/`,generated、
|
||||
不進 git),SQLAlchemy 2.0 typed ORM 定義 schema,
|
||||
設定經 pydantic-settings(`.env` 供應
|
||||
`SQLALCHEMY_DATABASE_URL` 與 `ANTHROPIC_API_KEY`)。
|
||||
歌詞全文入 DB(不進 git 故無版權疑慮)。
|
||||
- 衍生:`build-db` 建置工作儲存的同一動作產出人讀報表
|
||||
(`data/derived/`,進 git),與 SQLite 同交易語意。
|
||||
- 報表:論文引用的定案表進 `results/`——
|
||||
`codings.csv`(步驟 3 定案編碼)、`groups.csv`
|
||||
(步驟 4 定案編碼群)、`patterns.csv`(步驟 5c 定案
|
||||
樣態表)、`annotations.csv`(步驟 5d 定案歌×樣態
|
||||
矩陣)。衍生與報表為「可再生仍 commit」的例外,理由:
|
||||
引用穩定性、審稿人零門檻、撰稿期數字變動可 diff。
|
||||
- **資料模型**:`songs`(含 lyrics、`performer_gender`——
|
||||
演唱聲音之性別,由署名藝人之 Wikidata 性別推導後套用
|
||||
`data/manual/performer-gender-corrections.csv` 手工修正)、
|
||||
`chart_entries`、`artists`、`song_artists`、`codings`、
|
||||
`groups`(語意編碼群)、`patterns`(深讀樣態)、
|
||||
`annotations`(歌×樣態定案矩陣)。領域不變量(恰 1000
|
||||
筆榜單、每歌至少一 primary 歌手等)檢查內建於
|
||||
`build-db`,違規即建置失敗。
|
||||
- **歌手背景防火牆**:歌手背景資料只進人工解讀階段,
|
||||
**絕不進 LLM 輸入**——LLM 任務的 user message 維持
|
||||
歌詞-only 或管線中間產物-only,避免光環偏誤。
|
||||
- **Pilot 歌詞沿用(私人匯入,不進發布管線)**:先導研究
|
||||
捕捉檔(lyrics.json,684 首,2018–2025)以 `excludes/`
|
||||
的私人腳本匯入歌詞快取;讀者的重現路徑純粹是
|
||||
`fetch-lyrics`;沿用之事實記於
|
||||
`data/captures/lyrics-provenance.csv`(進 git)。
|
||||
- **子命令**(`pop-fem-audit-tools <cmd>`):`build-db`
|
||||
(`--codings`、`--groups`、`--gender-corrections`、
|
||||
`--patterns`、`--annotations`)、`cluster-keywords`、
|
||||
`export-llm-input`、`fetch-artists`、`fetch-lyrics`、
|
||||
`run-llm`(`--model` 於模型登錄表中擇一)、
|
||||
`tally-codings`、`tally-groups`、`tally-annotations`。
|
||||
- **Pilot 歌詞沿用**:先導歌詞沿用之來歷與細節詳見該檔
|
||||
(`pilot-study.md`)。
|
||||
|
||||
## 分析管線(五步驟,全部完成)
|
||||
|
||||
@@ -96,27 +48,19 @@
|
||||
階層式聚合分群 k=100,組名取 medoid;再併入研究者先驗
|
||||
主題詞 `women-power`(據實揭露的儀器介入),共 101 碼。
|
||||
3. **編碼(步驟 3)**:以定稿詞彙表對全 883 首編碼,逐
|
||||
標籤附引述,×3+多數決(`tally-codings`),定案
|
||||
`results/codings.csv`。「女性力量」候選集(wp 66 首、
|
||||
fe 144 首、wp∪fe 145 首)由此浮現。
|
||||
標籤附引述,×3+多數決,產出定案編碼表。「女性力量」
|
||||
候選集(wp 66 首、fe 144 首、wp∪fe 145 首)由此浮現。
|
||||
4. **語意編碼群(步驟 4)**:LLM 依編碼字面語意將 101 碼
|
||||
選入研究者指定的四個主題群(women-power/misogyny/
|
||||
masculine/vulnerable),×3+多數決(`tally-groups`),
|
||||
定案 `results/groups.csv`;wp/fe 與各編碼、各編碼群之
|
||||
關聯統計(BH-FDR 校正)入論文。
|
||||
masculine/vulnerable),×3+多數決,產出定案分群表;
|
||||
wp/fe 與各編碼、各編碼群之關聯統計(BH-FDR 校正)入
|
||||
論文。
|
||||
5. **女性主義問題之質性深讀(步驟 5)**:對 wp∪fe 145 首
|
||||
——5a 逐首盲讀(僅歌詞全文)×3;5b 逐首整合(收斂
|
||||
註記、主清單限兩讀以上);5c 樣態統整(全體基底+
|
||||
男聲/女聲/混合三個發話脈絡分組);5d 樣態標註——
|
||||
以分組樣態表逐首標註「有問題」的 111 首,×3+多數決
|
||||
(`tally-annotations`),定案 `results/patterns.csv` 與
|
||||
`results/annotations.csv`。
|
||||
|
||||
定義檔命名 `prompts/<步><次步>-<task>.md`,次步以字母標示
|
||||
(與論文正文的步驟編號一致);編號的所指是工序而非定義檔
|
||||
(確定性的第 2 步無定義檔仍佔編號);檔名
|
||||
不帶版本號——版本即 git 歷史;每次執行的定義檔快照隨
|
||||
`runs/` 自我完備,token 費用逐筆記於 `docs/run-costs.md`。
|
||||
以分組樣態表逐首標註「有問題」的 111 首,×3+多數決,
|
||||
產出定案樣態表與歌×樣態定案矩陣。
|
||||
|
||||
## 時程與剩餘工作
|
||||
|
||||
@@ -125,12 +69,3 @@
|
||||
污染」結論(步驟 5 素材已備)、信度說明、fe 與新自由
|
||||
主義敘事之詮釋標註。
|
||||
|
||||
## 其他已定案事項
|
||||
|
||||
- 歌詞受版權保護:完整歌詞不進 git
|
||||
(`data/captures/lyrics/` gitignored),論文與 repo 只留
|
||||
分析所引摘錄。
|
||||
- 工程支援(寫 script、style check、審稿)用公司訂閱的
|
||||
Claude Code;進論文的分析 token 由個人 API 帳號支付。
|
||||
- Positionality statement 寫入論文(53 歲、長年婦運者、
|
||||
資深工程師、離開學術圈多年、因整理榜單而起)。
|
||||
|
||||
+1
-1
@@ -48,7 +48,7 @@ LLM 合併嘗試(`01-02-01-merge`,詞彙表改採詞向量分群時
|
||||
| 2026-08-06 | 已廢棄 | — | claude-sonnet-4-6 | msgbatch_01N7bDbXRSfAVUzzaj2thKeR | 4 分 20 秒 | 781,037 | 39,315 | $1.47 | 現行(644 首全數有效,零攔阻;保留 1,481/送裁 1,699) | 03-02-arbitration |
|
||||
| 2026-08-06 | 3a-code | run3 | claude-sonnet-4-6 | msgbatch_01KnkCaGETnFJrPddrxZTYHA | 6 分 26 秒 | 1,625,458 | 363,840 | $5.17 | 現行(101 碼;883 首全數有效,零攔阻) | 03-code |
|
||||
| 2026-08-14 | 4-group | run1 | claude-sonnet-4-6 | msgbatch_01UPNedog6feQzJ9WVfSAxBD | 1 分 1 秒 | 3,129 | 470 | $0.01 | 已取代(僅 3 群;改納 women-power 群後重跑) | 04-group |
|
||||
| 2026-08-14 | 4-group | run1 | claude-sonnet-4-6 | msgbatch_01KntJgxdicStaMjNL12P3zi | 1 分 25 秒 | 4,170 | 558 | $0.01 | 已取代(改以 claude-fable-5 執行;vulnerable 輸出含詞彙表外碼 1 筆;歸檔備份於 excludes/experiments/) | 04-group |
|
||||
| 2026-08-14 | 4-group | run1 | claude-sonnet-4-6 | msgbatch_01KntJgxdicStaMjNL12P3zi | 1 分 25 秒 | 4,170 | 558 | $0.01 | 已取代(改以 claude-fable-5 執行;vulnerable 輸出含詞彙表外碼 1 筆;歸檔另行私人備份,不入版本庫) | 04-group |
|
||||
| 2026-08-14 | 4-group | run1 | claude-fable-5 | msgbatch_01Mr6goBb2Efa4YrqCprbP4U | 55 秒 | 5,553 | 2,049 | $0.08 | 現行(4 群;零違規碼;temperature 與 thinking 參數不適用於本模型,未送出) | 04-group |
|
||||
| 2026-08-14 | 4-group | run2 | claude-fable-5 | msgbatch_01C17xW3YBefThTYZ83g7KkL | 1 分 21 秒 | 5,553 | 1,891 | $0.08 | 現行(4 群;零違規碼) | 04-group |
|
||||
| 2026-08-14 | 4-group | run3 | claude-fable-5 | msgbatch_01DveEMYyjCAYe6wxpcCD87V | 2 分 9 秒 | 5,553 | 1,975 | $0.08 | 現行(4 群;零違規碼) | 04-group |
|
||||
|
||||
+2
-2
@@ -3,7 +3,7 @@
|
||||
# Authors:
|
||||
# imacat@mail.imacat.idv.tw (imacat), 2026/7/31
|
||||
|
||||
# The SQLAlchemy database URL.
|
||||
SQLALCHEMY_DATABASE_URL="postgresql://user:password@host/db"
|
||||
# The SQLAlchemy database URI.
|
||||
SQLALCHEMY_DATABASE_URI="postgresql://user:password@host/db"
|
||||
# The Anthropic API key
|
||||
ANTHROPIC_API_KEY=sk-ant-...
|
||||
|
||||
@@ -0,0 +1,10 @@
|
||||
Change Log
|
||||
==========
|
||||
|
||||
|
||||
version 1.0.0
|
||||
-------------
|
||||
|
||||
Released 2026/8/19
|
||||
|
||||
Initial release.
|
||||
@@ -13,6 +13,8 @@ This is a collection of supporting tools for the conference paper "流行音樂
|
||||
:maxdepth: 2
|
||||
:caption: Contents:
|
||||
|
||||
changelog
|
||||
|
||||
|
||||
Indices and tables
|
||||
==================
|
||||
|
||||
@@ -6,5 +6,5 @@
|
||||
"""Tools for A Feminist Audit of Pop Music."""
|
||||
|
||||
|
||||
VERSION: str = "0.0.0"
|
||||
VERSION: str = "1.0.0"
|
||||
"""The package version."""
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -24,8 +24,7 @@ keyword set for ``export-llm-input --extras`` is written as a JSON
|
||||
file holding the group name keywords plus every extra a-priori
|
||||
keyword the caller gives with the repeatable ``--extra-keyword``
|
||||
command-line option, as
|
||||
:attr:`KeywordsToMerge.KEYWORDS_TO_MERGE_JSON`; with no
|
||||
``--extra-keyword``, it holds the group names alone. No default
|
||||
:attr:`KeywordsToMerge.KEYWORDS_TO_MERGE_JSON`. No default
|
||||
extra keyword is ever injected; the caller supplies each one
|
||||
consciously. Finally, the command-line choices and the
|
||||
environment that produced the numbers -- neither recoverable from
|
||||
@@ -160,15 +159,8 @@ class KeywordPooler:
|
||||
if line.strip() == "":
|
||||
continue
|
||||
record: Any = json.loads(line)
|
||||
if not isinstance(record, dict) or "id" not in record:
|
||||
raise ValueError(
|
||||
f"{path}: record without \"id\": {line}")
|
||||
if "error" in record:
|
||||
continue
|
||||
if "text" not in record:
|
||||
raise ValueError(
|
||||
f"{path}: id {record['id']}: record without"
|
||||
" \"text\" or \"error\"")
|
||||
song_id: int = cls.__parse_song_id(record["id"], path)
|
||||
try:
|
||||
keywords: Any = json.loads(
|
||||
@@ -297,7 +289,8 @@ class KeywordGroups:
|
||||
class KeywordClusterer:
|
||||
"""The clusterer of the pooled keywords into coding groups."""
|
||||
|
||||
DEFAULT_MODEL: str = "sentence-transformers/all-mpnet-base-v2"
|
||||
DEFAULT_MODEL: ClassVar[str] \
|
||||
= "sentence-transformers/all-mpnet-base-v2"
|
||||
"""The sentence embedding model used when the caller names
|
||||
none."""
|
||||
|
||||
@@ -668,11 +661,7 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
parser.add_argument(
|
||||
"output_dir", type=Path,
|
||||
help="the output directory, created if missing, that"
|
||||
f" receives {PooledKeywords.SOURCE_KEYWORDS_TXT},"
|
||||
f" {KeywordGroups.RESULT_KEYWORDS_TXT},"
|
||||
f" {KeywordGroups.RESULT_GROUPS_CSV},"
|
||||
f" {KeywordsToMerge.KEYWORDS_TO_MERGE_JSON},"
|
||||
f" and {RunMeta.META_JSON}")
|
||||
" receives the run's output artifacts")
|
||||
parser.add_argument(
|
||||
"--model", default=model,
|
||||
help=f"the sentence embedding model (default \"{model}\")")
|
||||
@@ -698,16 +687,9 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
"""Pool the two tagging runs' keywords and cluster them.
|
||||
|
||||
Writes the five fixed-named artifacts under the output
|
||||
directory, creating it (with parents) if it does not exist:
|
||||
the pooled keyword text file; then the group membership CSV
|
||||
file, holding the clustering result alone; the group name
|
||||
keyword text file, holding the same group names as a readable
|
||||
list; the coding keyword set JSON file, holding the group
|
||||
names plus every extra keyword given via ``--extra-keyword``;
|
||||
and the run metadata JSON file, recording the command-line
|
||||
choices and the environment. Each file is written as soon as
|
||||
its content is computed, so when the input is rejected, or an
|
||||
Creates the output directory (with parents) if it does not
|
||||
exist. Each output artifact is written as soon as its
|
||||
content is computed, so when the input is rejected, or an
|
||||
extra keyword duplicates a group name or another extra
|
||||
keyword, the output directory holds whatever the steps before
|
||||
the failing one produced, and the error message names what
|
||||
@@ -719,8 +701,8 @@ def main(argv: list[str] | None = None) -> int:
|
||||
"""
|
||||
started: float = time.monotonic()
|
||||
args: argparse.Namespace = parse_args(argv)
|
||||
args.output_dir.mkdir(parents=True, exist_ok=True)
|
||||
try:
|
||||
args.output_dir.mkdir(parents=True, exist_ok=True)
|
||||
source: PooledKeywords = KeywordPooler(
|
||||
args.run_dir_1, args.run_dir_2, args.output_dir).run()
|
||||
clusters: KeywordGroups = KeywordClusterer(
|
||||
@@ -735,7 +717,7 @@ def main(argv: list[str] | None = None) -> int:
|
||||
f"Done. Clustered {len(source.keywords)} keywords into"
|
||||
f" {len(clusters.names)}. {elapsed} elapsed.",
|
||||
file=sys.stderr)
|
||||
except ClusterError as error:
|
||||
except (ClusterError, OSError) as error:
|
||||
print(f"error: {error}", file=sys.stderr)
|
||||
return 1
|
||||
return 0
|
||||
|
||||
@@ -11,26 +11,17 @@ project's lyrics-only firewall: the output carries only the
|
||||
lyrics text of each song, identified by an opaque song key; no
|
||||
title, artist, or chart data crosses into the LLM input.
|
||||
|
||||
With ``--extras``, each record's ``content`` becomes a JSON
|
||||
object serialized as a string, its ``lyrics`` key holding the
|
||||
song's lyrics followed by the keys of the given extras file in
|
||||
their file order, so a step that needs parameters alongside the
|
||||
lyrics can carry them without this module knowing what they mean.
|
||||
|
||||
With ``--extras-per-id``, the same merge happens per song: the
|
||||
given file maps a song ID to the extra keys of that one song, and
|
||||
the export is restricted to the song IDs the file names, so a step
|
||||
that revisits only some of the songs, each with its own parameters,
|
||||
gets exactly those records. The two options may be given together,
|
||||
in which case a record's keys are ``lyrics``, the shared extras'
|
||||
keys, then that song's own keys, each group in its file order.
|
||||
With ``--extras`` and ``--extras-per-id``, a record's content may
|
||||
carry extra parameters alongside the lyrics, so a step that needs
|
||||
them can get them without this module knowing what they mean; see
|
||||
the exporter's content-building step for how the two merge.
|
||||
"""
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
from typing import Any, ClassVar
|
||||
|
||||
import sqlalchemy as sa
|
||||
from sqlalchemy.orm import Session
|
||||
@@ -40,6 +31,270 @@ from ..models import Song
|
||||
from ..utils import format_duration
|
||||
|
||||
|
||||
class LlmInputExporter:
|
||||
"""The exporter of the LLM input JSONL file."""
|
||||
|
||||
__LYRICS_KEY: ClassVar[str] = "lyrics"
|
||||
"""The key holding the lyrics in a record's merged content,
|
||||
and the key forbidden in an extras file."""
|
||||
|
||||
def __init__(
|
||||
self, output_jsonl: Path, extras: Path | None = None,
|
||||
extras_per_id: Path | None = None) -> None:
|
||||
"""Set up the exporter of the LLM input JSONL file.
|
||||
|
||||
:param output_jsonl: The JSONL output file.
|
||||
:param extras: The extras JSON file, or None for none.
|
||||
:param extras_per_id: The per-ID extras JSON file, or
|
||||
None for none.
|
||||
"""
|
||||
self.__output_jsonl: Path = output_jsonl
|
||||
"""The JSONL output file."""
|
||||
self.__extras_path: Path | None = extras
|
||||
"""The extras JSON file, or None for none."""
|
||||
self.__extras_per_id_path: Path | None = extras_per_id
|
||||
"""The per-ID extras JSON file, or None for none."""
|
||||
|
||||
def run(self) -> int:
|
||||
"""Export the songs' lyrics to the output JSONL file.
|
||||
|
||||
:return: The number of songs exported.
|
||||
:raises OSError: When a file cannot be read or written.
|
||||
:raises sqlalchemy.exc.SQLAlchemyError: When the working
|
||||
store cannot be read.
|
||||
:raises ValueError: When an extras file is malformed, an
|
||||
exported song has no lyrics, or the per-ID extras
|
||||
name a song the working store does not have.
|
||||
"""
|
||||
session: Session = ds.get_db()
|
||||
try:
|
||||
extras: dict[str, Any] | None = None
|
||||
if self.__extras_path is not None:
|
||||
extras = self.__load_extras(self.__extras_path)
|
||||
extras_per_id: dict[str, dict[str, Any]] | None = None
|
||||
if self.__extras_per_id_path is not None:
|
||||
extras_per_id = self.__load_extras_per_id(
|
||||
self.__extras_per_id_path)
|
||||
lines: list[str] = self.__build_lines(
|
||||
session, extras, extras_per_id)
|
||||
finally:
|
||||
session.close()
|
||||
self.__write_output(lines)
|
||||
return len(lines)
|
||||
|
||||
@staticmethod
|
||||
def __no_duplicate_keys(
|
||||
pairs: list[tuple[str, Any]]) -> dict[str, Any]:
|
||||
"""Build a dict from JSON object pairs, rejecting
|
||||
duplicates.
|
||||
|
||||
:param pairs: The key-value pairs of a JSON object, in
|
||||
file order.
|
||||
:return: The pairs as a dict, in file order.
|
||||
:raises ValueError: When a key appears more than once.
|
||||
"""
|
||||
result: dict[str, Any] = {}
|
||||
key: str
|
||||
value: Any
|
||||
for key, value in pairs:
|
||||
if key in result:
|
||||
raise ValueError(
|
||||
f"duplicate key \"{key}\" in extras")
|
||||
result[key] = value
|
||||
return result
|
||||
|
||||
@classmethod
|
||||
def __load_json_object(
|
||||
cls, path: Path, label: str) -> dict[str, Any]:
|
||||
"""Load a single JSON object from a file, in file order.
|
||||
|
||||
:param path: The JSON file.
|
||||
:param label: The kind of file, for the error messages.
|
||||
:return: The object, in file order.
|
||||
:raises OSError: When the file cannot be read.
|
||||
:raises ValueError: When the file is not valid JSON, is
|
||||
not a JSON object, or has duplicate keys.
|
||||
"""
|
||||
with open(path, encoding="utf-8") as file:
|
||||
text: str = file.read()
|
||||
try:
|
||||
data: Any = json.loads(
|
||||
text, object_pairs_hook=cls.__no_duplicate_keys)
|
||||
except json.JSONDecodeError as error:
|
||||
raise ValueError(
|
||||
f"invalid JSON in {label} file {path}: {error}") \
|
||||
from error
|
||||
if not isinstance(data, dict):
|
||||
raise ValueError(
|
||||
f"{label} file {path} must contain a JSON object")
|
||||
return data
|
||||
|
||||
@classmethod
|
||||
def __load_extras(cls, path: Path) -> dict[str, Any]:
|
||||
"""Load the extras object from a JSON file.
|
||||
|
||||
:param path: The extras JSON file.
|
||||
:return: The extras, in file order.
|
||||
:raises OSError: When the file cannot be read.
|
||||
:raises ValueError: When the file is not valid JSON, is
|
||||
not a JSON object, has duplicate keys, or has a
|
||||
"lyrics" key.
|
||||
"""
|
||||
data: dict[str, Any] = cls.__load_json_object(
|
||||
path, "extras")
|
||||
if cls.__LYRICS_KEY in data:
|
||||
raise ValueError(
|
||||
f"extras file {path} must not have a"
|
||||
f" \"{cls.__LYRICS_KEY}\" key")
|
||||
return data
|
||||
|
||||
@classmethod
|
||||
def __load_extras_per_id(
|
||||
cls, path: Path) -> dict[str, dict[str, Any]]:
|
||||
"""Load the per-ID extras object from a JSON file.
|
||||
|
||||
:param path: The per-ID extras JSON file, mapping a song
|
||||
ID, as ``song-<N>``, to the extras of that one song.
|
||||
:return: The extras of each song ID, in file order, every
|
||||
song's own extras in their file order too.
|
||||
:raises OSError: When the file cannot be read.
|
||||
:raises ValueError: When the file is not valid JSON, is
|
||||
not a JSON object, has duplicate keys, has a song
|
||||
whose value is not a JSON object, or has a song with
|
||||
a "lyrics" key.
|
||||
"""
|
||||
data: dict[str, Any] = cls.__load_json_object(
|
||||
path, "per-ID extras")
|
||||
song_id: str
|
||||
extras: Any
|
||||
for song_id, extras in data.items():
|
||||
if not isinstance(extras, dict):
|
||||
raise ValueError(
|
||||
f"per-ID extras file {path}: id {song_id}"
|
||||
" must have a JSON object")
|
||||
if cls.__LYRICS_KEY in extras:
|
||||
raise ValueError(
|
||||
f"per-ID extras file {path}: id {song_id}"
|
||||
f" must not have a \"{cls.__LYRICS_KEY}\""
|
||||
" key")
|
||||
return data
|
||||
|
||||
def __build_lines(
|
||||
self, session: Session,
|
||||
extras: dict[str, Any] | None = None,
|
||||
extras_per_id: dict[str, dict[str, Any]] | None
|
||||
= None) -> list[str]:
|
||||
"""Build the JSONL lines of the exported songs' lyrics.
|
||||
|
||||
Every song is exported, unless per-ID extras are given,
|
||||
in which case only the songs they name are; see
|
||||
:meth:`__build_content` for how the extras merge into a
|
||||
record's content.
|
||||
|
||||
:param session: The database session.
|
||||
:param extras: The extra parameters merged into every
|
||||
record's content alongside the lyrics, in the order
|
||||
they are to appear, or None for none.
|
||||
:param extras_per_id: The extra parameters merged into
|
||||
the content of one record alone, keyed by that
|
||||
record's song ID and in the order they are to appear,
|
||||
restricting the export to the song IDs they name, or
|
||||
None for no such extras and no such restriction.
|
||||
:return: The JSON lines, one per exported song, ordered by
|
||||
song ID.
|
||||
:raises ValueError: When an exported song has no lyrics,
|
||||
or the per-ID extras name a song the working store
|
||||
does not have.
|
||||
"""
|
||||
lines: list[str] = []
|
||||
exported: set[str] = set()
|
||||
song: Song
|
||||
for song in session.scalars(
|
||||
sa.select(Song).order_by(Song.id)):
|
||||
song_id: str = f"song-{song.id}"
|
||||
if extras_per_id is not None \
|
||||
and song_id not in extras_per_id:
|
||||
continue
|
||||
if song.lyrics is None:
|
||||
raise ValueError(
|
||||
f"song {song.id} \"{song.title}\": no lyrics")
|
||||
song_extras: dict[str, Any] | None = None \
|
||||
if extras_per_id is None \
|
||||
else extras_per_id[song_id]
|
||||
content: str = self.__build_content(
|
||||
song.lyrics, extras, song_extras)
|
||||
record: dict[str, str] = {
|
||||
"id": song_id, "content": content}
|
||||
lines.append(json.dumps(record, ensure_ascii=False))
|
||||
exported.add(song_id)
|
||||
if extras_per_id is not None:
|
||||
missing: list[str] = sorted(
|
||||
set(extras_per_id) - exported)
|
||||
if len(missing) > 0:
|
||||
raise ValueError(
|
||||
"the per-ID extras name songs the working"
|
||||
f" store does not have: {', '.join(missing)}")
|
||||
return lines
|
||||
|
||||
@classmethod
|
||||
def __build_content(
|
||||
cls, lyrics: str, extras: dict[str, Any] | None,
|
||||
song_extras: dict[str, Any] | None) -> str:
|
||||
"""Build the content of one exported record.
|
||||
|
||||
Without extras of either kind, a record's content is the
|
||||
bare lyrics string. With ``--extras``, the content
|
||||
becomes a JSON object serialized as a string, its
|
||||
"lyrics" key holding the song's lyrics followed by the
|
||||
keys of the given extras file, in their file order. With
|
||||
``--extras-per-id``, the same merge happens per song: the
|
||||
song's own extra keys follow the lyrics instead. When
|
||||
both are given, a record's keys are "lyrics", the shared
|
||||
extras' keys, then that song's own keys, each group in
|
||||
its file order.
|
||||
|
||||
:param lyrics: The lyrics of the song.
|
||||
:param extras: The extra parameters shared by every
|
||||
record, in the order they are to appear, or None for
|
||||
none.
|
||||
:param song_extras: The extra parameters of this record
|
||||
alone, in the order they are to appear, or None for
|
||||
none.
|
||||
:return: The bare lyrics when there are no extras of
|
||||
either kind, or otherwise a JSON object serialized as
|
||||
a string, whose first key is "lyrics" holding the
|
||||
lyrics, followed by the shared extras' keys and then
|
||||
this record's own keys, each group in its given
|
||||
order.
|
||||
"""
|
||||
if extras is None and song_extras is None:
|
||||
return lyrics
|
||||
payload: dict[str, Any] = {cls.__LYRICS_KEY: lyrics}
|
||||
if extras is not None:
|
||||
payload.update(extras)
|
||||
if song_extras is not None:
|
||||
payload.update(song_extras)
|
||||
return json.dumps(payload, ensure_ascii=False)
|
||||
|
||||
def __write_output(self, lines: list[str]) -> None:
|
||||
"""Write the exported lines to the output JSONL file.
|
||||
|
||||
Creates the parent directory when it does not exist.
|
||||
|
||||
:param lines: The JSONL lines, in the output order.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
"""
|
||||
self.__output_jsonl.parent.mkdir(
|
||||
parents=True, exist_ok=True)
|
||||
with open(
|
||||
self.__output_jsonl, "w",
|
||||
encoding="utf-8") as file:
|
||||
line: str
|
||||
for line in lines:
|
||||
file.write(line + "\n")
|
||||
|
||||
|
||||
def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
"""Parse the command-line arguments.
|
||||
|
||||
@@ -56,227 +311,33 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
parser.add_argument(
|
||||
"--extras", type=Path, default=None,
|
||||
help="a JSON file holding a single JSON object of extra"
|
||||
" parameters; when given, each record's \"content\""
|
||||
" becomes a JSON object string with a \"lyrics\" key"
|
||||
" followed by the extras' keys, instead of the bare"
|
||||
" lyrics string")
|
||||
" parameters merged into every record's content")
|
||||
parser.add_argument(
|
||||
"--extras-per-id", type=Path, default=None,
|
||||
help="a JSON file holding a single JSON object that maps a"
|
||||
" song ID, as \"song-<N>\", to a JSON object of extra"
|
||||
" parameters for that one song; the song's object is"
|
||||
" merged into its \"content\" the same way as with"
|
||||
" --extras, and the export is restricted to the song"
|
||||
" IDs the file names")
|
||||
" parameters for that one song, restricting the"
|
||||
" export to the song IDs the file names")
|
||||
return parser.parse_args(argv)
|
||||
|
||||
|
||||
def __no_duplicate_keys(
|
||||
pairs: list[tuple[str, Any]]) -> dict[str, Any]:
|
||||
"""Build a dict from JSON object pairs, rejecting duplicates.
|
||||
|
||||
:param pairs: The key-value pairs of a JSON object, in file
|
||||
order.
|
||||
:return: The pairs as a dict, in file order.
|
||||
:raises ValueError: When a key appears more than once.
|
||||
"""
|
||||
result: dict[str, Any] = {}
|
||||
key: str
|
||||
value: Any
|
||||
for key, value in pairs:
|
||||
if key in result:
|
||||
raise ValueError(f"duplicate key \"{key}\" in extras")
|
||||
result[key] = value
|
||||
return result
|
||||
|
||||
|
||||
def __load_json_object(path: Path, label: str) -> dict[str, Any]:
|
||||
"""Load a single JSON object from a file, in file order.
|
||||
|
||||
:param path: The JSON file.
|
||||
:param label: The kind of file, for the error messages.
|
||||
:return: The object, in file order.
|
||||
:raises OSError: When the file cannot be read.
|
||||
:raises ValueError: When the file is not valid JSON, is not
|
||||
a JSON object, or has duplicate keys.
|
||||
"""
|
||||
with open(path, encoding="utf-8") as file:
|
||||
text: str = file.read()
|
||||
try:
|
||||
data: Any = json.loads(
|
||||
text, object_pairs_hook=__no_duplicate_keys)
|
||||
except json.JSONDecodeError as error:
|
||||
raise ValueError(
|
||||
f"invalid JSON in {label} file {path}: {error}") \
|
||||
from error
|
||||
if not isinstance(data, dict):
|
||||
raise ValueError(
|
||||
f"{label} file {path} must contain a JSON object")
|
||||
return data
|
||||
|
||||
|
||||
def load_extras(path: Path) -> dict[str, Any]:
|
||||
"""Load the extras object from a JSON file.
|
||||
|
||||
:param path: The extras JSON file.
|
||||
:return: The extras, in file order.
|
||||
:raises OSError: When the file cannot be read.
|
||||
:raises ValueError: When the file is not valid JSON, is not
|
||||
a JSON object, has duplicate keys, or has a "lyrics" key.
|
||||
"""
|
||||
data: dict[str, Any] = __load_json_object(path, "extras")
|
||||
if "lyrics" in data:
|
||||
raise ValueError(
|
||||
f"extras file {path} must not have a \"lyrics\" key")
|
||||
return data
|
||||
|
||||
|
||||
def load_extras_per_id(path: Path) -> dict[str, dict[str, Any]]:
|
||||
"""Load the per-ID extras object from a JSON file.
|
||||
|
||||
:param path: The per-ID extras JSON file, mapping a song ID,
|
||||
as ``song-<N>``, to the extras of that one song.
|
||||
:return: The extras of each song ID, in file order, every
|
||||
song's own extras in their file order too.
|
||||
:raises OSError: When the file cannot be read.
|
||||
:raises ValueError: When the file is not valid JSON, is not
|
||||
a JSON object, has duplicate keys, has a song whose value
|
||||
is not a JSON object, or has a song with a "lyrics" key.
|
||||
"""
|
||||
data: dict[str, Any] = __load_json_object(path, "per-ID extras")
|
||||
song_id: str
|
||||
extras: Any
|
||||
for song_id, extras in data.items():
|
||||
if not isinstance(extras, dict):
|
||||
raise ValueError(
|
||||
f"per-ID extras file {path}: id {song_id} must"
|
||||
" have a JSON object")
|
||||
if "lyrics" in extras:
|
||||
raise ValueError(
|
||||
f"per-ID extras file {path}: id {song_id} must not"
|
||||
" have a \"lyrics\" key")
|
||||
return data
|
||||
|
||||
|
||||
def __build_content(
|
||||
lyrics: str,
|
||||
extras: dict[str, Any] | None,
|
||||
song_extras: dict[str, Any] | None) -> str:
|
||||
"""Build the content of one exported record.
|
||||
|
||||
:param lyrics: The lyrics of the song.
|
||||
:param extras: The extra parameters shared by every record,
|
||||
in the order they are to appear, or None for none.
|
||||
:param song_extras: The extra parameters of this record
|
||||
alone, in the order they are to appear, or None for none.
|
||||
:return: The bare lyrics when there are no extras of either
|
||||
kind, or otherwise a JSON object serialized as a string,
|
||||
whose first key is ``"lyrics"`` holding the lyrics,
|
||||
followed by the shared extras' keys and then this
|
||||
record's own keys, each group in its given order.
|
||||
"""
|
||||
if extras is None and song_extras is None:
|
||||
return lyrics
|
||||
payload: dict[str, Any] = {"lyrics": lyrics}
|
||||
if extras is not None:
|
||||
payload.update(extras)
|
||||
if song_extras is not None:
|
||||
payload.update(song_extras)
|
||||
return json.dumps(payload, ensure_ascii=False)
|
||||
|
||||
|
||||
def build_lines(
|
||||
session: Session,
|
||||
extras: dict[str, Any] | None = None,
|
||||
extras_per_id: dict[str, dict[str, Any]] | None = None) \
|
||||
-> list[str]:
|
||||
"""Build the JSONL lines of the exported songs' lyrics.
|
||||
|
||||
Without extras of either kind, each record's ``content`` is
|
||||
the bare lyrics string. With extras, ``content`` is a JSON
|
||||
object serialized as a string, whose first key is ``"lyrics"``
|
||||
holding the lyrics string, followed by the shared extras' keys
|
||||
and then the song's own per-ID extras' keys, each group in its
|
||||
given order.
|
||||
|
||||
Every song is exported, unless per-ID extras are given, in
|
||||
which case only the songs they name are.
|
||||
|
||||
:param session: The database session.
|
||||
:param extras: The extra parameters merged into every
|
||||
record's content alongside the lyrics, in the order they
|
||||
are to appear, or None for none.
|
||||
:param extras_per_id: The extra parameters merged into the
|
||||
content of one record alone, keyed by that record's song
|
||||
ID and in the order they are to appear, restricting the
|
||||
export to the song IDs they name, or None for no such
|
||||
extras and no such restriction.
|
||||
:return: The JSON lines, one per exported song, ordered by
|
||||
song ID.
|
||||
:raises ValueError: When an exported song has no lyrics, or
|
||||
the per-ID extras name a song the working store does not
|
||||
have.
|
||||
"""
|
||||
lines: list[str] = []
|
||||
exported: set[str] = set()
|
||||
song: Song
|
||||
for song in session.scalars(sa.select(Song).order_by(Song.id)):
|
||||
song_id: str = f"song-{song.id}"
|
||||
if extras_per_id is not None and song_id not in extras_per_id:
|
||||
continue
|
||||
if song.lyrics is None:
|
||||
raise ValueError(
|
||||
f"song {song.id} \"{song.title}\": no lyrics")
|
||||
song_extras: dict[str, Any] | None = None \
|
||||
if extras_per_id is None else extras_per_id[song_id]
|
||||
content: str = __build_content(
|
||||
song.lyrics, extras, song_extras)
|
||||
record: dict[str, str] = {
|
||||
"id": song_id, "content": content}
|
||||
lines.append(json.dumps(record, ensure_ascii=False))
|
||||
exported.add(song_id)
|
||||
if extras_per_id is not None:
|
||||
missing: list[str] = sorted(set(extras_per_id) - exported)
|
||||
if len(missing) > 0:
|
||||
raise ValueError(
|
||||
"the per-ID extras name songs the working store"
|
||||
f" does not have: {', '.join(missing)}")
|
||||
return lines
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
"""Export the LLM input JSONL file from the working store.
|
||||
|
||||
Every song is exported, unless ``--extras-per-id`` is given,
|
||||
in which case only the songs its file names are.
|
||||
|
||||
:param argv: The command-line arguments, or None for
|
||||
``sys.argv``.
|
||||
:return: The exit status: 0 on success, non-zero on failure.
|
||||
"""
|
||||
started: float = time.monotonic()
|
||||
args: argparse.Namespace = parse_args(argv)
|
||||
session: Session = ds.get_db()
|
||||
lines: list[str]
|
||||
try:
|
||||
extras: dict[str, Any] | None = None
|
||||
if args.extras is not None:
|
||||
extras = load_extras(args.extras)
|
||||
extras_per_id: dict[str, dict[str, Any]] | None = None
|
||||
if args.extras_per_id is not None:
|
||||
extras_per_id = load_extras_per_id(args.extras_per_id)
|
||||
lines = build_lines(session, extras, extras_per_id)
|
||||
count: int = LlmInputExporter(
|
||||
args.output_jsonl, args.extras,
|
||||
args.extras_per_id).run()
|
||||
except (OSError, sa.exc.SQLAlchemyError, ValueError) as error:
|
||||
print(f"error: {error}", file=sys.stderr)
|
||||
return 1
|
||||
finally:
|
||||
session.close()
|
||||
args.output_jsonl.parent.mkdir(parents=True, exist_ok=True)
|
||||
with open(args.output_jsonl, "w", encoding="utf-8") as file:
|
||||
line: str
|
||||
for line in lines:
|
||||
file.write(line + "\n")
|
||||
elapsed: str = format_duration(time.monotonic() - started)
|
||||
print(f"Done. {len(lines)} songs exported."
|
||||
print(f"Done. {count} songs exported."
|
||||
f" {elapsed} elapsed.", file=sys.stderr)
|
||||
return 0
|
||||
|
||||
@@ -12,12 +12,10 @@ layer. The working store is only read, never written; the
|
||||
``build-db`` subcommand assembles the captured files into the
|
||||
store on the next rebuild.
|
||||
|
||||
Every fetched row is meant for later human verification: the
|
||||
description of the resolved item is recorded in the note column
|
||||
so that a bad match can be spotted. An unresolved artist or an
|
||||
error on one artist is noted on its row and does not fail the
|
||||
run. A row whose name is no longer an artist of the store is
|
||||
dropped from the snapshot and reported on the standard error.
|
||||
An unresolved artist or an error on one artist is noted on its
|
||||
row and does not fail the run. A row whose name is no longer an
|
||||
artist of the store is dropped from the snapshot and reported on
|
||||
the standard error.
|
||||
"""
|
||||
import argparse
|
||||
import csv
|
||||
@@ -33,7 +31,7 @@ import urllib.request
|
||||
from collections.abc import Container, Sequence
|
||||
from dataclasses import asdict, dataclass, field, fields
|
||||
from pathlib import Path
|
||||
from typing import Any, Literal, TextIO
|
||||
from typing import Any, ClassVar, Literal, TextIO
|
||||
|
||||
import sqlalchemy as sa
|
||||
from sqlalchemy.orm import Session
|
||||
@@ -43,71 +41,6 @@ from ..database import ds
|
||||
from ..models import Artist, Song, SongArtist
|
||||
from ..utils import format_duration
|
||||
|
||||
API_URL: str = "https://www.wikidata.org/w/api.php"
|
||||
"""The URL of the Wikidata API endpoint."""
|
||||
SPARQL_URL: str = "https://query.wikidata.org/sparql"
|
||||
"""The URL of the Wikidata Query Service SPARQL endpoint."""
|
||||
USER_AGENT: str = (
|
||||
f"pop-fem-audit-tools/{VERSION}"
|
||||
" (https://github.com/imacat/pop-fem-audit;"
|
||||
" mailto:imacat@mail.imacat.idv.tw)")
|
||||
"""The User-Agent header sent on every HTTP request."""
|
||||
TIMEOUT: float = 30.0
|
||||
"""The timeout of an API HTTP request, in seconds."""
|
||||
SPARQL_TIMEOUT: float = 90.0
|
||||
"""The timeout of a SPARQL HTTP request, in seconds.
|
||||
|
||||
Higher than the API timeout: the WDQS server aborts a slow
|
||||
query at 60 seconds, and a lower client timeout would race
|
||||
that server-side abort and misclassify a slow-but-answerable
|
||||
query as a client-side timeout instead of letting the server's
|
||||
own HTTP error response arrive and enter the retry path."""
|
||||
SLEEP_SECONDS: float = 1.0
|
||||
"""The delay between consecutive HTTP requests, in seconds."""
|
||||
MAX_ATTEMPTS: int = 5
|
||||
"""The maximum number of attempts on a transient error."""
|
||||
RETRY_SECONDS: float = 15.0
|
||||
"""The back-off unit on a transient error, in seconds;
|
||||
multiplied by the attempt number already made."""
|
||||
RETRY_STATUSES: frozenset[int] = frozenset({429, 500, 502, 503})
|
||||
"""The HTTP statuses that are retried with a back-off."""
|
||||
MAX_STAGE1_TITLES: int = 3
|
||||
"""The maximum number of charted titles used for the stage-1 song
|
||||
corroboration."""
|
||||
HUMAN_QID: str = "Q5"
|
||||
"""The Wikidata item ID of "human"."""
|
||||
ENSEMBLE_QID: str = "Q2088357"
|
||||
"""The Wikidata item ID of "musical ensemble"."""
|
||||
ORIGINAL_CAST_QID: str = "Q106497009"
|
||||
"""The Wikidata item ID of "original cast"."""
|
||||
GROUP_KEYWORDS: Sequence[str] = ("band", "group", "duo", "trio")
|
||||
"""The label keywords that suggest a musical ensemble, covering
|
||||
labels like "boy band" and "girl group"."""
|
||||
NOTE_NOT_FOUND: str = "not found"
|
||||
"""The note sentinel of an artist without a resolved Wikidata
|
||||
item, written to the snapshot and read back for the
|
||||
classification."""
|
||||
CORPUS_START_YEAR: int = 2016
|
||||
"""The first year of the corpus window: a member who left a group
|
||||
before it never performed a corpus song."""
|
||||
MIXED_GENDER: str = "mixed"
|
||||
"""The gender recorded for a group whose members do not share one
|
||||
gender."""
|
||||
TIME_YEAR_PATTERN: re.Pattern[str] = re.compile(r"^[+-]?\d+")
|
||||
"""The leading year of a Wikidata time value."""
|
||||
PINNED_QIDS: dict[str, str] = {}
|
||||
"""The last-resort pinned item IDs, keyed by the artist name.
|
||||
|
||||
An entry is for an artist the algorithm documented on
|
||||
``ArtistFetcher`` is structurally unable to resolve, with its
|
||||
justification recorded here. Currently empty: the only pin ever
|
||||
needed, "Pinkfong" (typed as a brand, which the type gate
|
||||
excludes by design), became moot when the store's artist entity
|
||||
behind that credit was identified as Hope Segoine.
|
||||
|
||||
A pinned name skips the candidate retrieval and corroboration
|
||||
steps; its item ID is used directly."""
|
||||
|
||||
|
||||
class ArtistType(enum.StrEnum):
|
||||
"""The decided artist type of a snapshot row."""
|
||||
@@ -147,15 +80,14 @@ class ArtistSnapshot:
|
||||
return asdict(self)
|
||||
|
||||
|
||||
SNAPSHOT_FIELDS: Sequence[str] = tuple(
|
||||
x.name for x in fields(ArtistSnapshot))
|
||||
"""The header columns of the Wikidata artist snapshot CSV file."""
|
||||
|
||||
|
||||
@dataclass
|
||||
class GroupMember:
|
||||
"""One has-part member of a Wikidata group item."""
|
||||
|
||||
__CORPUS_START_YEAR: ClassVar[int] = 2016
|
||||
"""The first year of the corpus window: a member who left a
|
||||
group before it never performed a corpus song."""
|
||||
|
||||
qid: str
|
||||
"""The item ID of the member."""
|
||||
start_years: list[int] = field(default_factory=list)
|
||||
@@ -181,7 +113,7 @@ class GroupMember:
|
||||
if len(self.end_years) == 0:
|
||||
return True
|
||||
last_end: int = max(self.end_years)
|
||||
if last_end >= CORPUS_START_YEAR:
|
||||
if last_end >= self.__CORPUS_START_YEAR:
|
||||
return True
|
||||
if len(self.start_years) == 0:
|
||||
return False
|
||||
@@ -229,20 +161,8 @@ class RetryExhausted(Exception):
|
||||
"""
|
||||
|
||||
|
||||
def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
"""Parse the command-line arguments.
|
||||
|
||||
:param argv: The command-line arguments, or None for
|
||||
``sys.argv``.
|
||||
:return: The parsed arguments.
|
||||
"""
|
||||
parser: argparse.ArgumentParser = argparse.ArgumentParser(
|
||||
description="Fetch the artist metadata from Wikidata"
|
||||
" into the capture layer.")
|
||||
parser.add_argument(
|
||||
"wikidata_csv", type=Path,
|
||||
help="the Wikidata artist snapshot CSV file")
|
||||
return parser.parse_args(argv)
|
||||
NOTE_NOT_FOUND: str = "not found"
|
||||
"""The note marking an artist that could not be resolved."""
|
||||
|
||||
|
||||
class ArtistFetcher:
|
||||
@@ -291,6 +211,71 @@ class ArtistFetcher:
|
||||
in the note.
|
||||
"""
|
||||
|
||||
__API_URL: ClassVar[str] = "https://www.wikidata.org/w/api.php"
|
||||
"""The URL of the Wikidata API endpoint."""
|
||||
__SPARQL_URL: ClassVar[str] \
|
||||
= "https://query.wikidata.org/sparql"
|
||||
"""The URL of the Wikidata Query Service SPARQL endpoint."""
|
||||
__USER_AGENT: ClassVar[str] = (
|
||||
f"pop-fem-audit-tools/{VERSION}"
|
||||
" (https://github.com/imacat/pop-fem-audit;"
|
||||
" mailto:imacat@mail.imacat.idv.tw)")
|
||||
"""The User-Agent header sent on every HTTP request."""
|
||||
__TIMEOUT: ClassVar[float] = 30.0
|
||||
"""The timeout of an API HTTP request, in seconds."""
|
||||
__SPARQL_TIMEOUT: ClassVar[float] = 90.0
|
||||
"""The timeout of a SPARQL HTTP request, in seconds.
|
||||
|
||||
Higher than the API timeout: the WDQS server aborts a slow
|
||||
query at 60 seconds, and a lower client timeout would race
|
||||
that server-side abort and misclassify a slow-but-answerable
|
||||
query as a client-side timeout instead of letting the
|
||||
server's own HTTP error response arrive and enter the retry
|
||||
path."""
|
||||
__SLEEP_SECONDS: ClassVar[float] = 1.0
|
||||
"""The delay between consecutive HTTP requests, in
|
||||
seconds."""
|
||||
__MAX_ATTEMPTS: ClassVar[int] = 5
|
||||
"""The maximum number of attempts on a transient error."""
|
||||
__RETRY_SECONDS: ClassVar[float] = 15.0
|
||||
"""The back-off unit on a transient error, in seconds;
|
||||
multiplied by the attempt number already made."""
|
||||
__RETRY_STATUSES: ClassVar[frozenset[int]] \
|
||||
= frozenset({429, 500, 502, 503})
|
||||
"""The HTTP statuses that are retried with a back-off."""
|
||||
__MAX_STAGE1_TITLES: ClassVar[int] = 3
|
||||
"""The maximum number of charted titles used for the
|
||||
stage-1 song corroboration."""
|
||||
__HUMAN_QID: ClassVar[str] = "Q5"
|
||||
"""The Wikidata item ID of "human"."""
|
||||
__ENSEMBLE_QID: ClassVar[str] = "Q2088357"
|
||||
"""The Wikidata item ID of "musical ensemble"."""
|
||||
__ORIGINAL_CAST_QID: ClassVar[str] = "Q106497009"
|
||||
"""The Wikidata item ID of "original cast"."""
|
||||
__GROUP_KEYWORDS: ClassVar[Sequence[str]] \
|
||||
= ("band", "group", "duo", "trio")
|
||||
"""The label keywords that suggest a musical ensemble,
|
||||
covering labels like "boy band" and "girl group"."""
|
||||
__MIXED_GENDER: ClassVar[str] = "mixed"
|
||||
"""The gender recorded for a group whose members do not
|
||||
share one gender."""
|
||||
__TIME_YEAR_PATTERN: ClassVar[re.Pattern[str]] \
|
||||
= re.compile(r"^[+-]?\d+")
|
||||
"""The leading year of a Wikidata time value."""
|
||||
__PINNED_QIDS: ClassVar[dict[str, str]] = {}
|
||||
"""The last-resort pinned item IDs, keyed by the artist name.
|
||||
|
||||
An entry is for an artist the algorithm documented on
|
||||
``ArtistFetcher`` is structurally unable to resolve, with its
|
||||
justification recorded here. Currently empty: the only pin
|
||||
ever needed, "Pinkfong" (typed as a brand, which the type
|
||||
gate excludes by design), became moot when the store's
|
||||
artist entity behind that credit was identified as Hope
|
||||
Segoine.
|
||||
|
||||
A pinned name skips the candidate retrieval and corroboration
|
||||
steps; its item ID is used directly."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
"""Construct the fetcher."""
|
||||
self.__sent: int = 0
|
||||
@@ -342,8 +327,8 @@ class ArtistFetcher:
|
||||
transient error are exhausted.
|
||||
:raises ValueError: On a JSON decoding error.
|
||||
"""
|
||||
if name in PINNED_QIDS:
|
||||
return PINNED_QIDS[name]
|
||||
if name in self.__PINNED_QIDS:
|
||||
return self.__PINNED_QIDS[name]
|
||||
candidates: list[str] = self.__candidates(name)
|
||||
if len(candidates) == 0:
|
||||
return None
|
||||
@@ -373,11 +358,11 @@ class ArtistFetcher:
|
||||
{{ ?item rdfs:label ?name }}
|
||||
UNION {{ ?item skos:altLabel ?name }}
|
||||
{{
|
||||
?item wdt:P31 wd:{HUMAN_QID}
|
||||
?item wdt:P31 wd:{self.__HUMAN_QID}
|
||||
}} UNION {{
|
||||
?item wdt:P31/wdt:P279* wd:{ENSEMBLE_QID}
|
||||
?item wdt:P31/wdt:P279* wd:{self.__ENSEMBLE_QID}
|
||||
}} UNION {{
|
||||
?item wdt:P31 wd:{ORIGINAL_CAST_QID}
|
||||
?item wdt:P31 wd:{self.__ORIGINAL_CAST_QID}
|
||||
}}
|
||||
}}
|
||||
"""
|
||||
@@ -403,7 +388,7 @@ class ArtistFetcher:
|
||||
transient error are exhausted.
|
||||
:raises ValueError: On a JSON decoding error.
|
||||
"""
|
||||
subset: Sequence[str] = titles[:MAX_STAGE1_TITLES]
|
||||
subset: Sequence[str] = titles[:self.__MAX_STAGE1_TITLES]
|
||||
if len(subset) == 0:
|
||||
return None
|
||||
query: str = f"""
|
||||
@@ -532,7 +517,7 @@ class ArtistFetcher:
|
||||
qid: str
|
||||
for qid in qids:
|
||||
member: MemberClaims = claims.get(qid, MemberClaims())
|
||||
if HUMAN_QID not in member.instance_of_ids:
|
||||
if self.__HUMAN_QID not in member.instance_of_ids:
|
||||
continue
|
||||
if len(member.gender_ids) == 0:
|
||||
return
|
||||
@@ -543,7 +528,7 @@ class ArtistFetcher:
|
||||
[x[1] for x in genders], any_language=True)
|
||||
unique: set[str] = {x[1] for x in genders}
|
||||
snapshot.gender = labels[genders[0][1]] \
|
||||
if len(unique) == 1 else MIXED_GENDER
|
||||
if len(unique) == 1 else self.__MIXED_GENDER
|
||||
basis: str = "gender derived from members: " + "; ".join(
|
||||
f"{x} {labels[y]}" for x, y in genders)
|
||||
snapshot.note = f"{snapshot.note}; {basis}" \
|
||||
@@ -677,7 +662,8 @@ class ArtistFetcher:
|
||||
or not isinstance(value.get("time"), str):
|
||||
continue
|
||||
match: re.Match[str] | None \
|
||||
= TIME_YEAR_PATTERN.match(value["time"])
|
||||
= ArtistFetcher.__TIME_YEAR_PATTERN.match(
|
||||
value["time"])
|
||||
if match is not None:
|
||||
years.append(int(match.group()))
|
||||
return years
|
||||
@@ -806,12 +792,13 @@ class ArtistFetcher:
|
||||
``ArtistType.GROUP`` for a musical ensemble, or the
|
||||
empty string for the human to decide.
|
||||
"""
|
||||
if HUMAN_QID in type_ids:
|
||||
if ArtistFetcher.__HUMAN_QID in type_ids:
|
||||
return ArtistType.SOLO
|
||||
qid: str
|
||||
for qid in type_ids:
|
||||
label: str = labels.get(qid, "").lower()
|
||||
if any(x in label for x in GROUP_KEYWORDS):
|
||||
if any(x in label
|
||||
for x in ArtistFetcher.__GROUP_KEYWORDS):
|
||||
return ArtistType.GROUP
|
||||
return ""
|
||||
|
||||
@@ -827,14 +814,15 @@ class ArtistFetcher:
|
||||
transient error are exhausted.
|
||||
:raises ValueError: On a JSON decoding error.
|
||||
"""
|
||||
url: str = (f"{SPARQL_URL}?"
|
||||
f"{urllib.parse.urlencode({'query': query})}")
|
||||
url: str = (
|
||||
f"{self.__SPARQL_URL}?"
|
||||
f"{urllib.parse.urlencode({'query': query})}")
|
||||
request: urllib.request.Request = urllib.request.Request(
|
||||
url, headers={
|
||||
"User-Agent": USER_AGENT,
|
||||
"User-Agent": self.__USER_AGENT,
|
||||
"Accept": "application/sparql-results+json"})
|
||||
body: bytes = self.__send(
|
||||
request, timeout=SPARQL_TIMEOUT)
|
||||
request, timeout=self.__SPARQL_TIMEOUT)
|
||||
data: Any = json.loads(body)
|
||||
bindings: Any = None
|
||||
if isinstance(data, dict) \
|
||||
@@ -868,13 +856,14 @@ class ArtistFetcher:
|
||||
transient error are exhausted.
|
||||
:raises ValueError: On a JSON decoding error.
|
||||
"""
|
||||
url: str = f"{API_URL}?{urllib.parse.urlencode(params)}"
|
||||
url: str \
|
||||
= f"{self.__API_URL}?{urllib.parse.urlencode(params)}"
|
||||
request: urllib.request.Request = urllib.request.Request(
|
||||
url, headers={"User-Agent": USER_AGENT})
|
||||
url, headers={"User-Agent": self.__USER_AGENT})
|
||||
return json.loads(self.__send(request))
|
||||
|
||||
def __send(self, request: urllib.request.Request,
|
||||
timeout: float = TIMEOUT) -> bytes:
|
||||
timeout: float = __TIMEOUT) -> bytes:
|
||||
"""Send an HTTP request, retrying on a transient error.
|
||||
|
||||
Consecutive requests are separated by a fixed delay. A
|
||||
@@ -892,7 +881,7 @@ class ArtistFetcher:
|
||||
transient error are exhausted.
|
||||
"""
|
||||
if self.__sent > 0:
|
||||
time.sleep(SLEEP_SECONDS)
|
||||
time.sleep(self.__SLEEP_SECONDS)
|
||||
self.__sent += 1
|
||||
attempt: int = 1
|
||||
reason: str | None
|
||||
@@ -905,10 +894,10 @@ class ArtistFetcher:
|
||||
reason = self.__retry_reason(error)
|
||||
if reason is None:
|
||||
raise
|
||||
if attempt >= MAX_ATTEMPTS:
|
||||
if attempt >= self.__MAX_ATTEMPTS:
|
||||
raise RetryExhausted(
|
||||
f"retries exhausted ({reason})") from error
|
||||
time.sleep(RETRY_SECONDS * attempt)
|
||||
time.sleep(self.__RETRY_SECONDS * attempt)
|
||||
attempt += 1
|
||||
|
||||
@staticmethod
|
||||
@@ -923,7 +912,7 @@ class ArtistFetcher:
|
||||
the error is not transient and must not be retried.
|
||||
"""
|
||||
if isinstance(error, urllib.error.HTTPError):
|
||||
if error.code not in RETRY_STATUSES:
|
||||
if error.code not in ArtistFetcher.__RETRY_STATUSES:
|
||||
return None
|
||||
return str(error)
|
||||
if isinstance(error, TimeoutError):
|
||||
@@ -968,92 +957,216 @@ class ArtistFetcher:
|
||||
return uri.rsplit("/", 1)[-1]
|
||||
|
||||
|
||||
def read_snapshot_rows(file: TextIO) -> list[dict[str, str]]:
|
||||
"""Read the current rows of a snapshot CSV file handle.
|
||||
@dataclass(frozen=True)
|
||||
class FetchCounts:
|
||||
"""The outcome counts of one snapshot update run."""
|
||||
|
||||
:param file: The open, seekable snapshot CSV file.
|
||||
:return: The rows, keyed by the column name.
|
||||
:raises OSError: When the file cannot be read.
|
||||
fetched: int
|
||||
"""The number of artists newly resolved."""
|
||||
not_found: int
|
||||
"""The number of artists left unresolved."""
|
||||
errors: int
|
||||
"""The number of artists that ended in an error."""
|
||||
|
||||
|
||||
class ArtistSnapshotUpdater:
|
||||
"""The updater of the Wikidata artist snapshot CSV file.
|
||||
|
||||
Fetches the metadata of every artist of the working store
|
||||
that the snapshot does not resolve yet, appends a row for
|
||||
each to the snapshot as it is fetched, and rewrites the
|
||||
snapshot sorted by artist name with its stale rows dropped.
|
||||
"""
|
||||
file.seek(0)
|
||||
reader: csv.DictReader[str] = csv.DictReader(file)
|
||||
return list(reader)
|
||||
|
||||
__SNAPSHOT_FIELDS: ClassVar[Sequence[str]] = tuple(
|
||||
x.name for x in fields(ArtistSnapshot))
|
||||
"""The header columns of the Wikidata artist snapshot CSV file."""
|
||||
|
||||
def read_artist_titles(session: Session,
|
||||
artist_id: int) -> list[str]:
|
||||
"""Read the charted song titles credited to an artist.
|
||||
def __init__(self, wikidata_csv: Path) -> None:
|
||||
"""Set up the updater.
|
||||
|
||||
:param session: The database session.
|
||||
:param artist_id: The artist ID.
|
||||
:return: The song titles credited to the artist, ordered by
|
||||
the song ID, with the duplicate titles removed.
|
||||
"""
|
||||
titles: Sequence[str] = session.scalars(
|
||||
sa.select(Song.title)
|
||||
.join(SongArtist, SongArtist.song_id == Song.id)
|
||||
.where(SongArtist.artist_id == artist_id)
|
||||
.order_by(Song.id)).all()
|
||||
return list(dict.fromkeys(titles))
|
||||
:param wikidata_csv: The Wikidata artist snapshot CSV
|
||||
file.
|
||||
"""
|
||||
self.__wikidata_csv: Path = wikidata_csv
|
||||
"""The Wikidata artist snapshot CSV file."""
|
||||
|
||||
def run(self) -> FetchCounts:
|
||||
"""Fetch every unresolved artist and update the snapshot.
|
||||
|
||||
def ensure_snapshot_header(file: TextIO) -> None:
|
||||
"""Write the snapshot CSV header row if the file is empty.
|
||||
:return: The counts of the run.
|
||||
:raises OSError: When the snapshot file, or its parent
|
||||
directory, cannot be read or written.
|
||||
:raises sqlalchemy.exc.SQLAlchemyError: When the working
|
||||
store cannot be read.
|
||||
"""
|
||||
session: Session = ds.get_db()
|
||||
try:
|
||||
return self.__run(session)
|
||||
finally:
|
||||
session.close()
|
||||
|
||||
:param file: The open, seekable snapshot CSV file.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
"""
|
||||
file.seek(0, os.SEEK_END)
|
||||
if file.tell() == 0:
|
||||
csv.writer(file).writerow(SNAPSHOT_FIELDS)
|
||||
def __run(self, session: Session) -> FetchCounts:
|
||||
"""Run the fetch loop with an open database session.
|
||||
|
||||
:param session: The database session.
|
||||
:return: The counts of the run.
|
||||
:raises OSError: When the snapshot file, or its parent
|
||||
directory, cannot be read or written.
|
||||
"""
|
||||
fetcher: ArtistFetcher = ArtistFetcher()
|
||||
fetched: int = 0
|
||||
not_found: int = 0
|
||||
errors: int = 0
|
||||
self.__wikidata_csv.parent.mkdir(
|
||||
parents=True, exist_ok=True)
|
||||
with open(self.__wikidata_csv, "a+", encoding="utf-8",
|
||||
newline="") as csv_file:
|
||||
done: set[str] = {
|
||||
x["name"] for x in
|
||||
self.__read_snapshot_rows(csv_file)
|
||||
if x["gender"] != ""}
|
||||
self.__ensure_snapshot_header(csv_file)
|
||||
names: set[str] = set()
|
||||
artist: Artist
|
||||
for artist in session.scalars(
|
||||
sa.select(Artist).order_by(Artist.id)):
|
||||
names.add(artist.name)
|
||||
if artist.name in done:
|
||||
continue
|
||||
titles: list[str] = self.__read_artist_titles(
|
||||
session, artist.id)
|
||||
snapshot: ArtistSnapshot = fetcher.fetch(
|
||||
artist.name, titles)
|
||||
self.__append_row(csv_file, snapshot)
|
||||
status: str = snapshot.qid
|
||||
if snapshot.note == NOTE_NOT_FOUND:
|
||||
not_found += 1
|
||||
status = NOTE_NOT_FOUND
|
||||
elif snapshot.note.startswith("error: "):
|
||||
errors += 1
|
||||
status = snapshot.note
|
||||
else:
|
||||
fetched += 1
|
||||
print(f"artist \"{artist.name}\": {status}",
|
||||
file=sys.stderr)
|
||||
self.__write_snapshot(csv_file, names)
|
||||
return FetchCounts(
|
||||
fetched=fetched, not_found=not_found, errors=errors)
|
||||
|
||||
@staticmethod
|
||||
def __read_snapshot_rows(file: TextIO) \
|
||||
-> list[dict[str, str]]:
|
||||
"""Read the current rows of a snapshot CSV file handle.
|
||||
|
||||
:param file: The open, seekable snapshot CSV file.
|
||||
:return: The rows, keyed by the column name.
|
||||
:raises OSError: When the file cannot be read.
|
||||
"""
|
||||
file.seek(0)
|
||||
reader: csv.DictReader[str] = csv.DictReader(file)
|
||||
return list(reader)
|
||||
|
||||
@staticmethod
|
||||
def __read_artist_titles(session: Session,
|
||||
artist_id: int) -> list[str]:
|
||||
"""Read the charted song titles credited to an artist.
|
||||
|
||||
:param session: The database session.
|
||||
:param artist_id: The artist ID.
|
||||
:return: The song titles credited to the artist, ordered
|
||||
by the song ID, with the duplicate titles removed.
|
||||
"""
|
||||
titles: Sequence[str] = session.scalars(
|
||||
sa.select(Song.title)
|
||||
.join(SongArtist, SongArtist.song_id == Song.id)
|
||||
.where(SongArtist.artist_id == artist_id)
|
||||
.order_by(Song.id)).all()
|
||||
return list(dict.fromkeys(titles))
|
||||
|
||||
@staticmethod
|
||||
def __ensure_snapshot_header(file: TextIO) -> None:
|
||||
"""Write the snapshot CSV header row if the file is
|
||||
empty.
|
||||
|
||||
:param file: The open, seekable snapshot CSV file.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
"""
|
||||
file.seek(0, os.SEEK_END)
|
||||
if file.tell() == 0:
|
||||
csv.writer(file).writerow(
|
||||
ArtistSnapshotUpdater.__SNAPSHOT_FIELDS)
|
||||
file.flush()
|
||||
|
||||
@staticmethod
|
||||
def __append_row(file: TextIO,
|
||||
snapshot: ArtistSnapshot) -> None:
|
||||
"""Append a snapshot row to a snapshot CSV file handle.
|
||||
|
||||
:param file: The open snapshot CSV file, opened for
|
||||
append.
|
||||
:param snapshot: The snapshot of an artist.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
"""
|
||||
csv.DictWriter(
|
||||
file, ArtistSnapshotUpdater.__SNAPSHOT_FIELDS).writerow(
|
||||
snapshot.to_row())
|
||||
file.flush()
|
||||
|
||||
@staticmethod
|
||||
def __write_snapshot(file: TextIO, names: Container[str]) \
|
||||
-> None:
|
||||
"""Rewrite a snapshot CSV file handle sorted by artist
|
||||
name.
|
||||
|
||||
def append_row(file: TextIO, snapshot: ArtistSnapshot) -> None:
|
||||
"""Append a snapshot row to a snapshot CSV file handle.
|
||||
The rows are ordered by the case-folded artist name,
|
||||
matching the convention of the derived ``artists.csv``.
|
||||
An artist keeps one row only, the last one of the file,
|
||||
so that a re-fetched artist replaces its earlier row. A
|
||||
row whose name is not an artist of the store is dropped
|
||||
and reported on the standard error.
|
||||
|
||||
:param file: The open snapshot CSV file, opened for append.
|
||||
:param snapshot: The snapshot of an artist.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
:param file: The open, seekable snapshot CSV file.
|
||||
:param names: The artist names of the working store.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be read or written.
|
||||
"""
|
||||
kept: dict[str, dict[str, str]] = {}
|
||||
row: dict[str, str]
|
||||
for row in ArtistSnapshotUpdater.__read_snapshot_rows(
|
||||
file):
|
||||
if row["name"] not in names:
|
||||
print(f"dropped stale row \"{row['name']}\":"
|
||||
" no such artist in the store",
|
||||
file=sys.stderr)
|
||||
continue
|
||||
kept[row["name"]] = row
|
||||
ordered: list[dict[str, str]] = sorted(
|
||||
kept.values(), key=lambda x: x["name"].casefold())
|
||||
file.seek(0)
|
||||
file.truncate()
|
||||
writer: csv.DictWriter[str] = csv.DictWriter(
|
||||
file, ArtistSnapshotUpdater.__SNAPSHOT_FIELDS)
|
||||
writer.writeheader()
|
||||
writer.writerows(ordered)
|
||||
|
||||
|
||||
def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
"""Parse the command-line arguments.
|
||||
|
||||
:param argv: The command-line arguments, or None for
|
||||
``sys.argv``.
|
||||
:return: The parsed arguments.
|
||||
"""
|
||||
csv.DictWriter(file, SNAPSHOT_FIELDS).writerow(
|
||||
snapshot.to_row())
|
||||
file.flush()
|
||||
|
||||
|
||||
def write_snapshot(file: TextIO, names: Container[str]) -> None:
|
||||
"""Rewrite a snapshot CSV file handle sorted by artist name.
|
||||
|
||||
The rows are ordered by the case-folded artist name, matching
|
||||
the convention of the derived ``artists.csv``. An artist
|
||||
keeps one row only, the last one of the file, so that a
|
||||
re-fetched artist replaces its earlier row. A row whose name
|
||||
is not an artist of the store is dropped and reported on the
|
||||
standard error.
|
||||
|
||||
:param file: The open, seekable snapshot CSV file.
|
||||
:param names: The artist names of the working store.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be read or written.
|
||||
"""
|
||||
kept: dict[str, dict[str, str]] = {}
|
||||
row: dict[str, str]
|
||||
for row in read_snapshot_rows(file):
|
||||
if row["name"] not in names:
|
||||
print(f"dropped stale row \"{row['name']}\":"
|
||||
" no such artist in the store", file=sys.stderr)
|
||||
continue
|
||||
kept[row["name"]] = row
|
||||
ordered: list[dict[str, str]] = sorted(
|
||||
kept.values(), key=lambda x: x["name"].casefold())
|
||||
file.seek(0)
|
||||
file.truncate()
|
||||
writer: csv.DictWriter[str] = csv.DictWriter(
|
||||
file, SNAPSHOT_FIELDS)
|
||||
writer.writeheader()
|
||||
writer.writerows(ordered)
|
||||
parser: argparse.ArgumentParser = argparse.ArgumentParser(
|
||||
description="Fetch the artist metadata from Wikidata"
|
||||
" into the capture layer.")
|
||||
parser.add_argument(
|
||||
"wikidata_csv", type=Path,
|
||||
help="the Wikidata artist snapshot CSV file")
|
||||
return parser.parse_args(argv)
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
@@ -1066,52 +1179,16 @@ def main(argv: list[str] | None = None) -> int:
|
||||
"""
|
||||
started: float = time.monotonic()
|
||||
args: argparse.Namespace = parse_args(argv)
|
||||
fetcher: ArtistFetcher = ArtistFetcher()
|
||||
fetched: int = 0
|
||||
not_found: int = 0
|
||||
errors: int = 0
|
||||
session: Session = ds.get_db()
|
||||
try:
|
||||
args.wikidata_csv.parent.mkdir(
|
||||
parents=True, exist_ok=True)
|
||||
with open(args.wikidata_csv, "a+", encoding="utf-8",
|
||||
newline="") as csv_file:
|
||||
done: set[str] = {x["name"] for x in
|
||||
read_snapshot_rows(csv_file)
|
||||
if x["gender"] != ""}
|
||||
ensure_snapshot_header(csv_file)
|
||||
names: set[str] = set()
|
||||
artist: Artist
|
||||
for artist in session.scalars(
|
||||
sa.select(Artist).order_by(Artist.id)):
|
||||
names.add(artist.name)
|
||||
if artist.name in done:
|
||||
continue
|
||||
titles: list[str] = read_artist_titles(
|
||||
session, artist.id)
|
||||
snapshot: ArtistSnapshot = fetcher.fetch(
|
||||
artist.name, titles)
|
||||
append_row(csv_file, snapshot)
|
||||
status: str = snapshot.qid
|
||||
if snapshot.note == NOTE_NOT_FOUND:
|
||||
not_found += 1
|
||||
status = "not found"
|
||||
elif snapshot.note.startswith("error: "):
|
||||
errors += 1
|
||||
status = snapshot.note
|
||||
else:
|
||||
fetched += 1
|
||||
print(f"artist \"{artist.name}\": {status}",
|
||||
file=sys.stderr)
|
||||
write_snapshot(csv_file, names)
|
||||
counts: FetchCounts \
|
||||
= ArtistSnapshotUpdater(args.wikidata_csv).run()
|
||||
except (OSError, sa.exc.SQLAlchemyError) as error:
|
||||
print(f"error: {error}", file=sys.stderr)
|
||||
return 1
|
||||
finally:
|
||||
session.close()
|
||||
attempted: int = fetched + not_found + errors
|
||||
attempted: int = counts.fetched + counts.not_found \
|
||||
+ counts.errors
|
||||
elapsed: str = format_duration(time.monotonic() - started)
|
||||
print(f"Done. Resolved {fetched}/{attempted} artists."
|
||||
f" {elapsed} elapsed.",
|
||||
print(f"Done. Resolved {counts.fetched}/{attempted}"
|
||||
f" artists. {elapsed} elapsed.",
|
||||
file=sys.stderr)
|
||||
return 0
|
||||
|
||||
@@ -30,8 +30,9 @@ import time
|
||||
import urllib.parse
|
||||
import urllib.request
|
||||
from collections.abc import Sequence
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
from typing import Any, ClassVar
|
||||
|
||||
import sqlalchemy as sa
|
||||
from sqlalchemy.orm import Session
|
||||
@@ -45,60 +46,6 @@ from ..models import (
|
||||
)
|
||||
from ..utils import format_duration
|
||||
|
||||
PROVENANCE_FIELDS: Sequence[str] = (
|
||||
"song_id", "source", "method", "acquired_at", "note")
|
||||
"""The header columns of the lyrics provenance CSV file."""
|
||||
USER_AGENT: str = ("pop-fem-audit-tools"
|
||||
" (https://github.com/imacat/pop-fem-audit)")
|
||||
"""The User-Agent header sent on every HTTP request."""
|
||||
TIMEOUT: float = 30.0
|
||||
"""The timeout of an HTTP request, in seconds."""
|
||||
SLEEP_SECONDS: float = 1.0
|
||||
"""The delay between consecutive HTTP requests, in seconds."""
|
||||
|
||||
|
||||
def __build_normalization() -> dict[int, str | None]:
|
||||
"""Build the lyrics normalization translation table.
|
||||
|
||||
:return: The codepoint-to-replacement mapping, a replacement
|
||||
of None meaning removal.
|
||||
"""
|
||||
table: dict[int, str | None] = {}
|
||||
codepoint: int
|
||||
for codepoint in range(0x80, 0xa0):
|
||||
try:
|
||||
table[codepoint] = bytes([codepoint]).decode("cp1252")
|
||||
except UnicodeDecodeError:
|
||||
table[codepoint] = None
|
||||
table[0x0435] = "e"
|
||||
table[0x03cc] = "ó"
|
||||
for codepoint in (0x2005, 0x205f, 0x200a):
|
||||
table[codepoint] = " "
|
||||
for codepoint in (0x200b, 0x200c, 0x200d, 0xfeff):
|
||||
table[codepoint] = None
|
||||
return table
|
||||
|
||||
|
||||
NORMALIZATION: dict[int, str | None] = __build_normalization()
|
||||
"""The codepoint-to-replacement mapping applied to fetched
|
||||
lyrics: cp1252-mojibake restoration for U+0080-U+009F (with the
|
||||
five byte values undefined in cp1252 removed), homoglyph
|
||||
restoration for the Cyrillic "e" and the Greek "o" with tonos,
|
||||
ASCII-space restoration for exotic space variants, and removal
|
||||
of zero-width characters. A replacement of None removes the
|
||||
codepoint."""
|
||||
|
||||
|
||||
def normalize_lyrics(text: str) -> str:
|
||||
"""Restore or remove watermark and mojibake characters.
|
||||
|
||||
:param text: The lyrics text as fetched from an API.
|
||||
:return: The text with the codepoints in
|
||||
:data:`NORMALIZATION` replaced or removed; every other
|
||||
character is unchanged.
|
||||
"""
|
||||
return text.translate(NORMALIZATION)
|
||||
|
||||
|
||||
def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
"""Parse the command-line arguments.
|
||||
@@ -122,6 +69,16 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
class LyricsFetcher:
|
||||
"""A fetcher of song lyrics from the public lyrics APIs."""
|
||||
|
||||
__USER_AGENT: ClassVar[str] = (
|
||||
"pop-fem-audit-tools"
|
||||
" (https://github.com/imacat/pop-fem-audit)")
|
||||
"""The User-Agent header sent on every HTTP request."""
|
||||
__TIMEOUT: ClassVar[float] = 30.0
|
||||
"""The timeout of an HTTP request, in seconds."""
|
||||
__SLEEP_SECONDS: ClassVar[float] = 1.0
|
||||
"""The delay between consecutive HTTP requests, in
|
||||
seconds."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
"""Construct the fetcher."""
|
||||
self.__sent: int = 0
|
||||
@@ -193,78 +150,216 @@ class LyricsFetcher:
|
||||
network, or decoding error.
|
||||
"""
|
||||
if self.__sent > 0:
|
||||
time.sleep(SLEEP_SECONDS)
|
||||
time.sleep(self.__SLEEP_SECONDS)
|
||||
self.__sent += 1
|
||||
request: urllib.request.Request = urllib.request.Request(
|
||||
url, headers={"User-Agent": USER_AGENT})
|
||||
url, headers={"User-Agent": self.__USER_AGENT})
|
||||
try:
|
||||
with urllib.request.urlopen(
|
||||
request, timeout=TIMEOUT) as response:
|
||||
request, timeout=self.__TIMEOUT) as response:
|
||||
return json.load(response)
|
||||
except (OSError, ValueError):
|
||||
return None
|
||||
|
||||
|
||||
def query_artist(session: Session, song_id: int) -> str:
|
||||
"""Find the artist name to query the APIs with.
|
||||
@dataclass(frozen=True)
|
||||
class LyricsFetchCounts:
|
||||
"""The outcome of one run of fetching the missing lyrics."""
|
||||
|
||||
:param session: The database session.
|
||||
:param song_id: The song ID.
|
||||
:return: The name of the primary-role artist with the lowest
|
||||
position.
|
||||
"""
|
||||
name: str | None = session.scalar(
|
||||
sa.select(Artist.name)
|
||||
.join(SongArtist, SongArtist.artist_id == Artist.id)
|
||||
.where(SongArtist.song_id == song_id,
|
||||
SongArtist.role == Role.PRIMARY)
|
||||
.order_by(SongArtist.position)
|
||||
.limit(1))
|
||||
assert name is not None
|
||||
return name
|
||||
fetched: int
|
||||
"""The number of songs newly fetched."""
|
||||
missed: int
|
||||
"""The number of songs every API missed."""
|
||||
|
||||
|
||||
def save_lyrics(lyrics_dir: Path, song_id: int,
|
||||
lyrics: str) -> None:
|
||||
"""Write the lyrics of a song into the cache directory.
|
||||
class LyricsFetchRunner:
|
||||
"""The orchestrator of one run of fetching missing lyrics."""
|
||||
|
||||
The cache directory is created when missing.
|
||||
__PROVENANCE_FIELDS: ClassVar[Sequence[str]] = (
|
||||
"song_id", "source", "method", "acquired_at", "note")
|
||||
"""The header columns of the lyrics provenance CSV file."""
|
||||
|
||||
The lyrics text is normalized with :func:`normalize_lyrics`
|
||||
before being written.
|
||||
@staticmethod
|
||||
def __build_normalization() -> dict[int, str | None]:
|
||||
"""Build the lyrics normalization translation table.
|
||||
|
||||
:param lyrics_dir: The lyrics cache directory.
|
||||
:param song_id: The song ID.
|
||||
:param lyrics: The lyrics text.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
"""
|
||||
lyrics_dir.mkdir(parents=True, exist_ok=True)
|
||||
(lyrics_dir / f"{song_id}.txt").write_text(
|
||||
normalize_lyrics(lyrics), encoding="utf-8")
|
||||
:return: The codepoint-to-replacement mapping, a
|
||||
replacement of None meaning removal.
|
||||
"""
|
||||
table: dict[int, str | None] = {}
|
||||
codepoint: int
|
||||
for codepoint in range(0x80, 0xa0):
|
||||
try:
|
||||
table[codepoint] = bytes(
|
||||
[codepoint]).decode("cp1252")
|
||||
except UnicodeDecodeError:
|
||||
table[codepoint] = None
|
||||
table[0x0435] = "e"
|
||||
table[0x03cc] = "ó"
|
||||
for codepoint in (0x2005, 0x205f, 0x200a):
|
||||
table[codepoint] = " "
|
||||
for codepoint in (0x200b, 0x200c, 0x200d, 0xfeff):
|
||||
table[codepoint] = None
|
||||
return table
|
||||
|
||||
__NORMALIZATION: ClassVar[dict[int, str | None]] \
|
||||
= __build_normalization()
|
||||
"""The codepoint-to-replacement mapping applied to fetched
|
||||
lyrics: cp1252-mojibake restoration for U+0080-U+009F (with
|
||||
the five byte values undefined in cp1252 removed), homoglyph
|
||||
restoration for the Cyrillic "e" and the Greek "o" with
|
||||
tonos, ASCII-space restoration for exotic space variants, and
|
||||
removal of zero-width characters. A replacement of None
|
||||
removes the codepoint."""
|
||||
|
||||
def append_provenance(path: Path, song_id: int,
|
||||
source: str) -> None:
|
||||
"""Append a provenance row for a fetched lyrics file.
|
||||
def __init__(self, lyrics_dir: Path,
|
||||
provenance_csv: Path) -> None:
|
||||
"""Set up the fetch run.
|
||||
|
||||
The CSV file is created with the header row when missing.
|
||||
:param lyrics_dir: The lyrics cache directory.
|
||||
:param provenance_csv: The lyrics provenance CSV file.
|
||||
"""
|
||||
self.__lyrics_dir: Path = lyrics_dir
|
||||
"""The lyrics cache directory."""
|
||||
self.__provenance_csv: Path = provenance_csv
|
||||
"""The lyrics provenance CSV file."""
|
||||
self.__fetcher: LyricsFetcher = LyricsFetcher()
|
||||
"""The fetcher of the public lyrics APIs."""
|
||||
|
||||
:param path: The lyrics provenance CSV file.
|
||||
:param song_id: The song ID.
|
||||
:param source: The source name of the fetched lyrics.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
"""
|
||||
is_new: bool = not path.exists()
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
with open(path, "a", encoding="utf-8",
|
||||
newline="") as file:
|
||||
writer: Any = csv.writer(file)
|
||||
if is_new:
|
||||
writer.writerow(PROVENANCE_FIELDS)
|
||||
writer.writerow([song_id, source, "api-fetch",
|
||||
datetime.date.today().isoformat(), ""])
|
||||
def run(self) -> LyricsFetchCounts:
|
||||
"""Fetch the missing lyrics of every song in the store.
|
||||
|
||||
Every song fetched or missed is reported on the standard
|
||||
error as an observable side effect.
|
||||
|
||||
:return: The number of songs fetched and missed.
|
||||
:raises OSError: When a cache file or the provenance CSV
|
||||
cannot be written.
|
||||
:raises sqlalchemy.exc.SQLAlchemyError: On a database
|
||||
error.
|
||||
"""
|
||||
fetched: int = 0
|
||||
missed: int = 0
|
||||
session: Session = ds.get_db()
|
||||
try:
|
||||
song: Song
|
||||
for song in session.scalars(
|
||||
sa.select(Song).order_by(Song.id)):
|
||||
if (self.__lyrics_dir
|
||||
/ f"{song.id}.txt").exists():
|
||||
continue
|
||||
if self.__fetch_one(session, song):
|
||||
fetched += 1
|
||||
else:
|
||||
missed += 1
|
||||
finally:
|
||||
session.close()
|
||||
return LyricsFetchCounts(fetched=fetched, missed=missed)
|
||||
|
||||
def __fetch_one(self, session: Session, song: Song) -> bool:
|
||||
"""Fetch and save the lyrics of one song.
|
||||
|
||||
The song is queried by its primary-role artist name; when
|
||||
every API misses and the song's full artist credit
|
||||
differs from that name, the same APIs are queried again
|
||||
with the artist credit.
|
||||
|
||||
:param session: The database session.
|
||||
:param song: The song to fetch.
|
||||
:return: True when a lyrics text was fetched and saved,
|
||||
False when every API missed on both queries.
|
||||
:raises OSError: When the cache file or the provenance
|
||||
CSV cannot be written.
|
||||
"""
|
||||
artist: str = self.__query_artist(session, song.id)
|
||||
result: tuple[str, str] | None = self.__fetcher.fetch(
|
||||
artist, song.title)
|
||||
if result is None and song.artist_credit != artist:
|
||||
result = self.__fetcher.fetch(
|
||||
song.artist_credit, song.title)
|
||||
if result is None:
|
||||
print(f"song {song.id} \"{song.title}\": miss",
|
||||
file=sys.stderr)
|
||||
return False
|
||||
lyrics: str
|
||||
source: str
|
||||
lyrics, source = result
|
||||
self.__save_lyrics(song.id, lyrics)
|
||||
self.__append_provenance(song.id, source)
|
||||
print(f"song {song.id} \"{song.title}\": {source}",
|
||||
file=sys.stderr)
|
||||
return True
|
||||
|
||||
@staticmethod
|
||||
def __query_artist(session: Session, song_id: int) -> str:
|
||||
"""Find the artist name to query the APIs with.
|
||||
|
||||
:param session: The database session.
|
||||
:param song_id: The song ID.
|
||||
:return: The name of the primary-role artist with the
|
||||
lowest position.
|
||||
"""
|
||||
name: str | None = session.scalar(
|
||||
sa.select(Artist.name)
|
||||
.join(SongArtist, SongArtist.artist_id == Artist.id)
|
||||
.where(SongArtist.song_id == song_id,
|
||||
SongArtist.role == Role.PRIMARY)
|
||||
.order_by(SongArtist.position)
|
||||
.limit(1))
|
||||
assert name is not None
|
||||
return name
|
||||
|
||||
def __save_lyrics(self, song_id: int, lyrics: str) -> None:
|
||||
"""Write the lyrics of a song into the cache directory.
|
||||
|
||||
The cache directory is created when missing.
|
||||
|
||||
The lyrics text is normalized with
|
||||
:meth:`normalize_lyrics` before being written.
|
||||
|
||||
:param song_id: The song ID.
|
||||
:param lyrics: The lyrics text.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
"""
|
||||
self.__lyrics_dir.mkdir(parents=True, exist_ok=True)
|
||||
(self.__lyrics_dir / f"{song_id}.txt").write_text(
|
||||
self.normalize_lyrics(lyrics), encoding="utf-8")
|
||||
|
||||
def __append_provenance(self, song_id: int,
|
||||
source: str) -> None:
|
||||
"""Append a provenance row for a fetched lyrics file.
|
||||
|
||||
The CSV file is created with the header row when
|
||||
missing.
|
||||
|
||||
:param song_id: The song ID.
|
||||
:param source: The source name of the fetched lyrics.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
"""
|
||||
is_new: bool = not self.__provenance_csv.exists()
|
||||
self.__provenance_csv.parent.mkdir(
|
||||
parents=True, exist_ok=True)
|
||||
with open(self.__provenance_csv, "a", encoding="utf-8",
|
||||
newline="") as file:
|
||||
writer: Any = csv.writer(file)
|
||||
if is_new:
|
||||
writer.writerow(self.__PROVENANCE_FIELDS)
|
||||
writer.writerow(
|
||||
[song_id, source, "api-fetch",
|
||||
datetime.date.today().isoformat(), ""])
|
||||
|
||||
@classmethod
|
||||
def normalize_lyrics(cls, text: str) -> str:
|
||||
"""Restore or remove watermark and mojibake characters.
|
||||
|
||||
:param text: The lyrics text as fetched from an API.
|
||||
:return: The text with the codepoints of the
|
||||
normalization table replaced or removed; every other
|
||||
character is unchanged.
|
||||
"""
|
||||
return text.translate(cls.__NORMALIZATION)
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
@@ -277,44 +372,15 @@ def main(argv: list[str] | None = None) -> int:
|
||||
"""
|
||||
started: float = time.monotonic()
|
||||
args: argparse.Namespace = parse_args(argv)
|
||||
fetcher: LyricsFetcher = LyricsFetcher()
|
||||
fetched: int = 0
|
||||
missed: int = 0
|
||||
session: Session = ds.get_db()
|
||||
try:
|
||||
song: Song
|
||||
for song in session.scalars(
|
||||
sa.select(Song).order_by(Song.id)):
|
||||
if (args.lyrics_dir / f"{song.id}.txt").exists():
|
||||
continue
|
||||
artist: str = query_artist(session, song.id)
|
||||
result: tuple[str, str] | None = fetcher.fetch(
|
||||
artist, song.title)
|
||||
if result is None and song.artist_credit != artist:
|
||||
result = fetcher.fetch(
|
||||
song.artist_credit, song.title)
|
||||
if result is None:
|
||||
missed += 1
|
||||
print(f"song {song.id} \"{song.title}\": miss",
|
||||
file=sys.stderr)
|
||||
continue
|
||||
lyrics: str
|
||||
source: str
|
||||
lyrics, source = result
|
||||
save_lyrics(args.lyrics_dir, song.id, lyrics)
|
||||
append_provenance(args.provenance_csv, song.id,
|
||||
source)
|
||||
fetched += 1
|
||||
print(f"song {song.id} \"{song.title}\": {source}",
|
||||
file=sys.stderr)
|
||||
counts: LyricsFetchCounts = LyricsFetchRunner(
|
||||
args.lyrics_dir, args.provenance_csv).run()
|
||||
except (OSError, sa.exc.SQLAlchemyError) as error:
|
||||
print(f"error: {error}", file=sys.stderr)
|
||||
return 1
|
||||
finally:
|
||||
session.close()
|
||||
attempted: int = fetched + missed
|
||||
attempted: int = counts.fetched + counts.missed
|
||||
elapsed: str = format_duration(time.monotonic() - started)
|
||||
print(f"Done. Fetched lyrics for {fetched}/{attempted}"
|
||||
f" songs. {elapsed} elapsed.",
|
||||
print(f"Done. Fetched lyrics for {counts.fetched}/"
|
||||
f"{attempted} songs. {elapsed} elapsed.",
|
||||
file=sys.stderr)
|
||||
return 0
|
||||
|
||||
@@ -28,27 +28,13 @@ import time
|
||||
from dataclasses import asdict, dataclass
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
from typing import Any, Self
|
||||
from typing import Any, ClassVar, Self
|
||||
|
||||
import anthropic
|
||||
|
||||
from ..config import get_settings
|
||||
from ..utils import format_duration
|
||||
|
||||
# claude-fable-5 accepts neither "temperature" nor "thinking";
|
||||
# a model's entry holds exactly the extra request parameters it
|
||||
# accepts.
|
||||
MODELS: dict[str, dict[str, Any]] = {
|
||||
"claude-sonnet-4-6": {
|
||||
"temperature": 0.0,
|
||||
"thinking": {"type": "disabled"},
|
||||
},
|
||||
"claude-fable-5": {},
|
||||
}
|
||||
DEFAULT_MODEL: str = "claude-sonnet-4-6"
|
||||
SCRIPT_VERSION: str = "run_llm.py 3.1.0"
|
||||
POLL_INTERVAL_SECONDS: float = 60.0
|
||||
|
||||
|
||||
class InputFormatError(Exception):
|
||||
"""An error in the JSONL input file."""
|
||||
@@ -136,13 +122,23 @@ class BatchResult:
|
||||
if x.type == "text")
|
||||
return cls(id=entry.custom_id, text=text,
|
||||
stop_reason=message.stop_reason,
|
||||
usage=usage_to_dict(message.usage))
|
||||
usage=cls.__usage_to_dict(message.usage))
|
||||
case "errored":
|
||||
return cls(id=entry.custom_id,
|
||||
error=result.error.error.type)
|
||||
case other:
|
||||
return cls(id=entry.custom_id, error=str(other))
|
||||
|
||||
@staticmethod
|
||||
def __usage_to_dict(usage: Any) -> dict[str, Any]:
|
||||
"""Convert a usage object to a plain dictionary.
|
||||
|
||||
:param usage: The usage object of a message.
|
||||
:return: The usage as a dictionary, without null entries.
|
||||
"""
|
||||
return {k: v for k, v in usage.model_dump().items()
|
||||
if v is not None}
|
||||
|
||||
def to_record(self) -> dict[str, Any]:
|
||||
"""Return this result as an archive JSONL record.
|
||||
|
||||
@@ -169,10 +165,409 @@ class BatchInfo:
|
||||
is still processing."""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ExecutionOutcome:
|
||||
"""The outcome of submitting and awaiting one batch."""
|
||||
|
||||
batch: BatchInfo
|
||||
"""The submitted batch's bookkeeping."""
|
||||
results: Results
|
||||
"""The batch's results, keyed by item ID."""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class RunOutcome:
|
||||
"""The outcome of one LLM definition file run."""
|
||||
|
||||
item_count: int
|
||||
"""The number of loaded input items."""
|
||||
dry_run: bool
|
||||
"""Whether this was a dry run."""
|
||||
dry_run_request: dict[str, Any] | None
|
||||
"""The first item's preview request, for a dry run; None for
|
||||
an actual run."""
|
||||
failed: list[str]
|
||||
"""The failed item IDs, in item order; always empty for a dry
|
||||
run."""
|
||||
|
||||
|
||||
class LLMRunner:
|
||||
"""The orchestrator of one LLM definition file run."""
|
||||
|
||||
# claude-fable-5 accepts neither "temperature" nor "thinking";
|
||||
# a model's entry holds exactly the extra request parameters
|
||||
# it accepts.
|
||||
MODELS: ClassVar[dict[str, dict[str, Any]]] = {
|
||||
"claude-sonnet-4-6": {
|
||||
"temperature": 0.0,
|
||||
"thinking": {"type": "disabled"},
|
||||
},
|
||||
"claude-fable-5": {},
|
||||
}
|
||||
"""The supported model IDs and their extra request
|
||||
parameters."""
|
||||
DEFAULT_MODEL: ClassVar[str] = "claude-sonnet-4-6"
|
||||
"""The default model ID."""
|
||||
__SCRIPT_VERSION: ClassVar[str] = "run_llm.py 3.1.0"
|
||||
"""The script version recorded into the archive metadata."""
|
||||
__POLL_INTERVAL_SECONDS: ClassVar[float] = 60.0
|
||||
"""The interval between batch status polls."""
|
||||
|
||||
def __init__(self, prompt: Path, input_path: Path,
|
||||
archive_dir: Path, model: str, max_tokens: int,
|
||||
dry_run: bool, replace: bool) -> None:
|
||||
"""Set up the run of one LLM definition file.
|
||||
|
||||
:param prompt: The prompt definition file, used as the
|
||||
system prompt.
|
||||
:param input_path: The JSONL input file with "id" and
|
||||
"content".
|
||||
:param archive_dir: The destination archive directory.
|
||||
:param model: The model ID, a key of :attr:`MODELS`.
|
||||
:param max_tokens: The maximum output tokens per request.
|
||||
:param dry_run: Whether to validate and archive without
|
||||
calling the API.
|
||||
:param replace: Whether to replace an already existing
|
||||
archive directory.
|
||||
"""
|
||||
self.__prompt: Path = prompt
|
||||
"""The prompt definition file."""
|
||||
self.__input: Path = input_path
|
||||
"""The JSONL input file."""
|
||||
self.__archive_dir: Path = archive_dir
|
||||
"""The destination archive directory."""
|
||||
self.__model: str = model
|
||||
"""The model ID."""
|
||||
self.__max_tokens: int = max_tokens
|
||||
"""The maximum output tokens per request."""
|
||||
self.__dry_run: bool = dry_run
|
||||
"""Whether to validate and archive without calling the
|
||||
API."""
|
||||
self.__replace: bool = replace
|
||||
"""Whether to replace an already existing archive
|
||||
directory."""
|
||||
|
||||
def run(self) -> RunOutcome:
|
||||
"""Load the input, archive the prompt, and run the batch.
|
||||
|
||||
Always writes ``prompt.md`` and ``meta.json`` into the
|
||||
archive directory. A dry run stops there, previewing the
|
||||
first item's request; an actual run also submits the
|
||||
batch, awaits it, and writes ``output.jsonl``.
|
||||
|
||||
:return: The outcome of the run.
|
||||
:raises InputFormatError: When the input file is
|
||||
malformed.
|
||||
:raises OSError: When the input or prompt file cannot be
|
||||
read, the archive directory already exists without
|
||||
``replace``, or an output file cannot be written.
|
||||
"""
|
||||
items: list[InputItem] = self.__load_items()
|
||||
prompt_text: str = self.__prompt.read_text(encoding="utf-8")
|
||||
archive_dir: Path = self.__create_archive_dir()
|
||||
(archive_dir / "prompt.md").write_bytes(
|
||||
self.__prompt.read_bytes())
|
||||
meta: dict[str, Any] = self.__build_meta(items)
|
||||
meta_path: Path = archive_dir / "meta.json"
|
||||
if self.__dry_run:
|
||||
self.__write_json(meta_path, meta)
|
||||
request: dict[str, Any] = self.__build_request(
|
||||
items[0], prompt_text)
|
||||
return RunOutcome(
|
||||
item_count=len(items), dry_run=True,
|
||||
dry_run_request=request, failed=[])
|
||||
client: anthropic.Anthropic = anthropic.Anthropic(
|
||||
api_key=get_settings().ANTHROPIC_API_KEY)
|
||||
outcome: ExecutionOutcome = self.__execute_run(
|
||||
client, items, prompt_text)
|
||||
item_ids: list[str] = [x.id for x in items]
|
||||
self.__write_jsonl(
|
||||
archive_dir / "output.jsonl",
|
||||
[outcome.results[x].to_record() for x in item_ids
|
||||
if x in outcome.results])
|
||||
meta["batch"] = outcome.batch
|
||||
meta["usage"] = self.__sum_usage(outcome.results)
|
||||
self.__write_meta(meta_path, meta)
|
||||
failed: list[str] = self.__find_failures(
|
||||
item_ids, outcome.results)
|
||||
return RunOutcome(
|
||||
item_count=len(items), dry_run=False,
|
||||
dry_run_request=None, failed=failed)
|
||||
|
||||
def __load_items(self) -> list[InputItem]:
|
||||
"""Load and validate the JSONL input items.
|
||||
|
||||
:return: The input items, in file order.
|
||||
:raises InputFormatError: When a line is malformed, an ID
|
||||
is duplicated, or the file contains no item.
|
||||
:raises OSError: When the file cannot be read.
|
||||
"""
|
||||
items: list[InputItem] = []
|
||||
seen: set[str] = set()
|
||||
with open(self.__input, encoding="utf-8") as file:
|
||||
for number, line in enumerate(file, start=1):
|
||||
if line.strip() == "":
|
||||
continue
|
||||
data: Any
|
||||
try:
|
||||
data = json.loads(line)
|
||||
except json.JSONDecodeError as error:
|
||||
raise InputFormatError(
|
||||
f"{self.__input}: line {number}: malformed"
|
||||
f" JSON: {error}")
|
||||
item: InputItem = InputItem.get_instance(
|
||||
data, self.__input, number)
|
||||
if item.id in seen:
|
||||
raise InputFormatError(
|
||||
f"{self.__input}: line {number}:"
|
||||
f" duplicated ID \"{item.id}\"")
|
||||
seen.add(item.id)
|
||||
items.append(item)
|
||||
if len(items) == 0:
|
||||
raise InputFormatError(f"{self.__input}: no input items")
|
||||
return items
|
||||
|
||||
def __create_archive_dir(self) -> Path:
|
||||
"""Create the archive directory.
|
||||
|
||||
Only this directory is ever created or removed; no other
|
||||
directory is ever touched.
|
||||
|
||||
:return: The created archive directory.
|
||||
:raises FileExistsError: When the archive directory
|
||||
already exists and ``replace`` is False.
|
||||
"""
|
||||
if self.__archive_dir.exists():
|
||||
if not self.__replace:
|
||||
raise FileExistsError(
|
||||
f"{self.__archive_dir} already exists; pass"
|
||||
" --replace to replace it")
|
||||
shutil.rmtree(self.__archive_dir)
|
||||
self.__archive_dir.mkdir(parents=True)
|
||||
return self.__archive_dir
|
||||
|
||||
def __build_meta(self, items: list[InputItem]) -> dict[str, Any]:
|
||||
"""Build the initial archive metadata.
|
||||
|
||||
:param items: The loaded input items.
|
||||
:return: The metadata, "batch" and "usage" not yet filled
|
||||
in for an actual run.
|
||||
"""
|
||||
return {
|
||||
"script_version": self.__SCRIPT_VERSION,
|
||||
"model": self.__model,
|
||||
"temperature": self.MODELS[self.__model].get(
|
||||
"temperature"),
|
||||
"thinking": self.MODELS[self.__model].get("thinking"),
|
||||
"max_tokens": self.__max_tokens,
|
||||
"prompt_path": str(self.__prompt),
|
||||
"prompt_sha256": self.__sha256_of(self.__prompt),
|
||||
"input_path": str(self.__input),
|
||||
"input_sha256": self.__sha256_of(self.__input),
|
||||
"item_count": len(items),
|
||||
"dry_run": self.__dry_run,
|
||||
"started_at": self.__now_iso(),
|
||||
"batch": None,
|
||||
"usage": {},
|
||||
}
|
||||
|
||||
def __build_request(self, item: InputItem, system_prompt: str) \
|
||||
-> dict[str, Any]:
|
||||
"""Build one Message Batches request for an input item.
|
||||
|
||||
:param item: The input item.
|
||||
:param system_prompt: The system prompt text.
|
||||
:return: The batch request with "custom_id" and "params".
|
||||
"""
|
||||
return {
|
||||
"custom_id": item.id,
|
||||
"params": {
|
||||
"model": self.__model,
|
||||
"max_tokens": self.__max_tokens,
|
||||
**self.MODELS[self.__model],
|
||||
"system": system_prompt,
|
||||
"messages": [
|
||||
{"role": "user", "content": item.content},
|
||||
],
|
||||
},
|
||||
}
|
||||
|
||||
def __execute_run(
|
||||
self, client: anthropic.Anthropic,
|
||||
items: list[InputItem], system_prompt: str) \
|
||||
-> ExecutionOutcome:
|
||||
"""Submit the batch of this run and await its results.
|
||||
|
||||
:param client: The Anthropic client.
|
||||
:param items: The input items.
|
||||
:param system_prompt: The system prompt text.
|
||||
:return: The submitted batch's bookkeeping and its
|
||||
results.
|
||||
"""
|
||||
requests: list[dict[str, Any]] = [
|
||||
self.__build_request(x, system_prompt) for x in items]
|
||||
info: BatchInfo = BatchInfo(
|
||||
batch_id=self.__submit_batch(client, requests),
|
||||
submitted_at=self.__now_iso())
|
||||
print(f"submitted batch {info.batch_id}", file=sys.stderr)
|
||||
batches: dict[str, Any] = self.__poll_batches(
|
||||
client, [info.batch_id])
|
||||
info.ended_at = batches[info.batch_id].ended_at.isoformat()
|
||||
results: Results = self.__collect_results(
|
||||
client, info.batch_id)
|
||||
return ExecutionOutcome(batch=info, results=results)
|
||||
|
||||
@staticmethod
|
||||
def __submit_batch(client: anthropic.Anthropic,
|
||||
requests: list[dict[str, Any]]) -> str:
|
||||
"""Submit one message batch.
|
||||
|
||||
:param client: The Anthropic client.
|
||||
:param requests: The batch requests.
|
||||
:return: The batch ID.
|
||||
"""
|
||||
return client.messages.batches.create(requests=requests).id
|
||||
|
||||
@classmethod
|
||||
def __poll_batches(cls, client: anthropic.Anthropic,
|
||||
batch_ids: list[str]) -> dict[str, Any]:
|
||||
"""Poll the batches until every one of them has ended.
|
||||
|
||||
Progress is printed to the standard error every poll.
|
||||
|
||||
:param client: The Anthropic client.
|
||||
:param batch_ids: The batch IDs to poll.
|
||||
:return: The final batch object of each batch, keyed by
|
||||
batch ID.
|
||||
"""
|
||||
while True:
|
||||
batches: dict[str, Any] = {
|
||||
x: client.messages.batches.retrieve(x)
|
||||
for x in batch_ids}
|
||||
pending: list[str] = [
|
||||
x for x in batch_ids
|
||||
if batches[x].processing_status != "ended"]
|
||||
for batch_id in batch_ids:
|
||||
status: str = batches[batch_id].processing_status
|
||||
print(f"batch {batch_id}: {status}", file=sys.stderr)
|
||||
if len(pending) == 0:
|
||||
return batches
|
||||
time.sleep(cls.__POLL_INTERVAL_SECONDS)
|
||||
|
||||
@staticmethod
|
||||
def __collect_results(client: anthropic.Anthropic,
|
||||
batch_id: str) -> Results:
|
||||
"""Collect the results of an ended batch.
|
||||
|
||||
:param client: The Anthropic client.
|
||||
:param batch_id: The batch ID.
|
||||
:return: The result records, keyed by custom ID.
|
||||
"""
|
||||
results: Results = {}
|
||||
for entry in client.messages.batches.results(batch_id):
|
||||
results[entry.custom_id] = BatchResult.get_instance(
|
||||
entry)
|
||||
return results
|
||||
|
||||
@staticmethod
|
||||
def __find_failures(item_ids: list[str],
|
||||
results: Results) -> list[str]:
|
||||
"""Find the item IDs that failed in a result set.
|
||||
|
||||
An item failed when it is missing from the results or when
|
||||
its record is a failure.
|
||||
|
||||
:param item_ids: The item IDs to check, in order.
|
||||
:param results: The result records, keyed by item ID.
|
||||
:return: The failed item IDs, in the given order.
|
||||
"""
|
||||
return [x for x in item_ids
|
||||
if x not in results or results[x].is_failure]
|
||||
|
||||
@staticmethod
|
||||
def __sum_usage(results: Results) -> dict[str, int]:
|
||||
"""Sum the token usage of every succeeded result.
|
||||
|
||||
:param results: The result records, keyed by item ID.
|
||||
:return: The summed integer usage fields.
|
||||
"""
|
||||
totals: dict[str, int] = {}
|
||||
for result in results.values():
|
||||
if result.usage is None:
|
||||
continue
|
||||
for key, value in result.usage.items():
|
||||
if isinstance(value, int):
|
||||
totals[key] = totals.get(key, 0) + value
|
||||
return totals
|
||||
|
||||
@staticmethod
|
||||
def __write_jsonl(path: Path,
|
||||
records: list[dict[str, Any]]) -> None:
|
||||
"""Write records to a file as JSON Lines.
|
||||
|
||||
:param path: The path of the file to write.
|
||||
:param records: The records, one per line.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
"""
|
||||
with open(path, "w", encoding="utf-8") as file:
|
||||
for record in records:
|
||||
file.write(
|
||||
json.dumps(record, ensure_ascii=False) + "\n")
|
||||
|
||||
@staticmethod
|
||||
def __write_json(path: Path, data: dict[str, Any]) -> None:
|
||||
"""Write data to a file as pretty-printed JSON.
|
||||
|
||||
:param path: The path of the file to write.
|
||||
:param data: The data to write.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
"""
|
||||
path.write_text(
|
||||
json.dumps(data, ensure_ascii=False, indent=2) + "\n",
|
||||
encoding="utf-8")
|
||||
|
||||
@classmethod
|
||||
def __write_meta(cls, path: Path, meta: dict[str, Any]) -> None:
|
||||
"""Write the metadata to the ``meta.json`` file.
|
||||
|
||||
The ``BatchInfo`` value under "batch" is written as a
|
||||
plain JSON object.
|
||||
|
||||
:param path: The path of the ``meta.json`` file.
|
||||
:param meta: The metadata to write.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
"""
|
||||
cls.__write_json(
|
||||
path, {**meta, "batch": asdict(meta["batch"])})
|
||||
|
||||
@staticmethod
|
||||
def __sha256_of(path: Path) -> str:
|
||||
"""Calculate the SHA-256 digest of a file.
|
||||
|
||||
:param path: The path of the file.
|
||||
:return: The hexadecimal SHA-256 digest.
|
||||
"""
|
||||
with open(path, "rb") as file:
|
||||
return hashlib.file_digest(file, "sha256").hexdigest()
|
||||
|
||||
@staticmethod
|
||||
def __now_iso() -> str:
|
||||
"""Return the current local time in ISO 8601 format.
|
||||
|
||||
:return: The current local time with the timezone offset.
|
||||
"""
|
||||
return datetime.now().astimezone().isoformat(
|
||||
timespec="seconds")
|
||||
|
||||
|
||||
def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
"""Parse the command-line arguments.
|
||||
|
||||
:param argv: The command-line arguments, or None for ``sys.argv``.
|
||||
:param argv: The command-line arguments, or None for
|
||||
``sys.argv``.
|
||||
:return: The parsed arguments.
|
||||
"""
|
||||
parser: argparse.ArgumentParser = argparse.ArgumentParser(
|
||||
@@ -180,7 +575,8 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
" and archive the result.")
|
||||
parser.add_argument(
|
||||
"prompt", type=Path,
|
||||
help="the prompt definition file, used as the system prompt")
|
||||
help="the prompt definition file, used as the system"
|
||||
" prompt")
|
||||
parser.add_argument(
|
||||
"input", type=Path,
|
||||
help="the JSONL input file with \"id\" and \"content\"")
|
||||
@@ -188,11 +584,13 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
"archive_dir", type=Path,
|
||||
help="the destination archive directory")
|
||||
parser.add_argument(
|
||||
"--model", choices=sorted(MODELS), default=DEFAULT_MODEL,
|
||||
help=f"the model ID (default {DEFAULT_MODEL})")
|
||||
"--model", choices=sorted(LLMRunner.MODELS),
|
||||
default=LLMRunner.DEFAULT_MODEL,
|
||||
help=f"the model ID (default {LLMRunner.DEFAULT_MODEL})")
|
||||
parser.add_argument(
|
||||
"--max-tokens", type=int, default=2048,
|
||||
help="the maximum output tokens per request (default 2048)")
|
||||
help="the maximum output tokens per request (default"
|
||||
" 2048)")
|
||||
parser.add_argument(
|
||||
"--dry-run", action="store_true",
|
||||
help="validate and archive without calling the API")
|
||||
@@ -202,327 +600,30 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
return parser.parse_args(argv)
|
||||
|
||||
|
||||
def load_items(path: Path) -> list[InputItem]:
|
||||
"""Load and validate the JSONL input items.
|
||||
|
||||
:param path: The path of the JSONL input file.
|
||||
:return: The input items, in file order.
|
||||
:raises InputFormatError: When a line is malformed, an ID is
|
||||
duplicated, or the file contains no item.
|
||||
:raises OSError: When the file cannot be read.
|
||||
"""
|
||||
items: list[InputItem] = []
|
||||
seen: set[str] = set()
|
||||
with open(path, encoding="utf-8") as file:
|
||||
for number, line in enumerate(file, start=1):
|
||||
if line.strip() == "":
|
||||
continue
|
||||
try:
|
||||
data: Any = json.loads(line)
|
||||
except json.JSONDecodeError as error:
|
||||
raise InputFormatError(
|
||||
f"{path}: line {number}: malformed JSON: {error}")
|
||||
item: InputItem = InputItem.get_instance(
|
||||
data, path, number)
|
||||
if item.id in seen:
|
||||
raise InputFormatError(
|
||||
f"{path}: line {number}: duplicated ID"
|
||||
f" \"{item.id}\"")
|
||||
seen.add(item.id)
|
||||
items.append(item)
|
||||
if len(items) == 0:
|
||||
raise InputFormatError(f"{path}: no input items")
|
||||
return items
|
||||
|
||||
|
||||
def build_request(item: InputItem, system_prompt: str,
|
||||
max_tokens: int, model: str) -> dict[str, Any]:
|
||||
"""Build one Message Batches request for an input item.
|
||||
|
||||
:param item: The input item.
|
||||
:param system_prompt: The system prompt text.
|
||||
:param max_tokens: The maximum output tokens.
|
||||
:param model: The model ID, a key of ``MODELS``.
|
||||
:return: The batch request with "custom_id" and "params".
|
||||
"""
|
||||
return {
|
||||
"custom_id": item.id,
|
||||
"params": {
|
||||
"model": model,
|
||||
"max_tokens": max_tokens,
|
||||
**MODELS[model],
|
||||
"system": system_prompt,
|
||||
"messages": [
|
||||
{"role": "user", "content": item.content},
|
||||
],
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def submit_batch(client: anthropic.Anthropic,
|
||||
requests: list[dict[str, Any]]) -> str:
|
||||
"""Submit one message batch.
|
||||
|
||||
:param client: The Anthropic client.
|
||||
:param requests: The batch requests.
|
||||
:return: The batch ID.
|
||||
"""
|
||||
return client.messages.batches.create(requests=requests).id
|
||||
|
||||
|
||||
def poll_batches(client: anthropic.Anthropic,
|
||||
batch_ids: list[str]) -> dict[str, Any]:
|
||||
"""Poll the batches until every one of them has ended.
|
||||
|
||||
Progress is printed to the standard error every poll.
|
||||
|
||||
:param client: The Anthropic client.
|
||||
:param batch_ids: The batch IDs to poll.
|
||||
:return: The final batch object of each batch, keyed by batch ID.
|
||||
"""
|
||||
while True:
|
||||
batches: dict[str, Any] = {
|
||||
x: client.messages.batches.retrieve(x) for x in batch_ids}
|
||||
pending: list[str] = [
|
||||
x for x in batch_ids
|
||||
if batches[x].processing_status != "ended"]
|
||||
for batch_id in batch_ids:
|
||||
status: str = batches[batch_id].processing_status
|
||||
print(f"batch {batch_id}: {status}", file=sys.stderr)
|
||||
if len(pending) == 0:
|
||||
return batches
|
||||
time.sleep(POLL_INTERVAL_SECONDS)
|
||||
|
||||
|
||||
def usage_to_dict(usage: Any) -> dict[str, Any]:
|
||||
"""Convert a usage object to a plain dictionary.
|
||||
|
||||
:param usage: The usage object of a message.
|
||||
:return: The usage as a dictionary, without null entries.
|
||||
"""
|
||||
return {k: v for k, v in usage.model_dump().items()
|
||||
if v is not None}
|
||||
|
||||
|
||||
def sum_usage(results: Results) -> dict[str, int]:
|
||||
"""Sum the token usage of every succeeded result.
|
||||
|
||||
:param results: The result records, keyed by item ID.
|
||||
:return: The summed integer usage fields.
|
||||
"""
|
||||
totals: dict[str, int] = {}
|
||||
for result in results.values():
|
||||
if result.usage is None:
|
||||
continue
|
||||
for key, value in result.usage.items():
|
||||
if isinstance(value, int):
|
||||
totals[key] = totals.get(key, 0) + value
|
||||
return totals
|
||||
|
||||
|
||||
def collect_results(client: anthropic.Anthropic,
|
||||
batch_id: str) -> Results:
|
||||
"""Collect the results of an ended batch.
|
||||
|
||||
:param client: The Anthropic client.
|
||||
:param batch_id: The batch ID.
|
||||
:return: The result records, keyed by custom ID.
|
||||
"""
|
||||
results: Results = {}
|
||||
for entry in client.messages.batches.results(batch_id):
|
||||
results[entry.custom_id] = BatchResult.get_instance(entry)
|
||||
return results
|
||||
|
||||
|
||||
def find_failures(item_ids: list[str],
|
||||
results: Results) -> list[str]:
|
||||
"""Find the item IDs that failed in a result set.
|
||||
|
||||
An item failed when it is missing from the results or when its
|
||||
record is a failure.
|
||||
|
||||
:param item_ids: The item IDs to check, in order.
|
||||
:param results: The result records, keyed by item ID.
|
||||
:return: The failed item IDs, in the given order.
|
||||
"""
|
||||
return [x for x in item_ids
|
||||
if x not in results or results[x].is_failure]
|
||||
|
||||
|
||||
def create_archive_dir(directory: Path, replace: bool) -> Path:
|
||||
"""Create the archive directory.
|
||||
|
||||
Only this directory is ever created or removed; no other
|
||||
directory is ever touched.
|
||||
|
||||
:param directory: The destination archive directory.
|
||||
:param replace: Whether to remove an already existing archive
|
||||
directory before creating it.
|
||||
:return: The created archive directory.
|
||||
:raises FileExistsError: When the archive directory already
|
||||
exists and ``replace`` is False.
|
||||
"""
|
||||
if directory.exists():
|
||||
if not replace:
|
||||
raise FileExistsError(
|
||||
f"{directory} already exists; pass --replace to"
|
||||
" replace it")
|
||||
shutil.rmtree(directory)
|
||||
directory.mkdir(parents=True)
|
||||
return directory
|
||||
|
||||
|
||||
def write_jsonl(path: Path, records: list[dict[str, Any]]) -> None:
|
||||
"""Write records to a file as JSON Lines.
|
||||
|
||||
:param path: The path of the file to write.
|
||||
:param records: The records, one per line.
|
||||
:return: None.
|
||||
"""
|
||||
with open(path, "w", encoding="utf-8") as file:
|
||||
for record in records:
|
||||
file.write(json.dumps(record, ensure_ascii=False) + "\n")
|
||||
|
||||
|
||||
def write_json(path: Path, data: dict[str, Any]) -> None:
|
||||
"""Write data to a file as pretty-printed JSON.
|
||||
|
||||
:param path: The path of the file to write.
|
||||
:param data: The data to write.
|
||||
:return: None.
|
||||
"""
|
||||
path.write_text(
|
||||
json.dumps(data, ensure_ascii=False, indent=2) + "\n",
|
||||
encoding="utf-8")
|
||||
|
||||
|
||||
def write_meta(path: Path, meta: dict[str, Any]) -> None:
|
||||
"""Write the metadata to the ``meta.json`` file.
|
||||
|
||||
The ``BatchInfo`` value under ``batch`` is written as a plain
|
||||
JSON object.
|
||||
|
||||
:param path: The path of the ``meta.json`` file.
|
||||
:param meta: The metadata to write.
|
||||
:return: None.
|
||||
"""
|
||||
write_json(path, {**meta, "batch": asdict(meta["batch"])})
|
||||
|
||||
|
||||
def sha256_of(path: Path) -> str:
|
||||
"""Calculate the SHA-256 digest of a file.
|
||||
|
||||
:param path: The path of the file.
|
||||
:return: The hexadecimal SHA-256 digest.
|
||||
"""
|
||||
with open(path, "rb") as file:
|
||||
return hashlib.file_digest(file, "sha256").hexdigest()
|
||||
|
||||
|
||||
def now_iso() -> str:
|
||||
"""Return the current local time in ISO 8601 format.
|
||||
|
||||
:return: The current local time with the timezone offset.
|
||||
"""
|
||||
return datetime.now().astimezone().isoformat(timespec="seconds")
|
||||
|
||||
|
||||
def execute_run(
|
||||
client: anthropic.Anthropic, items: list[InputItem],
|
||||
system_prompt: str, max_tokens: int, model: str,
|
||||
meta: dict[str, Any],
|
||||
) -> Results:
|
||||
"""Submit the batch of this run and await its results.
|
||||
|
||||
The batch ID and timestamps are recorded into the metadata as an
|
||||
observable side effect.
|
||||
|
||||
:param client: The Anthropic client.
|
||||
:param items: The input items.
|
||||
:param system_prompt: The system prompt text.
|
||||
:param max_tokens: The maximum output tokens per request.
|
||||
:param model: The model ID, a key of ``MODELS``.
|
||||
:param meta: The metadata to record the batch bookkeeping into.
|
||||
:return: The results of this run, keyed by item ID.
|
||||
"""
|
||||
requests: list[dict[str, Any]] = [
|
||||
build_request(x, system_prompt, max_tokens, model)
|
||||
for x in items]
|
||||
info: BatchInfo = BatchInfo(
|
||||
batch_id=submit_batch(client, requests),
|
||||
submitted_at=now_iso())
|
||||
meta["batch"] = info
|
||||
print(f"submitted batch {info.batch_id}", file=sys.stderr)
|
||||
batches: dict[str, Any] = poll_batches(client, [info.batch_id])
|
||||
info.ended_at = batches[info.batch_id].ended_at.isoformat()
|
||||
return collect_results(client, info.batch_id)
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
"""Run one LLM definition file against one input and archive it.
|
||||
|
||||
:param argv: The command-line arguments, or None for ``sys.argv``.
|
||||
:param argv: The command-line arguments, or None for
|
||||
``sys.argv``.
|
||||
:return: The exit status: 0 on success, non-zero on failure.
|
||||
"""
|
||||
started: float = time.monotonic()
|
||||
args: argparse.Namespace = parse_args(argv)
|
||||
try:
|
||||
items: list[InputItem] = load_items(args.input)
|
||||
prompt_text: str = args.prompt.read_text(encoding="utf-8")
|
||||
except (OSError, InputFormatError) as error:
|
||||
outcome: RunOutcome = LLMRunner(
|
||||
args.prompt, args.input, args.archive_dir, args.model,
|
||||
args.max_tokens, args.dry_run, args.replace).run()
|
||||
except (InputFormatError, OSError) as error:
|
||||
print(f"error: {error}", file=sys.stderr)
|
||||
return 1
|
||||
try:
|
||||
archive_dir: Path = create_archive_dir(
|
||||
args.archive_dir, args.replace)
|
||||
except FileExistsError as error:
|
||||
print(f"error: {error}", file=sys.stderr)
|
||||
return 1
|
||||
meta_path: Path = archive_dir / "meta.json"
|
||||
(archive_dir / "prompt.md").write_bytes(args.prompt.read_bytes())
|
||||
meta: dict[str, Any] = {
|
||||
"script_version": SCRIPT_VERSION,
|
||||
"model": args.model,
|
||||
"temperature": MODELS[args.model].get("temperature"),
|
||||
"thinking": MODELS[args.model].get("thinking"),
|
||||
"max_tokens": args.max_tokens,
|
||||
"prompt_path": str(args.prompt),
|
||||
"prompt_sha256": sha256_of(args.prompt),
|
||||
"input_path": str(args.input),
|
||||
"input_sha256": sha256_of(args.input),
|
||||
"item_count": len(items),
|
||||
"dry_run": args.dry_run,
|
||||
"started_at": now_iso(),
|
||||
"batch": None,
|
||||
"usage": {},
|
||||
}
|
||||
if args.dry_run:
|
||||
write_json(meta_path, meta)
|
||||
print(json.dumps(
|
||||
build_request(items[0], prompt_text, args.max_tokens,
|
||||
args.model),
|
||||
ensure_ascii=False, indent=2))
|
||||
elapsed: str = format_duration(time.monotonic() - started)
|
||||
print(f"Done. {len(items)} jobs finished."
|
||||
f" {elapsed} elapsed.", file=sys.stderr)
|
||||
return 0
|
||||
client: anthropic.Anthropic = anthropic.Anthropic(
|
||||
api_key=get_settings().ANTHROPIC_API_KEY)
|
||||
results: Results = execute_run(
|
||||
client, items, prompt_text, args.max_tokens, args.model,
|
||||
meta)
|
||||
item_ids: list[str] = [x.id for x in items]
|
||||
write_jsonl(
|
||||
archive_dir / "output.jsonl",
|
||||
[results[x].to_record() for x in item_ids if x in results])
|
||||
meta["usage"] = sum_usage(results)
|
||||
write_meta(meta_path, meta)
|
||||
failed: list[str] = find_failures(item_ids, results)
|
||||
if len(failed) > 0:
|
||||
print(f"error: failed items: {', '.join(failed)}",
|
||||
if not outcome.dry_run and len(outcome.failed) > 0:
|
||||
print(f"error: failed items: {', '.join(outcome.failed)}",
|
||||
file=sys.stderr)
|
||||
return 1
|
||||
if outcome.dry_run:
|
||||
print(json.dumps(
|
||||
outcome.dry_run_request, ensure_ascii=False, indent=2))
|
||||
elapsed: str = format_duration(time.monotonic() - started)
|
||||
print(f"Done. {len(items)} jobs finished."
|
||||
print(f"Done. {outcome.item_count} jobs finished."
|
||||
f" {elapsed} elapsed.", file=sys.stderr)
|
||||
return 0
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -6,49 +6,16 @@ r"""The majority tally of the three coding runs.
|
||||
|
||||
Settles the coding step: the same coding definition file is run
|
||||
three times independently, and this command counts the votes and
|
||||
writes the final coding table the paper cites, as the CSV file
|
||||
given as the fourth positional command-line argument. Only the
|
||||
keyword key sets of the three runs' archived ``output.jsonl``
|
||||
files take part in the tally; the lyric quotes never do. A
|
||||
(song, keyword) pair is written out when at least two of the
|
||||
three runs assign it, so three votes never tie, and it carries
|
||||
the lyric quotes of every run that assigned it, pooled,
|
||||
deduplicated, sorted by Unicode code point, and joined with a
|
||||
single ``|``: the three runs are peers, so the quote order
|
||||
follows the text alone. A quote carries the lyric line-break
|
||||
convention ``" / "`` where the lyric has a newline, applied once
|
||||
when the run records load, so the corrections, the coding table,
|
||||
and the database all share the one representation and nothing is
|
||||
ever converted back. No lyric of the 883-song corpus contains
|
||||
``" / "`` -- a corpus fact checked exhaustively, not a structural
|
||||
guarantee -- so the convention is unambiguous here. The three
|
||||
archives must cover exactly the same set of song IDs, every
|
||||
record must be a successful result, and every record's "text"
|
||||
must parse to a JSON object; otherwise the tally fails and
|
||||
nothing is written.
|
||||
|
||||
Two optional inputs guard the tally. ``--corrections`` names a
|
||||
CSV file of researcher-reviewed repairs, applied to each run's
|
||||
records before anything else happens: a keyword row renames or
|
||||
drops one keyword assignment of one song in one run, and an
|
||||
evidence row rewrites or drops one lyric quote string wherever it
|
||||
appears in that song's record for that run. Its two text fields
|
||||
carry the two characters ``\n`` where the text has a newline, so
|
||||
the file holds one row per line. Every row must match, so a
|
||||
stale row fails the run. ``--valid-keywords`` names a
|
||||
plain text file of the allowed keywords, one per line; once the
|
||||
corrections are in, every keyword left in any record must appear
|
||||
in it. The order is fixed and matters: the corrections come
|
||||
first, so a repair may reunite the votes of a misspelled keyword
|
||||
that the check would otherwise reject. With neither option, no
|
||||
record is touched and no vocabulary is checked.
|
||||
|
||||
The archives identify a song as ``song-<ID>``, where ``<ID>`` is
|
||||
the song's ID in the SQLite working store. The output table does
|
||||
not carry that ID: every song is looked up in the working store
|
||||
and written as its title and its stored artist credit instead, so
|
||||
this command runs after ``build-db``. The step is fully
|
||||
deterministic; no LLM call is made.
|
||||
writes the final coding table the paper cites. A (song, keyword)
|
||||
pair is written out when at least two of the three runs assign
|
||||
it, carrying the pooled, deduplicated lyric quotes of the runs
|
||||
that assigned it. ``--corrections`` names a CSV file of
|
||||
researcher-reviewed repairs to a run's records, applied before
|
||||
the tally. ``--valid-keywords`` names a plain text file of the
|
||||
allowed keywords that every record's keywords must appear in.
|
||||
The songs are named from the working store, so this command runs
|
||||
after ``build-db``. When any input is malformed, the tally fails
|
||||
and nothing is written; the error message names what failed.
|
||||
"""
|
||||
import argparse
|
||||
import csv
|
||||
@@ -127,11 +94,6 @@ class Correction:
|
||||
class CorrectionTable:
|
||||
"""The researcher-reviewed repairs of the runs' records."""
|
||||
|
||||
MANUAL_CORRECTIONS_CSV: ClassVar[str] \
|
||||
= "coding-corrections.csv"
|
||||
"""The correction table CSV file's conventional name under
|
||||
``data/manual/``."""
|
||||
|
||||
path: Path
|
||||
"""The correction table CSV file the repairs came from."""
|
||||
corrections: list[Correction]
|
||||
@@ -141,7 +103,7 @@ class CorrectionTable:
|
||||
class CorrectionsLoader:
|
||||
"""The loader of the researcher-reviewed correction table."""
|
||||
|
||||
__HEADER: tuple[str, str, str, str, str] = (
|
||||
__HEADER: ClassVar[tuple[str, str, str, str, str]] = (
|
||||
"Song ID", "Run", "Type", "To Be Replaced", "Correct Term")
|
||||
"""The header row the correction table CSV file must carry."""
|
||||
|
||||
@@ -164,10 +126,8 @@ class CorrectionsLoader:
|
||||
of the runs the command was given, and a known type. The
|
||||
file is read with the CSV reader, so a quoted field may
|
||||
hold a comma or a double quote. No field holds a line
|
||||
break: the two text fields carry the lyric line-break
|
||||
convention ``" / "`` where the text has a newline -- the
|
||||
same representation the loaded run records carry -- and
|
||||
are matched and applied verbatim. Nothing is written.
|
||||
break (see the line-break convention on
|
||||
``CodingTallier``). Nothing is written.
|
||||
|
||||
:return: The repairs, in file order.
|
||||
:raises TallyError: When the file cannot be read, the
|
||||
@@ -340,16 +300,16 @@ class TalliedCodings:
|
||||
class CodingTallier:
|
||||
"""The tallier of the three coding runs' keyword votes."""
|
||||
|
||||
__MAJORITY: int = 2
|
||||
__MAJORITY: ClassVar[int] = 2
|
||||
"""The number of runs that must assign a keyword to a song for
|
||||
that code to be settled."""
|
||||
__MAX_REPORTED_IDS: int = 10
|
||||
__MAX_REPORTED_IDS: ClassVar[int] = 10
|
||||
"""The number of song IDs an error message lists before
|
||||
summarizing the rest as a count."""
|
||||
__QUOTE_SEPARATOR: str = "|"
|
||||
__QUOTE_SEPARATOR: ClassVar[str] = "|"
|
||||
"""The separator between the distinct lyric quotes of one
|
||||
settled code."""
|
||||
__LINE_BREAK: str = " / "
|
||||
__LINE_BREAK: ClassVar[str] = " / "
|
||||
"""The lyric line-break convention replacing every LF inside a
|
||||
quote. Unambiguous for this corpus only: none of the 883
|
||||
songs' lyrics contains the three characters, checked
|
||||
@@ -616,11 +576,8 @@ class CodingTallier:
|
||||
if line.strip() == "":
|
||||
continue
|
||||
record: Any = cls.__parse_json(line, str(path))
|
||||
if not isinstance(record, dict) or "id" not in record:
|
||||
raise ValueError(
|
||||
f"{path}: record without \"id\": {line}")
|
||||
item_id: Any = record["id"]
|
||||
if "error" in record or "text" not in record:
|
||||
if "text" not in record:
|
||||
raise ValueError(
|
||||
f"{path}: id {item_id}: not a successful"
|
||||
" result")
|
||||
@@ -647,8 +604,8 @@ class CodingTallier:
|
||||
:param label: The location of the record, for the error
|
||||
message.
|
||||
:return: The lyric quotes of every keyword, in the given
|
||||
order, every LF inside a quote turned into the lyric
|
||||
line-break convention ``" / "``.
|
||||
order, every LF inside a quote turned into the
|
||||
line-break convention.
|
||||
:raises ValueError: When a keyword's value is not a list
|
||||
of strings.
|
||||
"""
|
||||
@@ -739,7 +696,7 @@ class CodingTallier:
|
||||
first: set[int] = set(runs[0])
|
||||
index: int
|
||||
records: dict[int, dict[str, list[str]]]
|
||||
for index, records in enumerate(runs):
|
||||
for index, records in enumerate(runs[1:], start=1):
|
||||
song_ids: set[int] = set(records)
|
||||
if song_ids == first:
|
||||
continue
|
||||
@@ -813,9 +770,6 @@ class CodingTallier:
|
||||
class CodingTable:
|
||||
"""The final coding table the paper cites."""
|
||||
|
||||
RESULT_CODINGS_CSV: ClassVar[str] = "codings.csv"
|
||||
"""The coding table CSV file's conventional name under
|
||||
``results/``."""
|
||||
__HEADER: ClassVar[tuple[str, str, str, str]] \
|
||||
= ("Song", "Artist Credit", "Keyword", "Quote")
|
||||
"""The header row of the coding table CSV file."""
|
||||
@@ -833,11 +787,8 @@ class CodingTable:
|
||||
endings, carrying the header row
|
||||
``Song,Artist Credit,Keyword,Quote`` and one row per
|
||||
settled keyword, in the row order. Every field is
|
||||
written verbatim; a quote carries the lyric line-break
|
||||
convention ``" / "`` where the lyric has a line break, so
|
||||
no field holds a line break and the file holds one row
|
||||
per line. The parent directory is created when it does
|
||||
not exist.
|
||||
written verbatim, so the file holds one row per line.
|
||||
The parent directory is created when it does not exist.
|
||||
|
||||
:param output_csv: The output CSV file.
|
||||
:return: None.
|
||||
@@ -970,8 +921,7 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
help="the third coding run's archive directory")
|
||||
parser.add_argument(
|
||||
"output_csv", type=Path,
|
||||
help="the output CSV file, by convention"
|
||||
f" results/{CodingTable.RESULT_CODINGS_CSV}")
|
||||
help="the output CSV file")
|
||||
parser.add_argument(
|
||||
"--valid-keywords", type=Path, default=None,
|
||||
help="a plain text file of the allowed keywords, one per"
|
||||
@@ -980,33 +930,23 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
parser.add_argument(
|
||||
"--corrections", type=Path, default=None,
|
||||
help="the researcher-reviewed correction table CSV file,"
|
||||
" by convention"
|
||||
f" data/manual/{CorrectionTable.MANUAL_CORRECTIONS_CSV},"
|
||||
" applied to the runs' records before the tally"
|
||||
" (default: no repair)")
|
||||
return parser.parse_args(argv)
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
r"""Settle the coding by a majority of the three coding runs.
|
||||
"""Settle the coding by a majority of the three coding runs.
|
||||
|
||||
Writes the final coding table as the given CSV file, holding
|
||||
the header row ``Song,Artist Credit,Keyword,Quote`` and one
|
||||
row per keyword at least two of the three runs assign, the
|
||||
song named by its title and its stored artist credit from the
|
||||
SQLite working store, and the keyword carrying the pooled,
|
||||
deduplicated, and sorted lyric quotes of the runs that
|
||||
assigned it, joined with a single ``|`` and carrying the
|
||||
lyric line-break convention ``" / "`` where the lyric has a
|
||||
newline, so the table holds one row per line. The records are
|
||||
repaired from the ``--corrections`` table and then checked
|
||||
against the ``--valid-keywords`` list, when either is given.
|
||||
Nothing is written when the three archives do not cover the
|
||||
same songs, a record is not a successful result, a record's
|
||||
"text" does not parse to a JSON object of quote string lists,
|
||||
a correction is invalid or matches nothing, a keyword is not
|
||||
in the valid keyword list, or a song is not in the working
|
||||
store; the error message names what failed.
|
||||
Writes the final coding table CSV file described in the
|
||||
module docstring. The records are repaired from the
|
||||
``--corrections`` table and then checked against the
|
||||
``--valid-keywords`` list, when either is given. Nothing is
|
||||
written when the three archives do not cover the same songs,
|
||||
a record is not a successful result, a correction is invalid
|
||||
or matches nothing, a keyword is not in the valid keyword
|
||||
list, or a song is not in the working store; the error message
|
||||
names what failed.
|
||||
|
||||
:param argv: The command-line arguments, or None for
|
||||
``sys.argv``.
|
||||
|
||||
@@ -8,46 +8,281 @@
|
||||
Settles the semantic code groups of step 4: the same group
|
||||
selection definition file is run three times independently, and
|
||||
this command counts the votes and writes the final group table
|
||||
the paper cites, as the CSV file given as the last positional
|
||||
command-line argument. A (group, keyword) pair is written out
|
||||
when at least two of the three runs select it, so three votes
|
||||
never tie. A selected item that is not in the valid keyword
|
||||
list is invalid and casts no vote; every dropped occurrence is
|
||||
reported on standard error. The group name is the record ID
|
||||
with its ``group-`` prefix dropped. The rows are ordered by the
|
||||
group name and then by the keyword, by Unicode code point, and
|
||||
the file carries the header row ``Group,Keyword,Votes`` with
|
||||
CRLF line endings per RFC 4180.
|
||||
|
||||
The three archives must cover exactly the same set of group IDs,
|
||||
every record ID must carry the ``group-`` prefix, and every
|
||||
record's "text" must parse to a JSON array of strings; otherwise
|
||||
the tally fails and nothing is written.
|
||||
the paper cites. A (group, keyword) pair is written out when at
|
||||
least two of the three runs select it. When an input is
|
||||
malformed, the tally fails and nothing is written; the error
|
||||
message names what failed.
|
||||
"""
|
||||
import argparse
|
||||
import csv
|
||||
import json
|
||||
import sys
|
||||
import time
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
from typing import Any, ClassVar
|
||||
|
||||
from ..utils import format_duration
|
||||
|
||||
GROUP_ID_PREFIX: str = "group-"
|
||||
"""The prefix every group record ID must carry; the group name
|
||||
is the rest of the ID."""
|
||||
MAJORITY: int = 2
|
||||
"""The number of runs that must select a keyword for a group for
|
||||
that pair to be settled."""
|
||||
HEADER: tuple[str, str, str] = ("Group", "Keyword", "Votes")
|
||||
"""The header row of the group table CSV file."""
|
||||
|
||||
|
||||
class TallyError(Exception):
|
||||
"""An error that fails the group tally."""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class TalliedGroups:
|
||||
"""The outcome of settling the group table."""
|
||||
|
||||
codes: int
|
||||
"""The number of settled (group, keyword) pairs."""
|
||||
groups: int
|
||||
"""The number of groups the three runs cover."""
|
||||
|
||||
|
||||
class GroupTallier:
|
||||
"""The tallier of the three group-selection runs' votes."""
|
||||
|
||||
__GROUP_ID_PREFIX: ClassVar[str] = "group-"
|
||||
"""The prefix every group record ID must carry; the group
|
||||
name is the rest of the ID."""
|
||||
__MAJORITY: ClassVar[int] = 2
|
||||
"""The number of runs that must select a keyword for a group
|
||||
for that pair to be settled."""
|
||||
__HEADER: ClassVar[tuple[str, str, str]] \
|
||||
= ("Group", "Keyword", "Votes")
|
||||
"""The header row of the group table CSV file."""
|
||||
|
||||
def __init__(self, run_dir_1: Path, run_dir_2: Path,
|
||||
run_dir_3: Path, valid_keywords_txt: Path,
|
||||
output_csv: Path) -> None:
|
||||
"""Set up the tallier of the three selection runs.
|
||||
|
||||
:param run_dir_1: The first selection run's archive
|
||||
directory, containing ``output.jsonl``.
|
||||
:param run_dir_2: The second selection run's archive
|
||||
directory, containing ``output.jsonl``.
|
||||
:param run_dir_3: The third selection run's archive
|
||||
directory, containing ``output.jsonl``.
|
||||
:param valid_keywords_txt: The plain text file of the
|
||||
allowed keywords, one per line.
|
||||
:param output_csv: The output group table CSV file.
|
||||
"""
|
||||
self.__run_dirs: list[Path] = [
|
||||
run_dir_1, run_dir_2, run_dir_3]
|
||||
"""The three runs' archive directories, in the given
|
||||
order."""
|
||||
self.__valid_keywords_txt: Path = valid_keywords_txt
|
||||
"""The plain text file of the allowed keywords, one per
|
||||
line."""
|
||||
self.__output_csv: Path = output_csv
|
||||
"""The output group table CSV file."""
|
||||
|
||||
def run(self) -> TalliedGroups:
|
||||
"""Load the three runs, tally their votes, and write the
|
||||
table.
|
||||
|
||||
:return: The settled code count and the group count.
|
||||
:raises TallyError: When a file cannot be read, a line is
|
||||
not a well-formed output record, a record is not a
|
||||
successful result, an ID lacks the ``group-`` prefix,
|
||||
a group has two records, a "text" does not parse to a
|
||||
JSON array of strings, or the three runs do not cover
|
||||
the same set of groups.
|
||||
:raises OSError: When the output file cannot be written.
|
||||
"""
|
||||
valid: set[str] = self.__load_valid_keywords(
|
||||
self.__valid_keywords_txt)
|
||||
runs: list[dict[str, set[str]]] = [
|
||||
self.__load_run(x) for x in self.__run_dirs]
|
||||
self.__drop_invalid(runs, self.__run_dirs, valid)
|
||||
rows: list[tuple[str, str, int]] = self.__tally(runs)
|
||||
self.__write_csv(rows)
|
||||
return TalliedGroups(codes=len(rows), groups=len(runs[0]))
|
||||
|
||||
@staticmethod
|
||||
def __load_valid_keywords(path: Path) -> set[str]:
|
||||
"""Load the valid keyword list.
|
||||
|
||||
:param path: The plain text file of the allowed keywords,
|
||||
one per line.
|
||||
:return: The allowed keywords.
|
||||
:raises TallyError: When the file cannot be read or holds
|
||||
no keyword.
|
||||
"""
|
||||
text: str
|
||||
try:
|
||||
text = path.read_text(encoding="utf-8")
|
||||
except OSError as error:
|
||||
raise TallyError(str(error)) from error
|
||||
keywords: set[str] = {x.strip() for x in text.split("\n")
|
||||
if x.strip() != ""}
|
||||
if len(keywords) == 0:
|
||||
raise TallyError(f"{path}: no keywords")
|
||||
return keywords
|
||||
|
||||
@classmethod
|
||||
def __load_run(cls, run_dir: Path) -> dict[str, set[str]]:
|
||||
"""Load and validate the selection records of one run.
|
||||
|
||||
:param run_dir: The run's archive directory, containing
|
||||
``output.jsonl``.
|
||||
:return: The selected keywords of every group of the run,
|
||||
keyed by the group name, the duplicates within one
|
||||
record's selection collapsed.
|
||||
:raises TallyError: When the file cannot be read, a line
|
||||
is not a well-formed output record, a record is not a
|
||||
successful result, an ID lacks the ``group-`` prefix,
|
||||
a group has two records, or a "text" does not parse
|
||||
to a JSON array of strings.
|
||||
"""
|
||||
path: Path = run_dir / "output.jsonl"
|
||||
text: str
|
||||
try:
|
||||
text = path.read_text(encoding="utf-8")
|
||||
except OSError as error:
|
||||
raise TallyError(str(error)) from error
|
||||
records: dict[str, set[str]] = {}
|
||||
line: str
|
||||
for line in text.split("\n"):
|
||||
if line.strip() == "":
|
||||
continue
|
||||
record: Any
|
||||
try:
|
||||
record = json.loads(line)
|
||||
except json.JSONDecodeError as error:
|
||||
raise TallyError(
|
||||
f"{path}: malformed JSON: {error}") from error
|
||||
item_id: Any = record["id"]
|
||||
if "text" not in record:
|
||||
raise TallyError(
|
||||
f"{path}: id {item_id}: not a successful"
|
||||
" result")
|
||||
if not isinstance(item_id, str) \
|
||||
or not item_id.startswith(
|
||||
cls.__GROUP_ID_PREFIX) \
|
||||
or item_id == cls.__GROUP_ID_PREFIX:
|
||||
raise TallyError(
|
||||
f"{path}: id {item_id}: not in the"
|
||||
f" \"{cls.__GROUP_ID_PREFIX}<name>\" form")
|
||||
group: str = item_id[len(cls.__GROUP_ID_PREFIX):]
|
||||
if group in records:
|
||||
raise TallyError(
|
||||
f"{path}: id {item_id}: duplicate record")
|
||||
records[group] = cls.__parse_selection(
|
||||
record["text"], f"{path}: id {item_id}")
|
||||
return records
|
||||
|
||||
@staticmethod
|
||||
def __parse_selection(text: Any, label: str) -> set[str]:
|
||||
"""Parse and validate the selected keywords of one record.
|
||||
|
||||
:param text: The "text" field of the record.
|
||||
:param label: The location of the record, for the error
|
||||
message.
|
||||
:return: The selected keywords, the duplicates collapsed.
|
||||
:raises TallyError: When the text does not parse to a
|
||||
JSON array of strings.
|
||||
"""
|
||||
selected: Any
|
||||
try:
|
||||
selected = json.loads(text)
|
||||
except json.JSONDecodeError as error:
|
||||
raise TallyError(
|
||||
f"{label}: \"text\" is malformed JSON:"
|
||||
f" {error}") from error
|
||||
if not isinstance(selected, list) \
|
||||
or not all(isinstance(x, str) for x in selected):
|
||||
raise TallyError(
|
||||
f"{label}: \"text\" does not parse to a JSON"
|
||||
" array of strings")
|
||||
return set(selected)
|
||||
|
||||
@staticmethod
|
||||
def __drop_invalid(runs: list[dict[str, set[str]]],
|
||||
run_dirs: list[Path],
|
||||
valid: set[str]) -> None:
|
||||
"""Drop the out-of-vocabulary selections of every run.
|
||||
|
||||
Every dropped occurrence is reported on standard error as
|
||||
an observable side effect.
|
||||
|
||||
:param runs: The runs' records, filtered in place.
|
||||
:param run_dirs: The run directories, for the messages.
|
||||
:param valid: The allowed keywords.
|
||||
:return: None.
|
||||
"""
|
||||
records: dict[str, set[str]]
|
||||
run_dir: Path
|
||||
for records, run_dir in zip(runs, run_dirs):
|
||||
group: str
|
||||
selected: set[str]
|
||||
for group, selected in records.items():
|
||||
keyword: str
|
||||
for keyword in sorted(selected - valid):
|
||||
print(
|
||||
f"note: {run_dir.name} group-{group}:"
|
||||
f" dropped out-of-vocabulary item"
|
||||
f" \"{keyword}\"", file=sys.stderr)
|
||||
records[group] = selected & valid
|
||||
|
||||
@classmethod
|
||||
def __tally(cls, runs: list[dict[str, set[str]]]) \
|
||||
-> list[tuple[str, str, int]]:
|
||||
"""Tally the keyword votes of the runs, group by group.
|
||||
|
||||
:param runs: The runs' records, all covering the same set
|
||||
of groups.
|
||||
:return: The settled rows, each the group name, the
|
||||
keyword, and the number of votes, ordered by the
|
||||
group name and then by the keyword, by Unicode code
|
||||
point.
|
||||
:raises TallyError: When the runs do not cover the same
|
||||
set of groups.
|
||||
"""
|
||||
groups: set[str] = set(runs[0])
|
||||
records: dict[str, set[str]]
|
||||
for records in runs[1:]:
|
||||
if set(records) != groups:
|
||||
raise TallyError(
|
||||
"the three runs do not cover the same"
|
||||
" groups: " + ", ".join(sorted(
|
||||
groups.symmetric_difference(
|
||||
set(records)))))
|
||||
rows: list[tuple[str, str, int]] = []
|
||||
group: str
|
||||
for group in sorted(groups):
|
||||
votes: dict[str, int] = {}
|
||||
for records in runs:
|
||||
keyword: str
|
||||
for keyword in records[group]:
|
||||
votes[keyword] = votes.get(keyword, 0) + 1
|
||||
rows.extend(
|
||||
(group, x, votes[x])
|
||||
for x in sorted(votes) if votes[x] >= cls.__MAJORITY)
|
||||
return rows
|
||||
|
||||
def __write_csv(self, rows: list[tuple[str, str, int]]) \
|
||||
-> None:
|
||||
"""Write the group table CSV file.
|
||||
|
||||
Writes an RFC 4180 CSV file, UTF-8, with CRLF line
|
||||
endings, carrying the header row ``Group,Keyword,Votes``
|
||||
and one row per settled (group, keyword) pair, in the row
|
||||
order. The parent directory is created when it does not
|
||||
exist.
|
||||
|
||||
:param rows: The settled rows, in the output order.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
"""
|
||||
self.__output_csv.parent.mkdir(parents=True, exist_ok=True)
|
||||
with open(self.__output_csv, "w", encoding="utf-8",
|
||||
newline="") as file:
|
||||
writer: Any = csv.writer(file)
|
||||
writer.writerow(self.__HEADER)
|
||||
writer.writerows(rows)
|
||||
|
||||
|
||||
def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
"""Parse the command-line arguments.
|
||||
|
||||
@@ -68,238 +303,34 @@ def parse_args(argv: list[str] | None) -> argparse.Namespace:
|
||||
"run_dir_3", type=Path,
|
||||
help="the third selection run's archive directory")
|
||||
parser.add_argument(
|
||||
"valid_keywords", type=Path,
|
||||
"output_csv", type=Path,
|
||||
help="the output CSV file")
|
||||
parser.add_argument(
|
||||
"--valid-keywords", type=Path, required=True,
|
||||
help="a plain text file of the allowed keywords, one per"
|
||||
" line")
|
||||
parser.add_argument(
|
||||
"output_csv", type=Path,
|
||||
help="the output CSV file, by convention"
|
||||
" results/groups.csv")
|
||||
return parser.parse_args(argv)
|
||||
|
||||
|
||||
def load_valid_keywords(path: Path) -> set[str]:
|
||||
"""Load the valid keyword list.
|
||||
|
||||
:param path: The plain text file of the allowed keywords, one
|
||||
per line.
|
||||
:return: The allowed keywords.
|
||||
:raises TallyError: When the file cannot be read or holds no
|
||||
keyword.
|
||||
"""
|
||||
text: str
|
||||
try:
|
||||
text = path.read_text(encoding="utf-8")
|
||||
except OSError as error:
|
||||
raise TallyError(str(error)) from error
|
||||
keywords: set[str] = {x.strip() for x in text.split("\n")
|
||||
if x.strip() != ""}
|
||||
if len(keywords) == 0:
|
||||
raise TallyError(f"{path}: no keywords")
|
||||
return keywords
|
||||
|
||||
|
||||
def load_run(run_dir: Path) -> dict[str, set[str]]:
|
||||
"""Load and validate the selection records of one run.
|
||||
|
||||
:param run_dir: The run's archive directory, containing
|
||||
``output.jsonl``.
|
||||
:return: The selected keywords of every group of the run,
|
||||
keyed by the group name, the duplicates within one
|
||||
record's selection collapsed.
|
||||
:raises TallyError: When the file cannot be read, a line is
|
||||
not a well-formed output record, a record is not a
|
||||
successful result, an ID lacks the ``group-`` prefix, a
|
||||
group has two records, or a "text" does not parse to a
|
||||
JSON array of strings.
|
||||
"""
|
||||
path: Path = run_dir / "output.jsonl"
|
||||
text: str
|
||||
try:
|
||||
text = path.read_text(encoding="utf-8")
|
||||
except OSError as error:
|
||||
raise TallyError(str(error)) from error
|
||||
records: dict[str, set[str]] = {}
|
||||
line: str
|
||||
for line in text.split("\n"):
|
||||
if line.strip() == "":
|
||||
continue
|
||||
record: Any
|
||||
try:
|
||||
record = json.loads(line)
|
||||
except json.JSONDecodeError as error:
|
||||
raise TallyError(
|
||||
f"{path}: malformed JSON: {error}") from error
|
||||
if not isinstance(record, dict) or "id" not in record:
|
||||
raise TallyError(
|
||||
f"{path}: record without \"id\": {line}")
|
||||
item_id: Any = record["id"]
|
||||
if "error" in record or "text" not in record:
|
||||
raise TallyError(
|
||||
f"{path}: id {item_id}: not a successful result")
|
||||
if not isinstance(item_id, str) \
|
||||
or not item_id.startswith(GROUP_ID_PREFIX) \
|
||||
or item_id == GROUP_ID_PREFIX:
|
||||
raise TallyError(
|
||||
f"{path}: id {item_id}: not in the"
|
||||
f" \"{GROUP_ID_PREFIX}<name>\" form")
|
||||
group: str = item_id[len(GROUP_ID_PREFIX):]
|
||||
if group in records:
|
||||
raise TallyError(
|
||||
f"{path}: id {item_id}: duplicate record")
|
||||
records[group] = _selection(
|
||||
record["text"], f"{path}: id {item_id}")
|
||||
if len(records) == 0:
|
||||
raise TallyError(f"{path}: no records")
|
||||
return records
|
||||
|
||||
|
||||
def _selection(text: Any, label: str) -> set[str]:
|
||||
"""Parse and validate the selected keywords of one record.
|
||||
|
||||
:param text: The "text" field of the record.
|
||||
:param label: The location of the record, for the error
|
||||
message.
|
||||
:return: The selected keywords, the duplicates collapsed.
|
||||
:raises TallyError: When the text does not parse to a JSON
|
||||
array of strings.
|
||||
"""
|
||||
if not isinstance(text, str):
|
||||
raise TallyError(f"{label}: \"text\" is not a string")
|
||||
selected: Any
|
||||
try:
|
||||
selected = json.loads(text)
|
||||
except json.JSONDecodeError as error:
|
||||
raise TallyError(
|
||||
f"{label}: \"text\" is malformed JSON:"
|
||||
f" {error}") from error
|
||||
if not isinstance(selected, list) \
|
||||
or not all(isinstance(x, str) for x in selected):
|
||||
raise TallyError(
|
||||
f"{label}: \"text\" does not parse to a JSON array"
|
||||
" of strings")
|
||||
return set(selected)
|
||||
|
||||
|
||||
def drop_invalid(runs: list[dict[str, set[str]]],
|
||||
run_dirs: list[Path],
|
||||
valid: set[str]) -> None:
|
||||
"""Drop the out-of-vocabulary selections of every run.
|
||||
|
||||
Every dropped occurrence is reported on standard error as an
|
||||
observable side effect.
|
||||
|
||||
:param runs: The runs' records, filtered in place.
|
||||
:param run_dirs: The run directories, for the messages.
|
||||
:param valid: The allowed keywords.
|
||||
:return: None.
|
||||
"""
|
||||
records: dict[str, set[str]]
|
||||
run_dir: Path
|
||||
for records, run_dir in zip(runs, run_dirs):
|
||||
group: str
|
||||
selected: set[str]
|
||||
for group, selected in records.items():
|
||||
keyword: str
|
||||
for keyword in sorted(selected - valid):
|
||||
print(
|
||||
f"note: {run_dir.name} group-{group}:"
|
||||
f" dropped out-of-vocabulary item"
|
||||
f" \"{keyword}\"", file=sys.stderr)
|
||||
records[group] = selected & valid
|
||||
|
||||
|
||||
def tally(runs: list[dict[str, set[str]]]) \
|
||||
-> list[tuple[str, str, int]]:
|
||||
"""Tally the keyword votes of the runs, group by group.
|
||||
|
||||
:param runs: The runs' records, all covering the same set of
|
||||
groups.
|
||||
:return: The settled rows, each the group name, the keyword,
|
||||
and the number of votes, ordered by the group name and
|
||||
then by the keyword, by Unicode code point.
|
||||
:raises TallyError: When the runs do not cover the same set
|
||||
of groups.
|
||||
"""
|
||||
groups: set[str] = set(runs[0])
|
||||
records: dict[str, set[str]]
|
||||
for records in runs[1:]:
|
||||
if set(records) != groups:
|
||||
raise TallyError(
|
||||
"the three runs do not cover the same groups: "
|
||||
+ ", ".join(sorted(
|
||||
groups.symmetric_difference(set(records)))))
|
||||
rows: list[tuple[str, str, int]] = []
|
||||
group: str
|
||||
for group in sorted(groups):
|
||||
votes: dict[str, int] = {}
|
||||
for records in runs:
|
||||
keyword: str
|
||||
for keyword in records[group]:
|
||||
votes[keyword] = votes.get(keyword, 0) + 1
|
||||
rows.extend(
|
||||
(group, x, votes[x])
|
||||
for x in sorted(votes) if votes[x] >= MAJORITY)
|
||||
return rows
|
||||
|
||||
|
||||
def write_csv(output_csv: Path,
|
||||
rows: list[tuple[str, str, int]]) -> None:
|
||||
"""Write the group table CSV file.
|
||||
|
||||
Writes an RFC 4180 CSV file, UTF-8, with CRLF line endings,
|
||||
carrying the header row ``Group,Keyword,Votes`` and one row
|
||||
per settled (group, keyword) pair, in the row order. The
|
||||
parent directory is created when it does not exist.
|
||||
|
||||
:param output_csv: The output CSV file.
|
||||
:param rows: The settled rows, in the output order.
|
||||
:return: None.
|
||||
:raises OSError: When the file cannot be written.
|
||||
"""
|
||||
output_csv.parent.mkdir(parents=True, exist_ok=True)
|
||||
with open(output_csv, "w", encoding="utf-8",
|
||||
newline="") as file:
|
||||
writer: Any = csv.writer(file)
|
||||
writer.writerow(HEADER)
|
||||
writer.writerows(rows)
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
"""Settle the code groups by a majority of the three runs.
|
||||
|
||||
Writes the final group table as the given CSV file, holding
|
||||
the header row ``Group,Keyword,Votes`` and one row per
|
||||
(group, keyword) pair at least two of the three runs select,
|
||||
ordered by the group name and then by the keyword. A
|
||||
selected item that is not in the valid keyword list casts no
|
||||
vote, each dropped occurrence reported on standard error.
|
||||
Nothing is written when the three archives do not cover the
|
||||
same groups, a record is not a successful result, an ID lacks
|
||||
the ``group-`` prefix, or a record's "text" does not parse to
|
||||
a JSON array of strings; the error message names what failed.
|
||||
|
||||
:param argv: The command-line arguments, or None for
|
||||
``sys.argv``.
|
||||
:return: The exit status: 0 on success, non-zero on failure.
|
||||
"""
|
||||
started: float = time.monotonic()
|
||||
args: argparse.Namespace = parse_args(argv)
|
||||
run_dirs: list[Path] = [
|
||||
args.run_dir_1, args.run_dir_2, args.run_dir_3]
|
||||
try:
|
||||
valid: set[str] = load_valid_keywords(args.valid_keywords)
|
||||
runs: list[dict[str, set[str]]] = [
|
||||
load_run(x) for x in run_dirs]
|
||||
drop_invalid(runs, run_dirs, valid)
|
||||
rows: list[tuple[str, str, int]] = tally(runs)
|
||||
write_csv(args.output_csv, rows)
|
||||
tallied: TalliedGroups = GroupTallier(
|
||||
args.run_dir_1, args.run_dir_2, args.run_dir_3,
|
||||
args.valid_keywords, args.output_csv).run()
|
||||
except (TallyError, OSError) as error:
|
||||
print(f"error: {error}", file=sys.stderr)
|
||||
return 1
|
||||
elapsed: str = format_duration(time.monotonic() - started)
|
||||
print(
|
||||
f"Done. Settled {len(rows)} codes across"
|
||||
f" {len(runs[0])} groups. {elapsed} elapsed.",
|
||||
f"Done. Settled {tallied.codes} codes across"
|
||||
f" {tallied.groups} groups. {elapsed} elapsed.",
|
||||
file=sys.stderr)
|
||||
return 0
|
||||
|
||||
@@ -14,12 +14,12 @@ class Settings(BaseSettings):
|
||||
"""The application name."""
|
||||
admin_email: str = "imacat@mail.imacat.idv.tw"
|
||||
"""The administrator email address."""
|
||||
SQLALCHEMY_DATABASE_URL: str
|
||||
SQLALCHEMY_DATABASE_URI: str
|
||||
"""The SQLAlchemy database URL."""
|
||||
ANTHROPIC_API_KEY: str
|
||||
"""The Anthropic API key."""
|
||||
|
||||
model_config = SettingsConfigDict(env_file=".env", extra="ignore")
|
||||
model_config = SettingsConfigDict(env_file=".env")
|
||||
"""The model configuration."""
|
||||
|
||||
|
||||
|
||||
@@ -7,10 +7,11 @@
|
||||
"""
|
||||
from functools import cached_property
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import sqlalchemy as sa
|
||||
from sqlalchemy.engine.interfaces import DBAPICursor, DBAPIConnection
|
||||
from sqlalchemy.orm import DeclarativeBase, sessionmaker, Session
|
||||
from sqlalchemy.pool import ConnectionPoolEntry
|
||||
|
||||
from .config import Settings, get_settings
|
||||
|
||||
@@ -29,7 +30,7 @@ class DataSource:
|
||||
:return: The database engine.
|
||||
"""
|
||||
settings: Settings = get_settings()
|
||||
return self.__create_engine(settings.SQLALCHEMY_DATABASE_URL)
|
||||
return self.__create_engine(settings.SQLALCHEMY_DATABASE_URI)
|
||||
|
||||
@cached_property
|
||||
def __session_local(self) -> sessionmaker:
|
||||
@@ -51,9 +52,6 @@ class DataSource:
|
||||
def __create_engine(cls, url: str) -> sa.Engine:
|
||||
"""Constructs and returns the database engine.
|
||||
|
||||
The foreign key enforcement is enabled on every connection
|
||||
of a SQLite engine.
|
||||
|
||||
:param url: The SQLAlchemy database URL.
|
||||
:return: The database engine.
|
||||
"""
|
||||
@@ -65,24 +63,10 @@ class DataSource:
|
||||
poolclass=sa.StaticPool)
|
||||
else:
|
||||
engine = sa.create_engine(url)
|
||||
if engine.url.get_backend_name() == "sqlite":
|
||||
sa.event.listen(engine, "connect",
|
||||
cls.__enable_sqlite_foreign_keys)
|
||||
if engine.dialect.name == "sqlite":
|
||||
cls.__enable_sqlite_foreign_keys(engine)
|
||||
return engine
|
||||
|
||||
@staticmethod
|
||||
def __enable_sqlite_foreign_keys(dbapi_connection: Any,
|
||||
_: Any) -> None:
|
||||
"""Enables the foreign key enforcement on a new connection.
|
||||
|
||||
:param dbapi_connection: The DBAPI connection.
|
||||
:param _: The connection record (unused).
|
||||
:return: None.
|
||||
"""
|
||||
cursor: Any = dbapi_connection.cursor()
|
||||
cursor.execute("PRAGMA foreign_keys=ON")
|
||||
cursor.close()
|
||||
|
||||
@staticmethod
|
||||
def __resolve_sqlite_relative_url(url: str) -> str:
|
||||
"""Resolves the SQLite relative URL to the instance folder.
|
||||
@@ -101,6 +85,31 @@ class DataSource:
|
||||
path = base / "instance" / path
|
||||
return f"sqlite:///{path}"
|
||||
|
||||
@staticmethod
|
||||
def __enable_sqlite_foreign_keys(engine: sa.Engine) -> None:
|
||||
"""Turns on the foreign key enforcement of SQLite.
|
||||
|
||||
The ``foreign_keys`` pragma is turned on for every
|
||||
connection of the engine, so that the ``ON DELETE``
|
||||
actions of the schema run.
|
||||
|
||||
:param engine: The SQLite database engine.
|
||||
:return: None.
|
||||
"""
|
||||
def on_connect(dbapi_connection: DBAPIConnection,
|
||||
_: ConnectionPoolEntry) -> None:
|
||||
"""Turns on the pragma on a new connection.
|
||||
|
||||
:param dbapi_connection: The DB-API connection.
|
||||
:param _: The connection record (unused).
|
||||
:return: None.
|
||||
"""
|
||||
cursor: DBAPICursor = dbapi_connection.cursor()
|
||||
cursor.execute("PRAGMA foreign_keys=ON")
|
||||
cursor.close()
|
||||
|
||||
sa.event.listen(engine, "connect", on_connect)
|
||||
|
||||
|
||||
ds: DataSource = DataSource()
|
||||
"""The data source."""
|
||||
|
||||
@@ -278,16 +278,18 @@ class TestBuildDB(unittest.TestCase):
|
||||
self.__annotations: Path = self.__dir / "annotations.csv"
|
||||
self.__write_chart(self.CHART_CSV)
|
||||
config.set_settings(config.Settings(
|
||||
SQLALCHEMY_DATABASE_URL="sqlite://",
|
||||
SQLALCHEMY_DATABASE_URI="sqlite://",
|
||||
ANTHROPIC_API_KEY="test-key"))
|
||||
self.__ds: DataSource = DataSource()
|
||||
self.addCleanup(self.__ds.engine.dispose)
|
||||
patchers: list[Any] = [
|
||||
mock.patch.object(build_db, "ds", self.__ds),
|
||||
mock.patch.object(
|
||||
build_db.SongImporter, "YEARS", [2016, 2017]),
|
||||
build_db.SongImporter, "_SongImporter__YEARS",
|
||||
[2016, 2017]),
|
||||
mock.patch.object(
|
||||
build_db.SongImporter, "RANKS_PER_YEAR", 2)]
|
||||
build_db.SongImporter,
|
||||
"_SongImporter__RANKS_PER_YEAR", 2)]
|
||||
for patcher in patchers:
|
||||
patcher.start()
|
||||
self.addCleanup(patcher.stop)
|
||||
@@ -1014,12 +1016,9 @@ class TestBuildDB(unittest.TestCase):
|
||||
"""Test that the coding CSV imports one row per song and
|
||||
keyword, storing the quote column verbatim."""
|
||||
self.__write_codings(self.CODINGS_CSV)
|
||||
status: int
|
||||
stderr: str
|
||||
status, stderr = self.__run_build(
|
||||
"--codings", str(self.__codings))
|
||||
self.assertEqual(status, 0)
|
||||
self.assertIn("3 codings", stderr)
|
||||
self.assertEqual(
|
||||
self.__run_build("--codings", str(self.__codings))[0],
|
||||
0)
|
||||
self.assertEqual(
|
||||
self.__stored_codings(),
|
||||
{("Hello", "longing"):
|
||||
@@ -1043,11 +1042,7 @@ class TestBuildDB(unittest.TestCase):
|
||||
def test_omitted_codings_leaves_table_empty(self) -> None:
|
||||
"""Test that an omitted coding option leaves no codings."""
|
||||
self.__write_codings(self.CODINGS_CSV)
|
||||
status: int
|
||||
stderr: str
|
||||
status, stderr = self.__run_build()
|
||||
self.assertEqual(status, 0)
|
||||
self.assertIn("0 codings", stderr)
|
||||
self.assertEqual(self.__run_build()[0], 0)
|
||||
self.assertEqual(self.__stored_codings(), {})
|
||||
|
||||
def test_codings_unknown_song_fails(self) -> None:
|
||||
@@ -1119,12 +1114,9 @@ class TestBuildDB(unittest.TestCase):
|
||||
"Song,Artist Credit,Keyword,Quote\n"
|
||||
"Shape of You,Ed Sheeran,attraction,I'm in love with"
|
||||
" your body\n")
|
||||
status: int
|
||||
stderr: str
|
||||
status, stderr = self.__run_build(
|
||||
"--codings", str(self.__codings))
|
||||
self.assertEqual(status, 0)
|
||||
self.assertIn("1 codings", stderr)
|
||||
self.assertEqual(
|
||||
self.__run_build("--codings", str(self.__codings))[0],
|
||||
0)
|
||||
self.assertEqual(
|
||||
self.__stored_codings(),
|
||||
{("Shape of You", "attraction"):
|
||||
@@ -1159,12 +1151,8 @@ class TestBuildDB(unittest.TestCase):
|
||||
"""Test that the group CSV imports one row per group and
|
||||
keyword, the votes stored as integers."""
|
||||
self.__write_groups(self.GROUPS_CSV)
|
||||
status: int
|
||||
stderr: str
|
||||
status, stderr = self.__run_build(
|
||||
"--groups", str(self.__groups))
|
||||
self.assertEqual(status, 0)
|
||||
self.assertIn("3 group members", stderr)
|
||||
self.assertEqual(
|
||||
self.__run_build("--groups", str(self.__groups))[0], 0)
|
||||
self.assertEqual(
|
||||
self.__stored_groups(),
|
||||
{("masculine", "dominance-and-power"): 3,
|
||||
@@ -1180,12 +1168,8 @@ class TestBuildDB(unittest.TestCase):
|
||||
self.__write_groups(
|
||||
"Group,Keyword,Votes\n"
|
||||
"vulnerable,longing-and-loss,3\n")
|
||||
status: int
|
||||
stderr: str
|
||||
status, stderr = self.__run_build(
|
||||
"--groups", str(self.__groups))
|
||||
self.assertEqual(status, 0)
|
||||
self.assertIn("1 group members", stderr)
|
||||
self.assertEqual(
|
||||
self.__run_build("--groups", str(self.__groups))[0], 0)
|
||||
self.assertEqual(
|
||||
self.__stored_groups(),
|
||||
{("vulnerable", "longing-and-loss"): 3})
|
||||
@@ -1407,14 +1391,11 @@ class TestBuildDB(unittest.TestCase):
|
||||
together, reflected in the final counts message."""
|
||||
self.__write_patterns(self.PATTERNS_CSV)
|
||||
self.__write_annotations(self.ANNOTATIONS_CSV)
|
||||
status: int
|
||||
stderr: str
|
||||
status, stderr = self.__run_build(
|
||||
"--patterns", str(self.__patterns),
|
||||
"--annotations", str(self.__annotations))
|
||||
self.assertEqual(status, 0)
|
||||
self.assertIn("2 patterns", stderr)
|
||||
self.assertIn("2 annotations", stderr)
|
||||
self.assertEqual(
|
||||
self.__run_build(
|
||||
"--patterns", str(self.__patterns),
|
||||
"--annotations", str(self.__annotations))[0],
|
||||
0)
|
||||
self.assertEqual(
|
||||
self.__stored_patterns(),
|
||||
{"M1": ("male", "Dominance",
|
||||
@@ -1432,12 +1413,7 @@ class TestBuildDB(unittest.TestCase):
|
||||
annotation tables empty."""
|
||||
self.__write_patterns(self.PATTERNS_CSV)
|
||||
self.__write_annotations(self.ANNOTATIONS_CSV)
|
||||
status: int
|
||||
stderr: str
|
||||
status, stderr = self.__run_build()
|
||||
self.assertEqual(status, 0)
|
||||
self.assertIn("0 patterns", stderr)
|
||||
self.assertIn("0 annotations", stderr)
|
||||
self.assertEqual(self.__run_build()[0], 0)
|
||||
self.assertEqual(self.__stored_patterns(), {})
|
||||
self.assertEqual(self.__stored_annotations(), {})
|
||||
|
||||
|
||||
@@ -36,7 +36,7 @@ class TestExportLlmInput(unittest.TestCase):
|
||||
self.__dir: Path = Path(tmp.name)
|
||||
self.__output: Path = self.__dir / "llm-input.jsonl"
|
||||
config.set_settings(config.Settings(
|
||||
SQLALCHEMY_DATABASE_URL="sqlite://",
|
||||
SQLALCHEMY_DATABASE_URI="sqlite://",
|
||||
ANTHROPIC_API_KEY="test-key"))
|
||||
self.__ds: DataSource = DataSource()
|
||||
self.addCleanup(self.__ds.engine.dispose)
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user