Initialize project structure, workflow rules, and research plan

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-30 18:51:10 +08:00
co-authored by Claude Fable 5
commit 9a42b463a3
9 changed files with 1548 additions and 0 deletions
+36
View File
@@ -0,0 +1,36 @@
# Project Conventions
## Analysis pipeline
- LLM analysis runs via Python scripts calling the Anthropic
Messages API: model `claude-sonnet-4-6`, `temperature=0`,
thinking disabled, Batch API where possible.
- Prompt definition files live in `prompts/<task>_v<N>.md` and
are passed verbatim as the system prompt.
- Every LLM step runs the same definition file twice, then a
separate arbitration step reconciles the two outputs
("2 runs + 1 arbitration"). If arbitration output is
unexpected, revise the definition file and repeat the whole
cycle; never patch results by hand.
- Every execution is archived self-contained under
`runs/<phase>/<date>-<prompt-version>/`: prompt snapshot,
raw outputs of both runs, arbitration output, and `meta.json`
(model ID, parameters, timestamps, batch IDs).
- Scripts read the API key from the `ANTHROPIC_API_KEY`
environment variable (`.env`, gitignored).
## Data rules
- `data/yearend_hot100_2016_2025.csv` is the only hand-placed
raw file; everything else in `data/` is script-derived.
- Full lyrics are copyrighted: they stay in `data/lyrics/`
(gitignored) and must never be committed or reproduced in
full anywhere in the repo.
## Documents
- `results/` holds final arbitrated tables (what the paper
cites); `runs/` holds raw audit records. The paper cites
`results/` only.
- Any change to a definition file, the codebook, or the plan is
recorded in `docs/decision_log.md` with date and reason.