Files
pop-fem-audit/CLAUDE.md
T

1.9 KiB

Project Conventions

Analysis pipeline

  • LLM analysis runs via Python scripts calling the Anthropic Messages API: model claude-sonnet-4-6, temperature=0, thinking disabled, Batch API where possible.
  • Prompt definition files live in prompts/<track>-<step>-<task>.md (e.g. 01-01-tag.md; no version suffix -- versions live in git history) and are passed verbatim as the system prompt.
  • LLM steps whose outputs are item-by-item comparable (convergence, coding) run the same definition file twice, then a separate arbitration step settles only the script-computed disagreements ("2 runs + 1 arbitration"). Free-generation steps run twice and both outputs are pooled, unarbitrated. If arbitration output is unexpected, revise the definition file and repeat that cycle; never patch results by hand.
  • The current execution of each step is archived self-contained under runs/<definition-file>/: prompt snapshot, raw outputs of both runs, arbitration output, and meta.json (model ID, parameters, timestamps, batch IDs). A rerun replaces the directory; superseded runs live in git history.
  • Scripts read the API key from the ANTHROPIC_API_KEY environment variable (.env, gitignored).

Data rules

  • data/source/ holds the immutable hand-placed raw files; data/captures/ is written only by the fetch commands and the private import script; data/manual/ is written only by the user's own hand; data/derived/ is written only by the build-db subcommand.
  • Full lyrics are copyrighted: they stay in data/captures/lyrics/ (gitignored) and must never be committed or reproduced in full anywhere in the repo.

Documents

  • results/ holds final arbitrated tables (what the paper cites); runs/ holds raw audit records. The paper cites results/ only.
  • Any change to a definition file, the codebook, or the plan is recorded in docs/decision-log.md with date and reason.