An LLM-as-judge harness that scores AI-generated campaign phone scripts against a weighted quality rubric with a real Haiku-vs-Opus benchmark.
python ai python3 prompt-engineering llm-eval evals llm-evaluation llm-as-judge llm-as-a-judge political-tech political-technology
-
Updated
Jun 30, 2026 - Python