Skip to content

fix(evaluator): retain valid feedback from partial judge ensembles - #511

Merged
codelion merged 1 commit into
algorithmicsuperintelligence:mainfrom
liuzhengyang699:fix/llm-judge-partial-scores
Oct 11, 2026
Merged

codelion merged 1 commit into
algorithmicsuperintelligence:mainfrom
liuzhengyang699:fix/llm-judge-partial-scores

Conversation

@liuzhengyang699

Copy link
Copy Markdown
Contributor

One malformed LLM judge response currently makes _llm_evaluate discard every judge's feedback, including valid responses already parsed. A missing metric is also implicitly scored as zero: with weights 0.25/0.75, a clarity score of 0.6 reported only by the first judge becomes 0.15.

Parse each judge independently and compute each metric's weighted mean over the judges that supplied it. Invalid JSON and non-object responses are skipped with a warning. Zero-weight judges do not introduce zero-valued metrics. If all responses are invalid, the measured program fitness remains unchanged.

The regression tests cover invalid responses in either position, missing metrics, zero-weight judges, and the resulting combined fitness through evaluate_program. They use mocked judge responses and a temporary evaluator, with no model calls.

Validation: 22 targeted tests passed. The official OPENAI_API_KEY=test-key-for-unit-tests python -m unittest discover tests suite passed 640 tests (2 skipped). The repository's isort/Black pre-commit hooks and git diff --check passed.

@codelion
codelion merged commit 30d9752 into algorithmicsuperintelligence:main Oct 11, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants