Files
windmill/ai_evals/package.json
T
hugocasaandClaude Opus 5.5 e117871bec weekly ai evals on current models, and claude 5.5/gpt-6 support (#11409)
* feat: run ai evals weekly on current models and post results to a dashboard

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat: add current flagship models, a reasoning flag and claude 5.5 defaults

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: never send a reasoning disable claude 5.5 or gpt-6-astra reject, and treat gpt-6 as a reasoning model

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: address review on gpt-6 support, chat completions tools and model metadata

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: leave tiered gpt-6 unpriced and drop the off sentinel on gpt-5 and o-series

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: point the ai_evals readme at the model registry instead of copying it

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor: encode the reasoning rules as per-family maps with a shared parity fixture

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: keep the chat completions tools rule open-ended past gpt-5.6 and scope the parity fixture

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 13:10:24 +02:00

22 lines
587 B
JSON

{
"name": "windmill-ai-evals",
"private": true,
"type": "module",
"scripts": {
"cli": "bun cli/index.ts",
"typecheck": "tsc -p tsconfig.json",
"test:frontend-graph": "cd ../frontend && node_modules/.bin/vitest run --project server --config ../ai_evals/adapters/frontend/vitest.unit.config.ts"
},
"dependencies": {
"@anthropic-ai/claude-agent-sdk": "^0.3.284",
"@anthropic-ai/sdk": "^0.129.0",
"commander": "^14.0.3",
"openai": "^6.9.1",
"yaml": "^2.8.3"
},
"devDependencies": {
"@types/bun": "latest",
"typescript": "^5.0.0"
}
}