Skip to content

Commit bc2a4f0

Browse files
committed
✨ Add SWE_Bench_GPT_5.5_Codex
1 parent a58e0be commit bc2a4f0

2 files changed

Lines changed: 675 additions & 0 deletions

File tree

README.md

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -34,6 +34,7 @@ Reproducible benchmarks for coding agents and models using Harbor
3434
| Model | Harness | Score | Cost |
3535
| ------------------------------ | ----------- | ---------------------------------------------------------------- | --------------- |
3636
| Opus 4.8 | Claude Code | [86.8%](./benchmarks/SWE_Bench_Opus_4.8_Claude_Code.md) | $395 |
37+
| GPT 5.5 | Codex | [79.8%](./benchmarks/SWE_Bench_Opus_4.8_Claude_Code.md) | $443 |
3738
| Sonnet 4.6 | Claude Code | [79.6%](https://www.anthropic.com/news/claude-sonnet-4-6) | N/A |
3839
| RedHatAI/Qwen3.6-35B-A3B-NVFP4 | Pi | [65.0%](./benchmarks/SWE_Bench_Qwen3.6_35b_NVFP4_Pi.md) | $51<sup>†</sup> |
3940
| RedHatAI/Qwen3.6-35B-A3B-NVFP4 | Qwen Code | [63.8%](./benchmarks/SWE_Bench_Qwen3.6_35b_NVFP4_Qwen_Code.md) | $37<sup>†</sup> |

0 commit comments

Comments
 (0)