Tag
SWE-Bench
2 articles

Development
SWE-Touch: The Edit You Make While an Agent Still Runs
SWE-Touch's 3 August 2026 benchmark found resolve rates fall 7.7 points on average when a user edits code an agent is still working on, and the agent often finishes anyway.

Development
FrontierCode: The Benchmark That Asks Whether AI Code Is Ready to Merge
A new benchmark built with more than 20 open-source maintainers deflates the record-breaking numbers behind coding agents: even the best model clears only 13% of the hardest tasks.