CodeScaleBench: Testing coding agents on large codebases and multi-repo software engineering tasks
The initial findings from CodeScaleBench, a new benchmark designed to evaluate coding agents against the true complexity of enterprise software development, including large codebases and multi-repository tasks.
Read more