Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Ā 

History

144 Commits
Ā 
Ā 
Ā 
Ā 
Ā 
Ā 

Repository files navigation

šŸ‘Øā€šŸ’» Awesome Code Benchmark

Awesome PRs Welcome

A comprehensive code domain benchmark review of LLM researches.

Oryx Video-ChatGPT

Table of Contents

Taxonomy

This list organizes code benchmarks by primary capability and software-engineering workflow. Each benchmark entry includes compact metadata for task type, granularity, interaction pattern, and evaluation method.

Dimension Values
Granularity Function, API, Query / Database, File, Project, UI, Repository, Workflow
Interaction Single-turn, Multi-turn, Agentic, Async, Multi-agent
Evaluation Unit Tests, Execution, Performance, Human, LLM-as-Judge, Security Exploit, Economic
Environment None, Sandbox, Terminal, Browser, IDE, CI
Freshness Static, Dynamic, Held-out, Contamination-resistant

Surveys

  1. Software Development Life Cycle Perspective A Survey of Benchmarks for Code Large Language Models and Agents from Xi’an Jiaotong University

  2. Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks from Zhejiang University

  3. A Survey on Large Language Model Benchmarks from Shenzhen Key Laboratory for High Performance Data Mining

šŸš€ Benchmark Categories

Repository & Agentic Software Engineering

Program Repair, Testing & Debugging

Security, Reliability & Robustness

Code Understanding, Search & Review

Performance Optimization

Frontend, UI & Visual-Interactive Development

Code Generation & Completion

Releases

Packages

Contributors