ice-score
[EACL 2024] ICE-Score: Instructing Large Language Models to Evaluate Code
Cyber
[EACL 2024] ICE-Score: Instructing Large Language Models to Evaluate Code
A list of LLM benchmark frameworks.
Source Code Data Augmentation for Deep Learning: A Survey.
PyArmadillo: an alternative approach to linear algebra in Python
X-Repo2Run: Configuraing Multilingual Docker Environment via Code Agent
Harbor is a framework for running agent evaluations and creating and using RL environments.
Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
A benchmark for LLMs on complicated tasks in the terminal
test-runner