Grok Bot template · Generative AI and LLMs
Evaluation Bigcode Evaluation Harness
Evaluates code generation models on 15+ benchmarks with pass@k metrics.
What it can do
The skills built into this template. Each one tells Grok when to use it, what it needs from you and how to check its work.
- Run standard code benchmark evaluation
- Run multi-language evaluation with MultiPL-E
- Evaluate instruction-tuned models
- Compare multiple models on the same benchmarks
- List available tasks
- Configure generation parameters for quantized or custom models
Apps it works with
Connect these in Grok for the best results. It also works without them: you paste the information in.
HuggingFace model repository accessDocker (for MultiPL-E evaluation)
The full template
For members
The complete Evaluation Bigcode Evaluation Harness template: its identity, every skill step by step, its limits and its first-run questions, ready to paste into a new Grok Bot. Members get it, and every other template here.
Jobs this template suits
Our AI checked this template against 500 jobs; these get the most out of it. Each job links to its learning path.