Skip to content
#

coding-benchmark

Here are 8 public repositories matching this topic...

Raw logs of Claude Code running on local Qwen3.5-27B (llama.cpp). Builds a Python todo app with 50 tests. Real-world performance data: 30 min, cache thrashing, 38 t/s generation.

  • Updated Apr 14, 2026
  • Python

Benchmark and tune every locally installed Ollama model on your own GPUs. Finds the best model per GPU setup plus its optimal temperature and context, with a no-spill VRAM guard, coding + agentic test suites, and a single-file HTML report. 100% local: no cloud, no API keys, no model judges.

  • Updated Sep 17, 2026
  • Python

Add this topic to your repo

To associate your repository with the coding-benchmark topic, visit your repo's landing page and select "manage topics."

Learn more