Lynn GitHub 镜像仓 · Primary repository: https://github.com/MerkyorLynn/Lynn · Downloads: https://download.merkyorlynn.com/download.html
-
Updated
Sep 11, 2026 - TypeScript
Lynn GitHub 镜像仓 · Primary repository: https://github.com/MerkyorLynn/Lynn · Downloads: https://download.merkyorlynn.com/download.html
Pokedex benchmark for LLMs
Benchmark of local LLMs (Qwen 3.6, Gemma 4) vs ChatGPT and Gemini on coding tasks. Apple M5, 32GB. Methodology, raw outputs, judge prompts, scores, and charts.
Saotri Bench — coding benchmark for evaluating LLM agents on multi-phase programming tasks with hidden requirements.
31 local and hosted LLMs tested with one difficult falling-block browser game prom31 local and hosted LLMs tested with one difficult falling-block browser game prompt. Runnable outputs and hands-on findings.pt. Runnable outputs and hands-on findings.
面向 AI IDE / 编码 Agent 的编程能力基准测试框架,内置防背题三防线:公私用例分离、泛化差距检测、参数化生成
Raw logs of Claude Code running on local Qwen3.5-27B (llama.cpp). Builds a Python todo app with 50 tests. Real-world performance data: 30 min, cache thrashing, 38 t/s generation.
Benchmark and tune every locally installed Ollama model on your own GPUs. Finds the best model per GPU setup plus its optimal temperature and context, with a no-spill VRAM guard, coding + agentic test suites, and a single-file HTML report. 100% local: no cloud, no API keys, no model judges.
To associate your repository with the coding-benchmark topic, visit your repo's landing page and select "manage topics."