Skip to content
#

consumer-gpu

Here are 51 public repositories matching this topic...

SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35B-A3B FP8 ~240 tok/s, 27B-int4 hybrid GDN+Mamba, Gemma4 26B/31B AWQ, 256K ctx. 321 patches: TurboQuant k8v4 KV, MTP/DFlash spec-decode, FULL cudagraph, hybrid GDN. vLLM pin dev424 + Control Center GUI.

  • Updated Sep 19, 2026
  • Python

llama.cpp fork for significantly improved performance on Ampere (especially RTX 3090 / 3090 Ti): TurboQuant KV cache, MTP speculative decoding with a 64K draft-vocabulary shortlist, custom SM86 + Qwen kernels. 90 tok/s over a 100K-token generation at temperature 1.

  • Updated Sep 19, 2026
  • C++

Verified AI infrastructure for regulated deployment. UltraCompress (our wedge): near-lossless 5-bit compression with SHA-256-reproducible reconstruction - prove the model in production is the one you validated. 23 architectures (0.6B-405B), Hermes-3-405B @ 1.0066x. OpenAI-compatible API. pip install ultracompress

  • Updated Jul 27, 2026
  • Python

MoziAI-35B: Local open-source financial AI LLM. 35B MoE model compressed to 15.5GB via MoziSmartBit quantization. 256K context, vision, tool calling, 140+ tok/s on consumer GPU. Based on Qwen3.6-35B-A3B.

  • Updated Sep 12, 2026
  • Jinja

Unofficial FreeToken fork: on one RTX 3060 12 GB, a 35B MoE at 250k of context or gpt-oss-120b; Flash-Next 125B on two. Half the RAM, image input. Runs on Turing: RTX 2060, RTX 20 series, sm_75.

  • Updated Sep 19, 2026
  • Python

Add this topic to your repo

To associate your repository with the consumer-gpu topic, visit your repo's landing page and select "manage topics."

Learn more