馃洜锔廩v0.4.1] Release Note: MTMD Template Compatibility, Vision/Video Input Improvements, and Expanded GGML Backend Bindings #182
JamePeng
announced in
Announcements
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
[0.4.1] Release Note: MTMD Template Compatibility, Vision/Video Input Improvements, and Expanded GGML Backend Bindings
This release focuses on MTMD compatibility, runtime reliability, and backend coverage.
Generic MTMD handling is expanded for more vision templates and media input formats, including Muse Glimmer, Qwen3-VL, and Qwen3.5 video inputs. It also fixes model-template resolution, standard Jinja helpers, media evaluation diagnostics, and the MTMD decoder-position ABI layout.
On the backend side,
ggml-backend.hbindings are significantly expanded to cover device, buffer, tensor, graph, event, and scheduler APIs, while the README and Wiki have been updated to reflect the current bundledllama.cppbehavior and build options.feat(mtmd): expand vision chat template media support
feat(ggml): expand backend API bindings
fix(mtmd): load and compile the model template for generic chat handlers
treating the precompiled fallback as an explicit template.
normalization, and extra template arguments use the same template.
fix(mtmd): provide actionable guidance for media evaluation failures
fix(mtmd): register standard Jinja chat template helpers
fix(mtmd): align decoder position struct with native ABI
docs(readme): align feature and build guidance with current behavior
guidance against the bundled llama.cpp implementation.
decoding boundaries, and the limits of cross-request state reuse.
describe batching, normalization, output ownership, and generation-state
cleanup.
current development test setup, and fix malformed Markdown fences and
project links.
docs(wiki): align backend options with bundled llama.cpp
obsolete cuBLAS compute variables with GGML_CUDA_CUBLAS_COMPUTE_TYPE.
Level Zero and oneDNN variable names.
WebGPU, and Hexagon build options with their validation boundaries.
guarantee that all operations avoid an available GPU backend.
Thanks to @KLL535 for the bug report and testing feedback in issue #180.
More information:
5c83af7...3e99f13
-- JamePeng
All reactions