Skip to content

Support dsv4 model - #1593

Open
WANDY666 wants to merge 229 commits into
mainfrom
support_dsv4_model
Open

WANDY666 wants to merge 229 commits into
mainfrom
support_dsv4_model

Conversation

@WANDY666

@WANDY666 WANDY666 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Add DeepSeek-V4 inference support, including the base model, MTP and DSpark speculative decoding, vision inputs, prefill/decode disaggregation, and GPU/CPU/disk KV caching.

The implementation handles DeepSeek-V4’s compressed-attention layout explicitly: compressed history can be shared between requests, while sliding-window KV and continuation state remain request-owned.

Model Architecture and Inference

  • Add model registration, checkpoint configuration handling, and weight loading for DeepSeek-V4.
  • Implement pre/post layers, transformer layers, manifold-constrained hyper-connections (mHC), and attention compressors.
  • Support sliding-window attention together with C4 and C128 compressed history.
  • Integrate the packed FP8 KV layout with FlashMLA sparse attention, including indexer KV storage and sparse index construction.
  • Implement the model-specific routing/top-k and clamped SwiGLU behavior through the existing MoE execution paths.
  • Add fused normalization/RoPE, packed KV writers, and compression/indexing kernels.
  • Support configuration normalization for both older checkpoint fields and the newer layer-type/compression-rate schema.

KV Layout

DeepSeek-V4 uses separate physical layouts for its attention components:

Component Physical page layout Ownership
SWA KV 128 slots per page Request-private
C4 compressed KV 64 slots per page Shareable compressed history
C128 compressed KV 2 slots per page Shareable compressed history
Compressor/indexer continuation state Checkpoint-specific layout Request-private runtime state

Compressed history is addressed through the existing request-to-token table using 256-token allocation pages. Closed compression groups derive their physical locations from the corresponding group-ending token slots.

Requests sharing a cached prefix reuse the compressed history but restore independent SWA and continuation state. Prefill retains the SWA pages needed for the entire current chunk, rather than recycling them before attention finishes.

Vision Support

  • Add DeepSeek-V4 image preprocessing, visual model execution, and visual embedding integration.
  • Support model-specific image-token encoding and multimodal message construction.
  • Integrate vision inputs with the existing visual-server execution path.
  • Preserve image-block boundaries during chunked prefill.
  • Adjust CPU-cache resume boundaries when a matched checkpoint would resume inside an image block.

This allows vision requests to use the same model inference, prefix-cache, and PD paths while keeping image blocks atomic where required.

Tokenizer and Response Protocol

  • Add the DeepSeek-V4 tokenizer and message encoder.
  • Support the model’s chat-template and thinking/reasoning request semantics.
  • Register DeepSeek-V4 reasoning-output parsing.
  • Support DSML tool-call parsing, including streaming output.

Prefill/Decode Disaggregation and DP Cache Transfer

Add model-specific pack/unpack and copy operations for DeepSeek-V4 caches.

PD Transfer

  • Transfer compressed history separately from the final continuation block.
  • Include SWA, compressor state, and indexer continuation in the final block.
  • Support arbitrary prompt endpoints, including partially completed compression groups.
  • Integrate the packed layout with both NCCL and NIXL transporters.
  • Support different transfer byte lengths for history-only blocks and full continuation blocks.
  • Keep actual transfer lengths separate from the request’s page-aligned reserved capacity.

P/D nodes must run compatible versions because the DeepSeek-V4 continuation layout is part of the transfer contract.

Cross-DP Prefix Reuse

  • Copy compressed history and restore request-private continuation state when fetching a cached prefix from another DP rank.
  • Preserve the distinction between shared history references and request-owned SWA resources.

CPU and Disk KV Cache

Support packed DeepSeek-V4 checkpoint pages through the existing multi-level cache pipeline.

Each CPU checkpoint contains compressed history and the tail state required to resume inference. CPU continuation checkpoints retain the final 256-token SWA region and the relevant C4/indexer state; aligned checkpoint boundaries close the C128 compression group.

  • Save checkpoints incrementally as prefill reaches cache-page boundaries.
  • Pack GPU cache data into staging buffers before asynchronous CPU storage.
  • Track packing completion separately from CPU-store completion so source pages can be safely reused.
  • Restore missing compressed history and request-private continuation state on cache hits.
  • Register restored checkpoints for subsequent GPU radix-cache reuse.
  • Retain matched CPU-page references until asynchronous load/store work completes.
  • Publish disk-cache entries only after the complete prefix group is ready, preserving cumulative hash-chain ordering.
  • Skip CPU-cache loading for prompt-logprobs requests, since skipped prompt computation cannot produce the requested log probabilities.

Validation

Regression coverage includes:

  • Packed SWA/C4/C128 indexing and logical compression boundaries.
  • Prefix-cache hits, request forks, checkpoint restoration, pause/abort, eviction, and resource reclamation.
  • Repeated speculative writes after rejected proposals.
  • CUDA Graph replay with updated request-private page mappings.
  • MTP acceptance-length synchronization and CPU/GPU metadata consistency.
  • Empty batches and uneven DP microbatches.
  • Vision-block boundaries during prefill and cache restoration.
  • Incremental CPU checkpoint storage with source-page release/reuse.
  • CPU/PD pack-unpack and two-GPU checkpoint transfer.

Related test batches completed with 133 passing startup/cache/vision tests and 107 passing speculative-decoding and associated regression tests. These results cover the targeted test suites, not a complete end-to-end accuracy or performance evaluation.

WANDY666 and others added 30 commits June 3, 2026 09:20
Root cause of the historical cudagraph accuracy drop (gsm8k 0.96 -> 0.74,
coherent-but-runaway generations; same 0.75 the pre-v5 fullslot_decode
experiments worked around): _capture_decode warms up via copy.copy(infer_state),
which SHARES decode_att_state. FlashMLASchedMeta is lazily planned at the first
kernel call and written back onto that shared state, so the warmup pass locks a
schedule planned for the dummy batch (seq=2); the capture pass then binds those
stale scheduler tensors and every replay runs real requests with a tile schedule
planned for near-empty kv (systematically under-read attention).

Fix: reset_sched_meta_for_capture() hook on the nsa decode att state, invoked in
both capture paths after warmup, so planning happens INSIDE the captured region
and re-plans on every replay from live tensors.

Validation (tp4, H200, prompt cache on): batch-1 greedy decode is now
character-identical to eager; per-layer probe shows embed+swa layers bitwise
equal under replay, benign rounding-class deltas only in compress layers,
argmax unchanged. gsm8k 100q/128: cold 0.960/111s, warm 0.960/23.3s 100% hits
(eager: 0.95-0.97, cold 141s / warm 50s). Batch-1 decode 20.4ms/token vs 142ms
eager. 41/41 unit tests green. Codex review GO (incl. overlap-path symmetry).

launch.sh: drop --disable_cudagraph, derive PYTHONPATH from the script dir
(hardcoded tree path made a worktree launch silently serve main-tree code).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Graph-sandwich prefill (graphs capture dense ops only; attention/compressor
run eagerly between segments) was already in-tree; enabling it exposed that
HOLD-pad rows read the racing HOLD slot, making their hiddens nondeterministic
and perturbing real rows via MoE expert batching (ulp-level, amplified
~1.9x/layer). Zero the pad rows' attention output.

Residual greedy-trajectory divergence vs eager equals the fp4 marlin MoE
kernel's own run-to-run reduction-order noise (eager-vs-eager control: 0/4
match), accepted statistically: gsm8k 100q cold 0.980/115.5s warm 0.960/25.9s
(eager-baseline parity); batch-1 TTFT 1.86x at 46 tokens.
Keep the selected DSV4 model stack in one squashed change, with excluded commits documented in revert.md. Preserve the MTP hidden input buffer owned by decode CUDA Graph capture while retaining eager-mode cleanup.
# Conflicts:
#	lightllm/common/basemodel/layer_weights/meta_weights/fused_moe/impl/deepgemm_impl.py
#	lightllm/common/basemodel/layer_weights/meta_weights/fused_moe/impl/marlin_impl.py
#	lightllm/common/basemodel/layer_weights/meta_weights/fused_moe/impl/triton_impl.py
#	lightllm/common/basemodel/triton_kernel/fused_moe/grouped_fused_moe_ep.py
#	lightllm/common/basemodel/triton_kernel/fused_moe/moe_silu_and_mul.py
#	lightllm/common/basemodel/triton_kernel/fused_moe/moe_silu_and_mul_mix_quant_ep.py
#	lightllm/common/req_manager/__init__.py
#	lightllm/common/state_cache_manager/__init__.py
This reverts commit fc8b7ec. Preserve the hc_post import required by the subsequent MTP hidden preparation path.

Record the rollback in revert.md.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants