-
-
Notifications
You must be signed in to change notification settings - Fork 21k
Pull requests: vllm-project/vllm
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[Parser] Move structural-tag capabilities into parser engines
deepseek
Related to DeepSeek models
glm
inkling
kimi
minimax
mistral
Related to Mistral models
qwen
Related to Qwen models
tool-calling
#53099
opened Aug 20, 2026 by
chaunceyjiang
Collaborator
Loading…
4 tasks
[ROCm][Quantization][MOE] Enable fused shared experts for block-quantized FP8
quantization
rocm
Related to AMD ROCm
#53097
opened Aug 20, 2026 by
xuebwang-amd
Contributor
Loading…
4 tasks
fix(dflash-speculator): guard capture() log on needs_capture() (#53031)
dflash
mrv2
Model Runner V2 specific
speculative-decoding
#53096
opened Aug 20, 2026 by
sharonyao1127
Loading…
[ROCm][Perf] Fuse DSA indexer QK preprocessing with AITER
deepseek
Related to DeepSeek models
performance
Performance-related issues
rocm
Related to AMD ROCm
#53094
opened Aug 20, 2026 by
sumin-hong
Loading…
4 tasks done
[MM] Remove text components from ProcessorInputs
build-docs
cohere
Related to Cohere models
deepseek
Related to DeepSeek models
documentation
Improvements or additions to documentation
glm
inkling
k3
kimi
llama
Related to Llama models
minimax
mistral
Related to Mistral models
multi-modality
Related to multi-modality (#4194)
qwen
Related to Qwen models
ready
ONLY add when PR is ready to merge/full CI is needed
#53093
opened Aug 20, 2026 by
DarkLight1337
Member
Loading…
4 tasks
[Perf][Spec Decode] Fuse target temperature in rejection sampler
mrv2
Model Runner V2 specific
performance
Performance-related issues
speculative-decoding
#53090
opened Aug 20, 2026 by
positive666
Contributor
•
Draft
upgrade tpu-inference to v0.27.0
ci/build
#53088
opened Aug 20, 2026 by
meiyeh123
Contributor
Loading…
[Bugfix][KV Offload] Bound primary HIT_PENDING waits
bug
Something isn't working
documentation
Improvements or additions to documentation
#53087
opened Aug 20, 2026 by
thunguo
Loading…
4 tasks done
[Core] Make sleep a pure memory-state transition
documentation
Improvements or additions to documentation
frontend
multi-modality
Related to multi-modality (#4194)
performance
Performance-related issues
[Perf][MoE] Share FlashInfer B12x MoE workspaces across layers
nvidia
#53081
opened Aug 20, 2026 by
lucifer1004
Contributor
•
Draft
[Bugfix] Separate target and draft scheduling budgets
bug
Something isn't working
scheduler
#53080
opened Aug 20, 2026 by
slippersss
Contributor
Loading…
4 tasks
[Bugfix][Core] Ignore stale KV-transfer completions for aborted requests
bug
Something isn't working
scheduler
#53079
opened Aug 20, 2026 by
lucifer1004
Contributor
Loading…
[Bugfix][KV Connector] Mooncake: heterogeneous-TP support for hybrid GDN/Mamba models
bug
Something isn't working
kv-connector
#53078
opened Aug 20, 2026 by
lucifer1004
Contributor
Loading…
[XPU] Add VLLM_XPU_DEVICE_OFFSET to place a worker on a non-zero card without masking
documentation
Improvements or additions to documentation
intel-gpu
Related to Intel GPU
[Benchmark] Add asyncio-based multi_turn benchmark v2 to fix deadlock
performance
Performance-related issues
#53075
opened Aug 20, 2026 by
Potterluo
Loading…
3 tasks done
[Bugfix][Core] Isolate hidden-state cache from DeepSeek-V4 MLA groups
bug
Something isn't working
deepseek
Related to DeepSeek models
DSv4
kv-cache-manager
kv-connector
#53074
opened Aug 20, 2026 by
my0901
Loading…
[Bugfix][KV Offload] Decouple shared-region creator ownership
bug
Something isn't working
#53073
opened Aug 20, 2026 by
Alex-ai-future
Contributor
Loading…
[Kernel][Perf] Support DSpark K=8 in fused GDN MTP decode
dflash
#53070
opened Aug 20, 2026 by
BabyDrangoner
Contributor
•
Draft
3 of 4 tasks
[Bugfix][CPU][RISC-V] Return zero below exp clamp bound
bug
Something isn't working
cpu
Related to CPU backends
#53069
opened Aug 20, 2026 by
HakureiPOI
•
Draft
3 of 4 tasks
[ROCm][Quantization] Implement Fp8Config shared-expert FSE compatibility check
quantization
rocm
Related to AMD ROCm
#53068
opened Aug 20, 2026 by
jin-amd
Contributor
Loading…
[XPU][MoE] Tune Triton fused MoE for Intel XPU
intel-gpu
Related to Intel GPU
performance
Performance-related issues
[Bugfix][Compilation] Guard empty splitting ops before early return
bug
Something isn't working
torch.compile
#53061
opened Aug 20, 2026 by
ActiveSky
Contributor
Loading…
Previous Next
ProTip!
What’s not been updated in a month: updated:<2026-07-20.