I build reliable AI agents, from tool orchestration and evaluation
to high-performance LLM serving.
Runtime
Tools
Evaluation
Serving
- Production AI Agents — evidence-driven workflows, tool orchestration, retrieval, structured evaluation, and failure recovery.
- Agent Platforms — sandbox runtimes, streaming events, asynchronous jobs, observability, and long-running task reliability.
- LLM Systems — vLLM scheduling, KV-cache-aware serving, quantized deployment, speculative decoding, and GPU performance analysis.
| System | What I worked on |
|---|---|
| Argus & Alice | Production AIOps agent and enterprise agent platform: evidence-based RCA, sandbox runtime, streaming orchestration, evaluation, and reliability. |
| DBOps Enterprise Copilot | A LangGraph Text-to-SQL agent with hybrid schema retrieval, graph-based join planning, SQL verification, and observable Kubernetes delivery. |
| LLM Serving Stack | Reproducible serving experiments covering vLLM scheduling, KV/prefix cache behavior, Qwen3 deployment, SLO analysis, and GPU topology. |
| DSpark Serving Benchmark | Speculative-decoding evaluation across models, quantization modes, concurrency levels, and adaptive runtime selection. |
- vLLM #51384 — bounded session-affinity scheduling for multi-turn workloads.
- vLLM #51381 — session identity propagation for KV Events.
Python · C++ · CUDA · LangGraph · FastAPI · vLLM · Qwen3 · Docker · Kubernetes · Prometheus
Pixel animation from Hazy Readme Cards, used under the MIT License.