Skip to content
View xuhuan51's full-sized avatar

Block or report xuhuan51

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
xuhuan51/README.md
Hi, I'm Guangli Liu. AI Agent Engineer.

I build reliable AI agents, from tool orchestration and evaluation
to high-performance LLM serving.

Runtime Tools Evaluation Serving

M.S. in Computer Science @ Zhejiang Normal University · Graduating 2027 · Open to opportunities

Animated developer workspace

What I Build

  • Production AI Agents — evidence-driven workflows, tool orchestration, retrieval, structured evaluation, and failure recovery.
  • Agent Platforms — sandbox runtimes, streaming events, asynchronous jobs, observability, and long-running task reliability.
  • LLM Systems — vLLM scheduling, KV-cache-aware serving, quantized deployment, speculative decoding, and GPU performance analysis.

Selected Work

System What I worked on
Argus & Alice Production AIOps agent and enterprise agent platform: evidence-based RCA, sandbox runtime, streaming orchestration, evaluation, and reliability.
DBOps Enterprise Copilot A LangGraph Text-to-SQL agent with hybrid schema retrieval, graph-based join planning, SQL verification, and observable Kubernetes delivery.
LLM Serving Stack Reproducible serving experiments covering vLLM scheduling, KV/prefix cache behavior, Qwen3 deployment, SLO analysis, and GPU topology.
DSpark Serving Benchmark Speculative-decoding evaluation across models, quantization modes, concurrency levels, and adaptive runtime selection.

Open Source

  • vLLM #51384 — bounded session-affinity scheduling for multi-turn workloads.
  • vLLM #51381 — session identity propagation for KV Events.

Working With

Python · C++ · CUDA · LangGraph · FastAPI · vLLM · Qwen3 · Docker · Kubernetes · Prometheus

From agent loops to GPU kernels.
Pixel animation from Hazy Readme Cards, used under the MIT License.

Popular repositories Loading

  1. dbops-enterprise-copilot dbops-enterprise-copilot Public

    企业级dbops-copilot

    Python 1

  2. Copilot_Hybrid-RAG Copilot_Hybrid-RAG Public

    Python

  3. llm-serving-stack llm-serving-stack Public

    Python

  4. dspark-spec-serving-benchmark dspark-spec-serving-benchmark Public

    Python

  5. vllm vllm Public

    Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python

  6. xuhuan51 xuhuan51 Public