LLM Inference Systems | GPU Optimization | Multimodal AI Research
Designed 3-tier KV-cache (GPU→CPU→SSD) with SGLang + RDMA achieving 92.27% hit rate and 8.3x TTFT reduction on 4xA800.
First author IEEE TNNLS paper on NVFP4-DiT — 4-bit quantization for video diffusion achieving 4x memory reduction.
M.Eng. Software Engineering @ NWPU (985/211), Top 1%, GPA 88/100. Chinese Government Scholarship.
InfiX.ai | Shenzhen | Apr 2026 – Jun 2026
Hong Kong Amram Group | Shenzhen | Jun 2025 – Sep 2025
Shaanxi Longong | Xi’an | Feb 2025 – Apr 2025
First author. 4x memory reduction with FP4 quantization for multimodal diffusion transformers.
GitHub →EfficientNet+FFT + MFCC + Transformer with attention fusion for multimodal deepfake detection.
Thesis PDF →Apache Spark + CKKS homomorphic encryption. 95% privacy preservation, 20% performance gain.
Thesis PDF →Added unit tests for KV-cache eviction priority strategies in sgl-project/sglang (PR #31121) — directly relevant to production LLM serving.
View PR →Opened PRs and doc/typo fixes across multiple open-source projects (SGLang, SDKs, developer tooling), building a consistent public contribution history.
GitHub →LLM inference systems, KV-cache architecture, PD disaggregation, RDMA-based distributed serving on multi-GPU clusters.
Visual + audio + temporal fusion, deepfake detection, audio-guided video generation, cross-modal attention mechanisms.
Low-precision quantization (FP8/FP4), CUDA/Triton kernel optimization, FlashAttention, GEMM acceleration for inference.
Open to collaboration on LLM serving, multimodal AI, and GPU optimization. Reach out via email or LinkedIn.