Skip to content
View yuchenwang3's full-sized avatar
:octocat:
Focusing
:octocat:
Focusing

Block or report yuchenwang3

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
yuchenwang3/README.md

Yuchen (Ean) Wang — handwritten typing signature

Agentic post-training & ML systems.
M.S. CS @illinois · Research intern @Accio-Lab / @alibaba.

Website CV Scholar LinkedIn Email WeChat · eangyc

Research

Occamy-1.0: official logo

35B-A3B agent model for long-horizon tool use.
Core contributor · Post-training algorithms, infrastructure & data

Report Demo Model Code

CineFlow: figure from the paper

Dependency-driven parallel video generation.
Co-first author · 1.7–5.5× end-to-end speedup in the reported evaluation

Project Paper

Dynamic Prefill: figure from the project report

Adaptive batching and prompt packing for LLM serving.
Up to 20% lower TTFT on the reported traces

Report Code

RareDx: co-first author; graph-grounded RL for rare disease diagnosis CUDA Attention RL for Legal Reasoning

Open-source contributions

11 merged · 1 adopted solution · 28 open · 17 projects

Highlights

huggingface organization avatar
HF Datasets

#8670 · merged Preserve source shards during streaming shuffle so four workers can share 19 shards

modelscope organization avatar
mcore-bridge

#211 · merged Score packed QSA within each document; 2.70× faster in an 8K synthetic selector benchmark

2 more PRs in this repository

#212 · open Preserve low-precision rounding in gated residual mixing

#213 · open Bound PLE backward's extra workspace through chunked token reductions

vllm-project organization avatar
vLLM

#54699 · merged Remove full-weight copies during MoE loading; conversion peak 7.88 → 3.94 GiB in the exact-shape TP2 benchmark

1 more PR in this repository

#58219 · open Clarify Qwen3 parser boundary tokens in custom grammars

NVIDIA-NeMo organization avatar
NeMo RL

#3943 · open Bypass driver tensor materialization in distillation; 4.4–5.3× faster transfers in a controlled Ray benchmark

3 more PRs in this repository

#4193 · open Render evaluation prompts as complete conversations

#4176 · open Unify worker selection through configuration

#2962 · open Sanitize non-finite async log probabilities

NVIDIA organization avatar
Megatron-LM

#5396 · open Fuse GDN Q/K normalization to remove an extra backward activation buffer

#5463 · open Enable selective Mamba mixer recompute to save activation memory without full-layer recomputation

5 more PRs in this repository

#7864 · open Preserve native Adam step counters across checkpoint restoration

#7881 · open Reuse packed chunkwise CP metadata across GDN and KDA layers

#5400 · open Route GDN input projections to Adam

#5431 · open Exclude GDN input projections from global clipping

#5395 · open Skip gradient clipping for Muon

Dao-AILab organization avatar
FlashAttention

#2507 · merged Prevent redundant backward-kernel recompilation with stable cache keys, without CPU–GPU sync
Solution adopted by the PR author

flashinfer-ai organization avatar
FlashInfer

#4984 · merged Restore K/V calibration in FP8 KV prefill, correcting silently mis-scaled attention outputs

sgl-project organization avatar
SGLang

#39765 · open Publish Mamba cache updates before dependent batches capture stale KV mappings

3 more PRs in this repository

#40103 · open Reject developer messages silently dropped by templates

#38063 · open Explain cold MXFP4 JIT startup

#31621 · open Honor weight-check exclusions during reset

NousResearch organization avatar
Hermes Agent

#100693 · open Resolve local schema references so nested tool arguments reach handlers as objects, not JSON strings

3 more PRs in this repository

#113511 · open Control partial-stream continuation for batch evaluation

#113538 · open Clarify API retry budgets and streaming defaults

#102549 · open Make SSH reconnects race-safe

More contributions

Megatron Bridge · 2 contributions
NVIDIA-NeMo organization avatar
Megatron Bridge

#6315 · open Add Bridge-local Qwen4-Exp text-decoder support; GPU integration is pending

#6312 · open Make HF/Megatron comparison failures return a nonzero exit status

slime · 1 contribution
THUDM organization avatar
slime

#2412 · open Score fan-out rollout samples as a flat group

ms-swift · 4 contributions
modelscope organization avatar
ms-swift

#9598 · merged Add order-preserving packing

#9602 · merged Warm up NCCL before training

#9599 · merged Pass through Muon Nesterov settings

#9591 · merged Expose Muon coefficient selection

vime · 1 contribution
vllm-project organization avatar
vime

#337 · merged Forward recompute flags; fix hybrid models

Emerging Optimizers · 1 contribution
NVIDIA-NeMo organization avatar
Emerging Optimizers

#230 · merged Keep Muon scale-invariant

NeMo Gym · 1 contribution
NVIDIA-NeMo organization avatar
NeMo Gym

#2726 · merged Preserve HTTP errors across process boundaries
Co-author

verl · 2 contributions
verl-project organization avatar
verl

#7906 · open Track response truncation across context limits

#7597 · open Validate actor FSDP strategy

TRL · 1 contribution
huggingface organization avatar
TRL

#7294 · open Fix async checkpoint resume after stale rollout drops

All contributions ↗ Engineering notes ↗

Academic service

Reviewer, WSDM 2027.

On GitHub

Follow on GitHub Stars on my repositories Explore my open pull requests

GitHub contribution rhythm over 26 weeks Primary languages of my public non-fork repositories

Contribution trail

Snake animation of my GitHub contribution history

Pinned Loading

  1. Accio-Lab/occamy Accio-Lab/occamy Public

    Occamy model repo

    JavaScript 250 7

  2. vllm-project/vllm vllm-project/vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 93.6k 23.2k

  3. sgl-project/sglang sgl-project/sglang Public

    SGLang is a high-performance serving framework for large language models and multimodal models.

    Python 37k 9.5k

  4. NVIDIA/Megatron-LM NVIDIA/Megatron-LM Public

    Ongoing research training transformer models at scale

    Python 18.1k 4.6k

  5. verl-project/verl verl-project/verl Public

    verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

    Python 23.8k 4.7k

  6. huggingface/datasets huggingface/datasets Public

    🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools

    Python 22k 3.5k