Skip to content
View guqiong96's full-sized avatar

Block or report guqiong96

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Popular repositories Loading

  1. Lvllm Lvllm Public

    LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, su…

    Python 462 41

  2. Lsglang Lsglang Public

    Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parallel architecture, suitable for MOE model hybrid inference.

    Python 142 15

  3. Lvllmds4-x Lvllmds4-x Public archive

    CPU-GPU hybrid inference for DeepSeek-V4 on NVIDIA SM80+ (A100/RTX 4090 etc.), forked from yhfgyyf/vllm-deepseek-v4-sm89.

    Python 60 12

  4. Lvllmds4 Lvllmds4 Public archive

    A fork of jasl/vllm (codex/ds4-sm120-min-enable) with CPU-GPU hybrid inference support for DeepSeek-V4 on SM120+.

    Python 31 3

  5. lktransformers lktransformers Public archive

    The complete NUMA-optimized branch of the ktransformers project

    Python 26 10