Skip to content

Survey, architecture, and workload definition #1505

Description

@kms12425-ctrl

Background

Phase 1 defines the research baseline for efficient memory allocation and SLO-aware scheduling under dynamic multi-model inference workloads.

Scope

  • Survey LLM serving, memory abstraction, and SLO-aware scheduling literature.
  • Compare kvcached, prism-research, and Aegaeon against vLLM-HUST + SAGE design assumptions.
  • Define SAGE vs vLLM-HUST interface boundaries.
  • Define interactive, long-generation, and multi-model mixed workloads.

Deliverables

  • Survey report with comparison matrix.
  • System design specification.
  • Workload and baseline definitions.

Acceptance Criteria

  • Survey of >=20 relevant papers/systems.
  • Explicit comparison section covering kvcached, prism-research, and Aegaeon.
  • Completed architecture and module boundary diagrams.
  • Stable reproduction of baseline instances and workload replay.

Dependencies

Notes

Focus on dynamic peak load, static one-model-per-GPU inefficiency, and multi-model resource sharing.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions