Background
Phase 1 defines the research baseline for efficient memory allocation and SLO-aware scheduling under dynamic multi-model inference workloads.
Scope
- Survey LLM serving, memory abstraction, and SLO-aware scheduling literature.
- Compare kvcached, prism-research, and Aegaeon against vLLM-HUST + SAGE design assumptions.
- Define SAGE vs vLLM-HUST interface boundaries.
- Define interactive, long-generation, and multi-model mixed workloads.
Deliverables
- Survey report with comparison matrix.
- System design specification.
- Workload and baseline definitions.
Acceptance Criteria
- Survey of >=20 relevant papers/systems.
- Explicit comparison section covering kvcached, prism-research, and Aegaeon.
- Completed architecture and module boundary diagrams.
- Stable reproduction of baseline instances and workload replay.
Dependencies
Notes
Focus on dynamic peak load, static one-model-per-GPU inefficiency, and multi-model resource sharing.
Background
Phase 1 defines the research baseline for efficient memory allocation and SLO-aware scheduling under dynamic multi-model inference workloads.
Scope
Deliverables
Acceptance Criteria
Dependencies
Notes
Focus on dynamic peak load, static one-model-per-GPU inefficiency, and multi-model resource sharing.