Topanga AI Computing offers four inference runtime variants: vLLM, NVIDIA NIM, Predefined Image Command, and Custom. The first three start from a built-in launch commands and only need a few env vars. Use Custom (the path used here) when you want full control via your own model-definition.yaml—for example, downloading the model on first start, picking specific vLLM flags, or running pre-start actions.
You will need a HuggingFace access token with permission to the model you plan to serve. For gated models like Gemma, you must also accept the license on HuggingFace first.
You will also need a GPU resource preset large enough for the model. A 7B model in bfloat16 needs ~16 GB GPU memory plus headroom for the KV cache; budget at least 0.2 H200 or equivalent.
In Data → Models, create a new folder with usage type Model. Leave it empty for now. Both the YAML and the downloaded weights will live here.
The folder must be created as a
modelfolder (notgeneral/data), otherwise it will not show up in the "Model Storage To Mount" picker on the Start Service form.
This file tells Backend.AI how to start the inference container, which port to expose, and how to health-check it. Save it locally as model-definition.yaml:
models:
- name: "gemma-7b-it"
model_path: "/models"
service:
start_command:
- /bin/bash
- -lc
- >
set -eo pipefail;
MODEL_DIR=/models;
MODEL_ID=google/gemma-7b-it;
HF_TOKEN_VALUE="$(printenv HF_TOKEN || true)";
if [ -z "$HF_TOKEN_VALUE" ]; then
echo "HF_TOKEN is not set";
exit 1;
fi;
if [ ! -f "$MODEL_DIR/config.json" ]; then
echo "Downloading $MODEL_ID into $MODEL_DIR ...";
if command -v hf >/dev/null 2>&1; then
hf download "$MODEL_ID" --local-dir "$MODEL_DIR" --token "$HF_TOKEN_VALUE";
else
huggingface-cli download --local-dir "$MODEL_DIR" --token "$HF_TOKEN_VALUE" "$MODEL_ID";
fi;
else
echo "Model already present in $MODEL_DIR, skipping download.";
fi;
exec vllm serve "$MODEL_DIR"
--host 0.0.0.0
--port 8000
--served-model-name gemma-7b-it
--dtype bfloat16
--max-model-len 4096
--generation-config vllm
port: 8000
health_check:
path: /v1/models
max_retries: 500What each field does:
name— the model identifier inside Topanga. Match this to--served-model-nameso the same string works in API calls.model_path— where the model folder is mounted inside the container./modelsis the Topanga default.service.start_command— the process the container runs. Here it (1) checks forHF_TOKEN, (2) downloads weights to/modelson first start (idempotent — re-uses the cache on restart), and (3) execsvllm servewith an OpenAI-compatible API on port 8000.service.port— the container port the service listens on. Topanga's AppProxy maps this to the public endpoint.service.health_check.path— endpoint the proxy polls./v1/modelsis reliable for vLLM because it only returns 200 once the model is fully loaded.service.health_check.max_retries—500is intentionally high to tolerate slow first-time HuggingFace downloads (multi-GB weights). Lower it to ~20 once weights are cached.
Other useful health-check options (defaults shown):
interval: 10.0— seconds between checks.max_wait_time: 15.0— per-check HTTP timeout.expected_status_code: 200.initial_delay: 60.0— wait after container start before first check. Bump this to300.0for 70B+ models.
Open the model folder you just created and upload the YAML at the root.
Go to Serving and click Start Service
Key fields to set:
- Service Name — any identifier (used in URLs and logs).
- Open To Public — leave off for token-protected access (recommended). When off, you'll generate an API token after launch.
- Inference Runtime Variant — choose Custom so Topanga uses your
model-definition.yaml. - Model Storage To Mount — pick the folder you created in step 1.
- Model Definition File Path —
model-definition.yaml(default). - Number of Replicas —
1for a single GPU; increase only if you have the resources and need throughput. - Environment / Version — pick the vLLM image. Here, we use the version 0.20.2.
- Resource Allocation — assign at least 0.5 H200 GPU. For Gemma-7B in bf16 with
max-model-len 4096, this would be a safe minimum. You can increase the model-len as well. - Environment Variables — add
HF_TOKEN=<your_hf_token>. This is read by thestart_commandto authenticate the download.
Select Create at the bottom.
Tip: there's a Validate button on the launcher that runs the start command in a test container and shows the log. Use it once before going live to catch YAML syntax errors or missing env vars early.
After creation you'll be redirected to the service detail page.
Status flow:
- DEGRADED — initial state during
initial_delay. Model is loading; no traffic is routed yet. First-run downloads (multi-GB weights) can keep you here for several minutes. - HEALTHY —
/v1/modelsreturned 200 within the timeout. The endpoint is now serving traffic. - UNHEALTHY — too many consecutive failed health checks. Common causes: insufficient GPU/RAM, malformed
model-definition.yaml, missingHF_TOKEN, or HuggingFace gating. Open the routing's container log to see why. Fix and click Clear Error and Retry.
Time-to-UNHEALTHY at startup with defaults is initial_delay + interval × (max_retries + 1). With max_retries: 500 and interval: 10s you get ~83 minutes of grace, which is what you want for a first-time download.
While the service runs, you can see its kernel under Sessions.
You can check the logs of the container by selecting the See Container Logs button.
Wait for the server to start.
The next step is to generate an API token. This token will be used to authenticate your requests to the API.
First, click on the Endpoint Name.
Scroll down to find the Generated Token section and click on the Generate Token button.
Here, you need to specify the expiration date of the token. The default is 7 days. The, click Generate.
Copy the token and save it. You will need this token to authenticate your requests to the API.
In the left column menu, under Playground, select Chat to open the WebUI's chat against your endpoint.
Select the serving service from the dropdown menu and then choose the model you want to use.
After that, you need to choose one of the API tokens you generated in the previous step. After that, please click on Refresh Model Information.
This will take you to the chat interface. You can now chat with the model.
Get the public endpoint from the Service Endpoint URL (port 10602 for this service).
export TOPANGA_API_KEY="<your_topangas_api_key>"
curl -H "Content-Type: application/json" \
-H "Authorization: Bearer $TOPANGA_API_KEY" \
https://topanga.carc.usc.edu:10602/v1/chat/completions \
-d '{
"model": "gemma-7b-it",
"messages": [{"role": "user", "content": "Hello, who are you?"}],
"max_tokens": 100
}'The model field must match --served-model-name from your start_command (here, gemma-7b-it).
Because vLLM exposes the OpenAI-compatible API, you can also point the official OpenAI Python SDK at the same URL by setting base_url to https://topanga.carc.usc.edu:10602/v1 and using Bearer $TOPANGA_API_KEY as the bearer.
- Stuck in DEGRADED: Open the container log. Likely an OOM during model load, a missing/invalid
HF_TOKEN, or you forgot to accept the model license on HuggingFace. HF_TOKEN is not setin logs: You forgot to addHF_TOKENunder Environment Variables on the Start Service form.- Health check failing on
/v1/models: vLLM hasn't finished loading. Either raiseinitial_delayandmax_retries, or use/healthinstead — note that/healthreturns 200 earlier in vLLM's startup, before the model is ready. - API call returns 401/403: Token expired or wrong header. The header is
Authorization: Bearer <token>.
When you're done, terminate the service from the Serving tab. Idle services keep replicas alive and consume GPU quota.



















