Skip to content

Repository files navigation

ATLAS-model

Surgical Anatomy Recognition with Context Learning using Foundation Representations

This repository contains the model implementation and training code for ATLAS, a clip-level semantic segmentation model for surgical anatomy recognition in minimally invasive surgery (MIS). It is trained and evaluated on the ATLAS-120k dataset — over 120,000 annotated frames from 100 surgical videos spanning 14 procedures and 42 anatomical classes.

ATLAS builds on the EoMT (Encoder-only Mask Transformer) architecture, using SurgeNet-pretrained DINOv3 surgical foundation backbones as the encoder. On top of standard mask classification it adds surgical-context learning: per-procedure recognition, learnable phase tokens, and a phase contrastive objective over short video clips.

Part of the ATLAS project ecosystem.

Features

  • Clip-level semantic segmentation of surgical anatomy (operates on short video clips, not just single frames).
  • DINOv3 surgical backbones (ViT-S / ViT-B / ViT-L) pretrained on SurgeNet.
  • Phase tokens & phase contrastive learning — learnable tokens that capture surgical phase/context, trained with a contrastive loss across temporally neighbouring frames.
  • Procedure recognition — auxiliary classification head over the 14 procedure types.
  • Temporal query learning — queries are propagated across frames within a clip.
  • Built on PyTorch Lightning with a config-driven LightningCLI interface.

Repository structure

.
├── main.py                  # LightningCLI entry point (fit / validate / test)
├── train.sh                 # SLURM + Apptainer training launcher
├── configs/                 # Experiment configs (one per backbone size)
│   ├── ATLAS_DINOv3_vits.yaml
│   ├── ATLAS_DINOv3_vitb.yaml
│   └── ATLAS_DINOv3_vitl.yaml
├── models/
│   ├── atlas.py             # ATLAS model (EoMT + phase/procedure/temporal heads)
│   ├── vit.py               # DINOv3 ViT encoder wrapper
│   └── scale_block.py       # Feature upscaling block
├── training/                # Lightning modules, losses, LR schedule
├── datasets/                # ATLAS dataset + Lightning data module + transforms
├── utils/
│   └── extract_model_weights.py   # Export weights from a Lightning checkpoint to .pth
├── Dockerfile               # Container definition
├── DOCKER.md                # How to build the container for train.sh
└── requirements.txt

Installation

Requires Python 3.10 and a CUDA-capable GPU.

pip install -r requirements.txt

Key dependencies: torch==2.7.0, torchvision==0.22.0, lightning==2.5.1, timm==1.0.15, transformers==4.56.1, wandb==0.19.10.

For a reproducible environment (and for running on HPC via Apptainer), see DOCKER.md.

Data & checkpoints

  • Dataset: download and process the ATLAS-120k data with the ATLAS repository. Training expects the packaged dataset (e.g. atlas.zip), passed via --data.path.
  • Backbone checkpoints: SurgeNet-pretrained DINOv3 weights (e.g. DINOv3-vits-256-surgenet2M.pth), passed via --model.ckpt_path.

Download the model weights here

You can download the model weights using the provided links below.

Model Variant Download
DINOv1 ViT-b Download
DINOv2 ViT-b Download
DINOv3 ViT-s Download
DINOv3 ViT-b Download
DINOv3 ViT-l Download

Training

Quick start

Training is driven by main.py and a YAML config. Pick the backbone size via the config file:

python main.py fit \
  -c configs/ATLAS_DINOv3_vits.yaml \
  --data.path /path/to/atlas.zip \
  --model.ckpt_path /path/to/DINOv3-vits-256-surgenet2M.pth \
  --trainer.devices 1 \
  --data.batch_size 24 \
  --trainer.max_epochs 3

Any config value can be overridden from the command line (CLI args take precedence over the YAML). Three backbone sizes are provided:

Config Backbone
configs/ATLAS_DINOv3_vits.yaml DINOv3 ViT-Small
configs/ATLAS_DINOv3_vitb.yaml DINOv3 ViT-Base
configs/ATLAS_DINOv3_vitl.yaml DINOv3 ViT-Large

On a SLURM cluster

train.sh wraps the command above for a SLURM + Apptainer setup. Edit the user-defined paths at the top of the file (data, container .sif, backbone checkpoints, results directory), set MODEL_VARIANT to vits, vitb, or vitl, and submit:

sbatch train.sh

It runs a single model variant with a single seed and logs to Weights & Biases.

Exporting weights

To extract a plain .pth state dict from a Lightning .ckpt checkpoint (e.g. to reuse the trained backbone elsewhere):

python utils/extract_model_weights.py /path/to/checkpoint.ckpt

Finetuned weights

The weights of the ATLAS model trained on the ATLAS-120k dataset can be found below.

Model Download
ATLAS ViT-s Download
ATLAS ViT-b Download
ATLAS ViT-l Download

Citation

If you use ATLAS-120k in your research, please cite:

@misc{dejong2026atlas,
      title={Surgical Anatomy Recognition with Context Learning using Foundation Representations}, 
      author={Ronald L. P. D. de Jong and Tim J. M. Jaspers and Raf A. H. Vervoort and Aron F. H. A. Bakker and Yiping Li and Jip L. Tolenaar and Jelle P. Ruurda and Willem M. Brinkman and Josien P. W. Pluim and Marcel Breeuwer and Daan de Geus and Fons van der Sommen},
      year={2026},
      eprint={2606.22124},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2606.22124}, 
}

About

No description, website, or topics provided.

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages