CalfVO

Monocular visual odometry without calibration or test-time optimization.

Project page | Paper | Code

Given uncalibrated video, CalfVO predicts a metric camera trajectory in a single forward pass at 53 FPS, with no camera intrinsics, no bundle adjustment and no loop closure. A transformer regresses relative camera poses over overlapping image windows together with separate rotation and translation confidences, supervised by camera poses alone, and a confidence-weighted module aggregates the overlapping predictions into one trajectory. Scale is recovered from learned priors, accurately enough that the trajectories are evaluated without any alignment to the ground truth.

The two checkpoints have different licences

This repository holds two models. They differ only in the frozen image encoder, but they are not under the same terms, because their encoders are not.

File Encoder Licence Commercial use
calfvo_croco.pth CroCo v2, as distributed with DUSt3R CC BY-NC-SA 4.0 No
calfvo_dinov2.pth DINOv2 ViT-L/14 Apache-2.0 Yes

calfvo_croco.pth contains Naver's CroCo v2 encoder weights. Those are CC BY-NC-SA 4.0, which means non-commercial and share-alike, and the file inherits them even though the training code is MIT.

calfvo_dinov2.pth contains only Meta's DINOv2 ViT-L/14, which is Apache-2.0. Together with the MIT code, that file can be used commercially. It is the weaker of the two on most metrics.

Which one the paper reports

The paper reports calfvo_croco.pth. Use that one to compare against CalfVO.

Results

Both are step 57000 of the same schedule on the same data. t_rel and r_rel are evo's relative pose error at a one-frame delta over all pairs, in metres and degrees per frame. Lower is better.

Test set croco t_rel dinov2 t_rel croco r_rel dinov2 r_rel
ScanNet (97 seq) 0.0054 0.0061 0.1634 0.2010
TUM RGB-D (9 seq) 0.0052 0.0054 0.3355 0.3593
TartanAir v1 (18 seq) 0.1027 0.0919 0.2815 0.3962
KITTI 09-10 (2 seq) 0.1620 0.1867 0.1231 0.2141
EuRoC MAV (11 seq) 0.0207 0.0177 0.3133 0.2584

CroCo is better on rotation for four of the five sets, by 7% to 74%. Translation is split, and DINOv2 wins on TartanAir and EuRoC. CroCo v2 is pretrained on cross-view completion, which is a geometric correspondence task, and that may be why it recovers rotation between frames better than DINOv2's semantic features do.

Usage

Clone the code, then load a checkpoint into the model its config builds.

import torch
from huggingface_hub import hf_hub_download
from src.model import get_model
from src.utils import io_utils

config = io_utils.read_yaml_file("configs/calfvo_croco.yaml")
config["model"]["num_views"] = config["num_views"]

model = get_model(config["model"])
path = hf_hub_download("voviktyl/CalfVO", "calfvo_croco.pth")
ckpt = torch.load(path, map_location="cpu")
model.load_state_dict(ckpt["model_state_dict"])
model.eval()

Pair calfvo_dinov2.pth with configs/calfvo_dinov2.yaml. The two checkpoints are not interchangeable. CroCo emits a 14x14 token grid at 224 px and DINOv2 emits 16x16, so the positional tables differ in size and loading one under the other's config raises an error instead of failing quietly.

Each file holds every weight the model runs with, the frozen encoder included, so one file is the whole model. Building the model still pulls the encoder from its own repository first, and the checkpoint then overwrites it, so a first run downloads both. The optimizer state is stripped, so these are for inference and cannot resume training.

To run the model over a video and write out a trajectory:

python scripts/run_vo_demo.py \
    --config configs/calfvo_croco.yaml \
    --video /path/to/clip.mp4 \
    --output_dir output/demo/vo

The checkpoint for the config's backbone is fetched from here and cached. Pass --checkpoint to use a file of your own instead.

Input

Eight-frame RGB windows at 224x224. The transform centre-crops each frame to its shortest edge before resizing, so any aspect ratio works and no intrinsics are needed. The model was trained on consecutive frames of roughly 30 fps footage and predicts the motion between whichever frames it is given, so use --stride to subsample a clip that moves faster than that.

Training data

ScanNet, TartanAir v2, KITTI odometry 00-08 and Virtual KITTI 2, mixed 32/36/16/16%. Evaluation holds out the ScanNet test split, TartanAir v1, TUM RGB-D, KITTI 09-10 and EuRoC MAV.

Citation

@misc{yugay2026calfvo,
      title={Monocular Visual Odometry without Calibration or Test-time Optimization},
      author={Vladimir Yugay and Duy-Kien Nguyen and Theo Gevers and Cees G. M. Snoek and Martin R. Oswald},
      year={2026},
      eprint={2510.03348},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2510.03348},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for voviktyl/CalfVO