# anypoint-datasets — the public data bucket

Bucket root mirrors `data/` in the repo. Synced with
`scripts/upload/hf_upload.py` / `scripts/download/download_bundle.py`, both
driven by `scripts/upload/hf_upload_manifest.yaml`. Companion:
**TorridFish/anypoint-checkpoints** (private) — stage-1 weights and their
frozen-encoder latent caches.

## What is here

| prefix | size | what it is |
|---|---|---|
| `feature_store/` | 376 GB | Phase-2 per-point store, 2305 train-mix scenes. Per scene: `feat.safetensors` (bf16 `[N_down, 1232]`), `coord.npy`, `objects.npz` (one int32 index array per mask key), `voxels_1_5.npz`, `voxels_1_0.npz`. Root `manifest.json` is the validity gate. |
| `feature_store_eval/` | 103 GB | Same format, **587 eval scenes**: the grounding and QA val sets plus both scene-caption val sets. Real-3DQA's 67 scenes are a subset. |
| `feature_store_hm3d_color/` | 189 GB | 134 five-room HM3D assemblies — `xr_scene_caption`. |
| `feature_store_hm3dqa_color/` | 154 GB | 160 three-room HM3D assemblies — `xr_qa`. Separate store because it shares scene ids with the row above, and a store is keyed by scene. |
| `grounded-3d-llm-proposal-top100/` | 0.1 GB | Packed proposals. Read at run time even on the store path: `local_scope="all"` counts columns from the npz header, and eval's `--proposal_npz_dir` is a cache key, not a path. |
| `scanrefer_gt_masks/`, `scannet_scene_masks/` | 3.7 GB + 224 MB | The GT and scene-level masks the QA and scene-caption prompts name. |
| `g3dllm_proposal_mappings/`, `grounding_box_table_g3dllm_val.pt` | 19 MB | proposal↔GT mapping and the box table the grounding/Scan2Cap evaluators resolve. |
| `sysobj_*.json` (21) | 0.37 GB | The canonical train mix and the eval val sets. The store's `data_paths_hash` is over exactly this path list, so substituting one invalidates the store it belongs to. |
| `eval/scan2cap/annotations/` | 9 MB | Scan2Cap references and corpus. |

Total: **~830 GB**.

## What changed on 2026-08-21

**Eval now runs from this bucket.** It previously could not: its masks and globals
lived in `precomputed_voxel_1_5_g3dllm/` and `precomputed_voxel_1_5_eval_g3d/`,
which were dropped on 2026-08-20 to save space. Three things closed that gap —
the eval store grew from 141 to 587 scenes, the two HM3D stores arrived, and the
3.9 GB of masks/mappings/box table above are now carried. Every eval entry point
reads a store: the sweep, Real-3DQA's driver, both XR drivers, the smoke scripts.

**The store cannot be rebuilt from this bucket**, and that is unchanged: its
inputs are the source Sonata archives (`scannet_sonata_feat` 198 GB, `sonata_feat`
47 GB, `hm3d_sonata_feat*_color` 201 GB), none of which are mirrored here. The
local copies are the only ones. The store is the artefact now.

A note on the HM3D pair, because the trade is not obvious: their *sources* are
201 GB against 343 GB of store, and a rebuild is ~6 minutes of CPU — so mirroring
the sources would have been cheaper. They are here because the sources are not,
and without them the XR benchmarks cannot run from this bucket at all.

`voxels.npz` (the pre-multigrid untagged name) is not synced: it is a hardlink to
`voxels_1_5.npz` locally, and every reader uses the tagged name.

## Using it

```bash
export HF_TOKEN=hf_...            # or: hf auth login -- required, checked up front
python scripts/download/download_bundle.py --dry-run          # where will it land?
python scripts/download/download_bundle.py --bucket datasets
python scripts/download/download_bundle.py --only feature_store_eval   # one piece
```

The pull checks each store against its own `manifest.json` afterwards and reports
`store complete: N scenes`, because an interrupted pull can leave a directory that
looks valid — the manifest is the validity gate and does not necessarily arrive
last.

Stores live under `data/` (`paths.store_root()`, override `STORE_ROOT`) — a store
is shared *input* that every run and every benchmark reads, so it sits beside the
datasets rather than in a private cache. Then:

```bash
# Stage 1 reads the store directly; Stage 2's latent cache is built FROM it
VOXEL_SIZE=1.0 bash src/run_train_ddp.sh
TRAIN_STAGE=2 WRAPPER_CKPT_PATH=<stage1-ckpt> bash src/run_train_ddp.sh

# Eval: no voxelization prestep exists on this path
bash eval/submit_eval.sh
```

`VOXEL_SIZE` *selects* a grid rather than triggering a rebuild; all four stores
carry 1.0 and 1.5. A grid a store lacks raises and names the command that derives
it (23 s for 2305 scenes, no features re-read).

## What does not travel

The source Sonata archives (above), and `voxel_point_coords` (A6 / coord_rel),
which is absent by design: it is `coord[perm]` from the store, verified bit-exact.
The archived 1.7 TB coordrel root was deleted on 2026-08-21 and could not be
reproduced anyway — it was built from an older Sonata version with a different
point *set* — so its recorded numbers are not comparable to anything derived here.
