TartanIMU Challenge - Unified Carrier-Conditioned Model

One unified model, applied identically to four platforms (car, dog, drone, human). It reads a 1.0 s window of raw 6-axis IMU and predicts that window's mean body-frame velocity. No platform label, no platform recovery, no ground-truth orientation or position, no internet. Every model in the blend runs on every window; there is no per-platform routing or switching. Re-executes in a few minutes and reproduces submission.csv exactly.

Files

  • infer.py - self-contained offline inference: builds the model, loads frozen weights from weights/, reads the test .npz files, writes submission.csv.
  • model.py, features.py, config.py - model definitions and gravity-aware feature engineering.
  • weights/*.pt - the frozen weights (all ensemble branches + the platform classifier).
  • weights/norm.npz - per-channel input normalisation statistics.
  • submission.csv - the exact submission that scored on the leaderboard.
  • requirements.txt - pinned dependencies.

Run

pip install -r requirements.txt
python infer.py --data_root /path/to/data --out submission.csv

--data_root contains test/<traj_id>.npz (each with imu of shape (N,6) = [ax,ay,az,gx,gy,gz] in the body frame at 200 Hz), index/test_windows.csv (traj_id, win_idx, window_id) and sample_submission.csv. The script windows each trajectory into non-overlapping 200-frame blocks and emits one velocity per window, keyed by window_id.

Method

The six raw IMU channels are extended to nine by a gravity-aware split (a causal EMA of the accelerometer tracks gravity; subtracting it isolates linear acceleration). The prediction is a fixed-weight blend of a small set of decorrelated branches that all read the same window:

  • carrier-conditioned Mixture-of-Experts models (MosaicIMU-style): a shared spectrogram+TCN encoder feeds a learnable prototype router (K experts, cosine similarity) that infers the platform from the signal itself and soft-blends expert heads - an internal, input-driven adaptation with no external label;
  • transformer encoders over downsampled convolutional features;
  • a multi-resolution short-time-FFT spectrogram branch.

Frequency/spectrogram views capture each platform's distinct rhythm (a car's smooth non-holonomic motion, the pulse-and-stop of a walking human or trotting dog, a drone's free 6-DOF motion). All adaptation is internal and input-driven (the MoE prototype router + FiLM); no platform label is inferred or used.

Rule compliance

  • One unified model, evaluated identically on all four platforms; no per-platform expert routing by external label. Platform adaptation is internal and input-driven (MoE prototype router + FiLM).
  • Inference consumes raw 6-axis IMU only. No ground truth and no supplied platform label are read.
  • No attempt to recover the anonymised platform identity.
  • Fully offline, well under the 16 GB / 2 h re-execution limits (a few minutes on one GPU).
  • Re-running infer.py reproduces submission.csv exactly (max abs diff 0.0).
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support