The big idea
Motion is universal โ but the models built for it weren't. Inertia-1 brings the whole landscape under one roof.
01A fragmented field
Datasets disagree on the basics โ sampling rate, window length, sensor modality, body placement, even signal format โ and every task gets its own bespoke model. Findings rarely carry from one setup to the next.
02One unified exploration
Inertia-1 studies the full lifecycle of motion models โ data, sensing, objectives, and scale โ inside a single, controlled space instead of isolated one-offs.
03A general representation
The payoff: one representation that adapts across placements, devices, and tasks โ the same backbone, working far beyond the setting it was trained on.
Head Chest Back Arm Wrist Hand Hip Thigh Knee Shin Ankle Accelerometer Gyroscope Magnetometer Triaxial ENMO 0.2 Hz 1 Hz 5 Hz 20 Hz 10 s window 30 s window 60 s window 2 hr window Frequency domain Time domain Activity recognition Gait detection Longitudinal health Head Chest Back Arm Wrist Hand Hip Thigh Knee Shin Ankle Accelerometer Gyroscope Magnetometer Triaxial ENMO 0.2 Hz 1 Hz 5 Hz 20 Hz 10 s window 30 s window 60 s window 2 hr window Frequency domain Time domain Activity recognition Gait detection Longitudinal health
What we found
Beyond benchmarks, Inertia-1 surfaces the choices that decide whether a motion model actually works in the real world.
Pretrain once on the wrist, then point the model anywhere. It holds up on body placements โ and even sensor types like gyroscope and magnetometer โ that it never saw during training. No retraining for each new spot on the body.
Wrist accelerometer Other sensors ยท gyro, mag Other placements
Fused representation higher accuracy ยท cleaner motion clusters
Stack on more streams โ extra placements, gyroscope, magnetometer โ and the learned representation gets both more accurate and cleaner, with activities separating into tighter clusters. The streams are complementary: each one catches something the others miss.
Also worth knowing
How you capture motion shapes what a model can do with it. A few practical rules of thumb from the study.
Pretrained models stay strong even at a low 1 Hz for activity recognition; finer-grained health signals benefit from higher sampling rates.
30โ60 second windows hit the sweet spot across most tasks โ long enough to capture context, short enough to stay sharp.
Full triaxial input consistently beats collapsed vector-magnitude summaries โ the extra axes carry signal worth keeping.
Time-domain modeling preserves gait and health cues better than frequency-domain reconstruction.
How it works
The general representation comes together in three clean steps.
01
Pretrain at scale
Learn from planetary-scale accelerometry โ over 18 million hours across global cohorts โ with self-supervision, no labels required.
02
Transfer across settings
Adapt the same representation to new placements, devices, and sampling rates with light tuning โ or none at all.
03
Deploy across tasks
Power activity, mobility, and health applications from one backbone โ from fitness tracking to clinical screening.
Capabilities
The same representation spans the full spectrum of motion understanding.
Inertia-1 is a first step toward a unified motion foundation model โ and an open invitation to collaborators with motion data, new tasks, or a shared interest in where the field is headed.