Axis Robotics has launched Axis Sim Dataset V1, one of many largest open-source simulation datasets for Franka arm manipulation, with the total dataset, coaching code, and benchmarks publicly obtainable. V1 is constructed from greater than 50,000 human-teleoperated simulation trajectories throughout 207 manipulation duties and 60,000+ scene variants on a simulated Franka Analysis 3 arm.
This dataset drew over 160,000 downloads, making it essentially the most downloaded open-source simulation Franka manipulation dataset on Hugging Face. In benchmarks, continuous pretraining on V1 lifted π0.5 and beat a volume-matched RoboCasa baseline, with each end result open and verifiable.

Axis Robotics is constructing the final word compounding knowledge engine for Bodily AI, a vertically built-in system spanning large-scale simulation, selfish real-world seize, humanoid loco-manipulation, and human-gated DAgger post-training. The corporate raised $12 million in seed funding led by Hack VC, with participation from Nomad Capital, Pi Community Ventures, 10K Ventures, and angel traders.
A Guess Towards “Clear Knowledge Solely”
A standard assumption in robotics is that demonstrations have to be near-optimal to start with — filter right down to professional trajectories, standardize the setup, and discard something noisy earlier than it’s protected to mimic. Axis’s thesis runs the opposite means: knowledge high quality lives on the distribution stage, not the one trajectory. When a big and various sufficient crowd produces noisy, suboptimal trajectories and their errors are uncorrelated, the noise averages out and a working coverage survives throughout coaching.
Axis Sim Dataset V1 places that thesis to a public check. Its trajectories span pick-and-place, stacking, pouring, articulated-object manipulation, and gear use, all collected via Axis’s browser-based teleoperation platform, Axis Hub, by a distributed crowd moderately than a single professional group. The dataset was constructed with researchers from UC Berkeley, Johns Hopkins, the College of Michigan, and different establishments.
Outcomes That Scale
On LIBERO-Plus, continuous pretraining on V1 lifts π0.5 from 83.9% to 88.8% success and outperforms a volume-matched RoboCasa365 baseline by 37.3%. Efficiency improves persistently as pretraining knowledge scales from 25% to 100% of the dataset, with no saturation in sight, proof that the positive aspects come from range and protection moderately than a one-off bump. The most important enhancements seem underneath digital camera, sensor-noise, and structure perturbations, the precise axes Axis randomizes throughout era.

The group says V2 is already underway, scaling to 1.2 million trajectories throughout 1,200 duties, with cross-embodiment generalization and outcomes throughout a number of VLA fashions exhibiting that suboptimal simulation knowledge trains sturdy insurance policies.
The Engine Behind the Dataset
The dataset is one output of a bigger, actively compounding knowledge engine. The place a standard knowledge vendor collects to a hard and fast spec and stops, Axis makes use of mannequin efficiency and failure circumstances to find out what must be collected subsequent, so each coaching spherical informs the subsequent. That engine runs on a hybrid technique throughout 4 knowledge strains, and all 4 now run at scale:
- Simulation: over 200,000 distributed contributors on Axis Hub, a top-3 dApp on Base, producing 4.7M+ trajectories throughout 13 embodiments.
- Selfish: a managed community of 1,000+ full-time, QC-trained collectors capturing first-person exercise in actual houses and companies throughout 14 industries: 200,000+ hours already banked and rising by 4,000+ hours day by day, with Vicon-verified hand pose.
- Loco-manipulation: 500+ hours combining mobility and dexterity on actual humanoids (Unitree G1, Booster T2) via hardware-agnostic teleoperation.
- Human-gated DAgger post-training: 500+ hours of human-in-the-loop correction focused at deployment edge circumstances.
Each job and trajectory is recorded on-chain on Base for provenance, and contributors are rewarded for verified work high quality.
From Open Knowledge to Industrial Deployment
Past open-sourcing simulation knowledge, Axis works immediately with robotic embodiment corporations to construct personalized, embodiment-specific knowledge pipelines and mannequin priors.
As Booster Robotics’ first sim-data associate, Axis rebuilt Booster’s actual workspace as a task-aligned digital twin, had distributed contributors accumulate 42,000+ simulation episodes on it, and distilled them right into a Booster-specific mannequin prior. With simply 30 real-robot demos, that prior reached 87.5% success versus 37.5% for an out-of-the-box π0.5, matching π0.5 utilizing half the real-world demonstrations.
Different companions span embodiment corporations (Feagine Robotics), mannequin corporations (Manycore Tech, Dexmal) and industrial automation (Lotus Vehicles, Geely Auto). Axis additionally provides on-chain robotics networks: BitRobot on Solana and OpenRoboto on Bittensor.
Redefining Bodily AI’s Knowledge Basis
“The way forward for Bodily AI isn’t a static dataset you obtain as soon as,” stated Chris Feng, founding father of Axis Robotics. “It’s an engine that retains producing the info the mannequin wants subsequent. Scale will get you broad protection. Range retains the noise unbiased. The closed loop turns each failure into progress. That’s what compounds.”
Axis was based by researchers from UC Berkeley, CMU, Georgia Tech, and SJTU, alongside serial founders who’ve scaled shopper platforms to over 30 million customers. Its analysis is suggested by Jiachen Li, Assistant Professor at Georgia Tech.
Paper Hyperlink: https://arxiv.org/abs/2607.21588
Venture Web page: https://axisaiorg.github.io/AXIS-V1/
Dataset Hyperlink: https://huggingface.co/datasets/axisrobotics/Franka-Dataset
Github Codebase: https://github.com/AxisAIOrg/Axis-V1-Coaching
















