Axis Robotics has launched Axis Sim Dataset V1, one of many largest open-source simulation datasets for Franka arm manipulation, with the total dataset, coaching code, and benchmarks publicly out there. V1 is constructed from greater than 50,000 human-teleoperated simulation trajectories throughout 207 manipulation duties and 60,000+ scene variants on a simulated Franka Analysis 3 arm.
This dataset drew over 160,000 downloads, making it essentially the most downloaded open-source simulation Franka manipulation dataset on Hugging Face. In benchmarks, continuous pretraining on V1 lifted π0.5 and beat a volume-matched RoboCasa baseline, with each consequence open and verifiable.
Axis Robotics is constructing the final word compounding information engine for Bodily AI, a vertically built-in system spanning large-scale simulation, selfish real-world seize, humanoid loco-manipulation, and human-gated DAgger post-training. The corporate raised $12 million in seed funding led by Hack VC, with participation from Nomad Capital, Pi Community Ventures, 10K Ventures, and angel traders.
A Wager Towards “Clear Information Solely”
A standard assumption in robotics is that demonstrations should be near-optimal to start with — filter all the way down to knowledgeable trajectories, standardize the setup, and discard something noisy earlier than it’s secure to mimic. Axis’s thesis runs the opposite manner: information high quality lives on the distribution stage, not the one trajectory. When a big and various sufficient crowd produces noisy, suboptimal trajectories and their errors are uncorrelated, the noise averages out and a working coverage survives throughout coaching.
Axis Sim Dataset V1 places that thesis to a public take a look at. Its trajectories span pick-and-place, stacking, pouring, articulated-object manipulation, and gear use, all collected via Axis’s browser-based teleoperation platform, Axis Hub, by a distributed crowd moderately than a single knowledgeable workforce. The dataset was constructed with researchers from UC Berkeley, Johns Hopkins, the College of Michigan, and different establishments.
Outcomes That Scale
On LIBERO-Plus, continuous pretraining on V1 lifts π0.5 from 83.9% to 88.8% success and outperforms a volume-matched RoboCasa365 baseline by 37.3%. Efficiency improves constantly as pretraining information scales from 25% to 100% of the dataset, with no saturation in sight, proof that the beneficial properties come from variety and protection moderately than a one-off bump. The most important enhancements seem underneath digital camera, sensor-noise, and structure perturbations, the precise axes Axis randomizes throughout era.

The workforce says V2 is already underway, scaling to 1.2 million trajectories throughout 1,200 duties, with cross-embodiment generalization and outcomes throughout a number of VLA fashions exhibiting that suboptimal simulation information trains sturdy insurance policies.
The Engine Behind the Dataset
The dataset is one output of a bigger, actively compounding information engine. The place a standard information vendor collects to a set spec and stops, Axis makes use of mannequin efficiency and failure instances to find out what ought to be collected subsequent, so each coaching spherical informs the subsequent. That engine runs on a hybrid technique throughout 4 information strains, and all 4 now run at scale:
Simulation: over 200,000 distributed contributors on Axis Hub, a top-3 dApp on Base, producing 4.7M+ trajectories throughout 13 embodiments.
Selfish: a managed community of 1,000+ full-time, QC-trained collectors capturing first-person exercise in actual houses and companies throughout 14 industries: 200,000+ hours already banked and rising by 4,000+ hours on daily basis, with Vicon-verified hand pose.
Loco-manipulation: 500+ hours combining mobility and dexterity on actual humanoids (Unitree G1, Booster T2) via hardware-agnostic teleoperation.
Human-gated DAgger post-training: 500+ hours of human-in-the-loop correction focused at deployment edge instances.
Each job and trajectory is recorded on-chain on Base for provenance, and contributors are rewarded for verified work high quality.
From Open Information to Business Deployment
Past open-sourcing simulation information, Axis works straight with robotic embodiment corporations to construct personalized, embodiment-specific information pipelines and mannequin priors.
As Booster Robotics’ first sim-data associate, Axis rebuilt Booster’s actual workspace as a task-aligned digital twin, had distributed contributors acquire 42,000+ simulation episodes on it, and distilled them right into a Booster-specific mannequin prior. With simply 30 real-robot demos, that prior reached 87.5% success versus 37.5% for an out-of-the-box π0.5, matching π0.5 utilizing half the real-world demonstrations.
Different companions span embodiment corporations (Feagine Robotics), mannequin corporations (Manycore Tech, Dexmal) and industrial automation (Lotus Vehicles, Geely Auto). Axis additionally provides on-chain robotics networks: BitRobot on Solana and OpenRoboto on Bittensor.
Redefining Bodily AI’s Information Basis
“The way forward for Bodily AI isn’t a static dataset you obtain as soon as,” stated Chris Feng, founding father of Axis Robotics. “It’s an engine that retains producing the info the mannequin wants subsequent. Scale will get you broad protection. Variety retains the noise unbiased. The closed loop turns each failure into progress. That’s what compounds.”
Axis was based by researchers from UC Berkeley, CMU, Georgia Tech, and SJTU, alongside serial founders who’ve scaled client platforms to over 30 million customers. Its analysis is suggested by Jiachen Li, Assistant Professor at Georgia Tech.
Paper Hyperlink: https://arxiv.org/abs/2607.21588
Venture Web page: https://axisaiorg.github.io/AXIS-V1/
Dataset Hyperlink: https://huggingface.co/datasets/axisrobotics/Franka-Dataset
Github Codebase: https://github.com/AxisAIOrg/Axis-V1-Coaching








