RoboStudio
The studio behind our robotics annotation work — vision, 3D, and force/torque/tactile in one tool. No platform we surveyed combines all three; that gap is what we built RoboStudio to close.
Vision tools ignore touch.
Every open-source annotation tool we evaluated does 2D, video, or 3D well. None of them label force, torque, or tactile contact — the signals that determine whether a grasp actually worked, not just whether it looked right on camera.
RoboStudio is built on CVAT — open-source, MIT-licensed, and genuinely good at vision and 3D — extended with a custom signal workspace for the one thing it doesn't do. That extension, not the annotation UI around it, is the part nobody else has built.
| Vision & 2D/video | Common |
| 3D / point cloud | Several platforms |
| Force / torque / tactile | None we found |
What Runs in RoboStudio
Eight labeling disciplines spanning perception, manipulation, and physical safety — the same capabilities described on our Robotics Annotation page, run through this tool.
Multi-sensor ingest
Synchronised LiDAR, depth camera, and IMU streams brought onto one timeline before any label is applied.
3D scene & pose
Object pose estimation, grasp-point labeling, and scene-graph annotation of spatial relationships.
Trajectory & kinematics
Joint-angle sequences and end-effector paths, sharing the same episode schema as our agentic work.
Human demonstration
Teleoperation and kinesthetic teaching capture, labeled for imitation learning.
Sim-to-real validation
Synthetic data from Isaac Sim, MuJoCo, or Gazebo checked against real-world behavior.
Safety & collision
Human-robot interaction safety-zone violations, mapped to ISO 10218 and ISO/TS 15066.
Temporal segmentation
Continuous operation logs segmented into discrete labeled sub-tasks.
Physical-world ontology
A standardized object, material, and affordance taxonomy — 13 terms and growing in production use.
Force/torque/tactile — no open-source base, built for this specifically
6-DOF force/torque signal labeling, tactile heatmaps, and contact-event detection, with interval-quality scoring that includes onset error — a metric we haven't seen any surveyed platform report. This is the module that answers "did the grasp actually work," not just "did it look right on camera."
Built on CVAT. Extended for what CVAT doesn't do.
RoboStudio runs on CVAT (MIT-licensed) unmodified at the server — we didn't fork the annotation engine, only the workspace UI, for the signal panels this page describes. Every label produced here flows into Datum, the same data platform behind our text, image, and video work, so export and lineage are consistent no matter which studio produced the label.
Have robot data nobody else can label?
If it involves force, torque, or touch alongside vision, talk to us before you talk to anyone else.