Machine readable product instructions

Is your product robot ready?

Axiom3 turns a single human demonstration into a tag anchored 4D instruction manifest. It is a machine readable record of how to use your product that a robot, or an AR app today, can resolve straight from a small printed marker on your label.

Proof of concept · one capture, two destinations

Left: a Fourier GR-1 humanoid performing the captured task in simulation. Its wrists and all five fingers on each hand follow the recorded hands to under a millimetre, and the drawer is driven by that hand — this is the plan a robot inherits. It is solved seated because the demonstration was performed seated. Everything above the wrists is the robot's own solution rather than captured motion: this recording carries hand data only. Right: what a user in AR glasses is shown — ghost hands demonstrating the task on the product itself, anchored to the printed tag. Hand pose sits within a median 1.4 cm of an independent measurement, which is the gap a robot's own sensors close on approach. Two steps, left hand then right, segmented from the capture rather than authored by hand. Best viewed full screen.

Your product already has instructions on the back for a human. Axiom3 adds the version a machine can read. We film an expert using your product once, and produce a 4D dataset that captures the exact hand motions, step by step, anchored to a small printed tag.

Today, a customer points their phone or headset at it and watches ghost hands show them exactly what to do, registered onto their real counter. Tomorrow, general purpose robots read the same tag and already know how to use your product. The tag is a permanent hook: you author it once, and it gets more valuable as the robots get better.

How it works

One demonstration in. A permanent, machine readable manifest out.

1

Capture once

An expert performs the task one time while we record the full motion, with a small printed tag placed near the product. We reconstruct both hands in precise 3D through the entire task, anchored to that tag.

2

Anchor to the product, not the room

Every motion is re-expressed in the coordinate frame of the printed tag. The room, the lighting, and where the table sat never travel with the data. Only the relationship between the hands and the product does. That's why the same tag works in any kitchen, any factory, anywhere.

3

Segment into steps

The continuous motion is cut into named steps with a bill of materials ("cake mix also needs 3 eggs, ½ cup oil, 1 cup water") and marked contact moments such as grasp, pour, and stir.

4

Publish a manifest

The result is a standard JSON manifest plus an animated 3D file. You print the tag on the label; the tag id resolves to the hosted manifest. One tag, one URL, re-publishable as the product changes.

5

Consume it two ways

A human watches ghost hand AR guidance today. A robot reads the same tag as a manipulation prior tomorrow. Same asset, two audiences.

Two horizons

Pays for itself today. Appreciates as robots arrive.

Today · Humans

Ghost-hand AR guidance

An AR app on a phone, Quest 3, or Vision Pro anchors ghost hands to the tag and walks the user through the task in place, registered onto their real environment. The instruction playback path exists and works today. Try the live viewer →

The recorded hands projected onto footage of the same dresser wearing only the tag no gloves, no trackers, nobody working. The tag was solved in every frame of the clip.

Tomorrow · Robots

A manipulation prior

A manipulation capable robot reads the tag, fetches the manifest, and uses the step plan, reference trajectories, and object frames as a prior for its own policy. It refines the last centimetre with its own sensors. You supply the plan and the intent, which is the expensive part to author per product.

Why it's valuable before the robots arrive

A moat on a 2-3 cm footprint.

📱

Pays for itself today

Premium AR "how-to-use-it" content drives measurable engagement and fewer support calls, all on the same label real estate you already own.

🔒

A moat, printed once

A 2-3 cm tag is a permanent, machine readable hook. Author it once and you're robot ready the day the hardware ships, while competitors are still filming.

📈

Improves without re-shooting

As retargeting and robot policies advance, the same captured manifest yields better robot behaviour. The asset appreciates.

🎨

It doesn't have to be visible

If a printed square is wrong for your packaging, the same marker can be produced in infrared absorbing ink: invisible to the eye, sharp to a camera that sees near infrared. The trade is on the reading end a stock phone camera has an infrared cut filter and will not see it, so an invisible tag suits robots and purpose built readers, which carry infrared capable cameras anyway. Visible tags stay the right call for phone based AR. We will tell you which fits your product rather than sell you the option.

The alternative

Why not just teleoperate a robot?

Manufacturers ask this, so here is the difference in plain terms. Teleoperation was built to teach a robot general manipulation. Axiom3 was built to document your product.

Teleoperation

Drive a humanoid through the task

An operator puppets a robot while the robot records itself.

  • You need the robot. A humanoid on site, plus the people to run it, for as long as collection takes. There is no version of this approach without one.
  • Your product is nowhere in the data. Everything is logged in room coordinates. Move the unit, or pull a different one off the line, and the recording no longer points at anything.
  • It records an operator, not a user. What gets captured is a person compensating for latency, a reduced-finger hand and no sense of touch, rather than how your product is actually handled.
  • It is tied to the robot that recorded it. The data describes one machine's joints, so the next generation of hardware means collecting all over again.
  • Every SKU starts from zero. New sessions, on the robot, with an operator, for every product in the line.
  • Nothing in it reaches your customers. It produces robot training data. Your buyers, your installers and your support line see none of it.
Axiom3

Capture a person using your product

Filmed once, anchored to a printed tag on the product itself.

  • No robot required. You ship us the unit. That is the whole hardware ask on your side.
  • Everything is measured against your product. The printed tag is the origin, so every hand position is expressed relative to the product rather than the room. Move it, replace it, or ship it to a customer and the manifest still refers to the right thing.
  • The motion is a real one. Somebody actually operating your product, at their own speed, with ten working fingers.
  • It outlives any one robot. The capture describes the task, not a machine, so it does not expire when the next generation of hardware ships.
  • Two audiences, one session. The same capture drives the AR view your customers open on a phone today and the reference a robot reads off the tag tomorrow.
  • One session per product, at one fixed price. Which is the only reason documenting a whole catalogue is affordable at all.

Your product lands in one of three complexity bands, and the band is the price. See the pricing →

What we deliver

Everything needed to make one product robot-ready.

  • A printed anchor-tag spec for your label, with the exact print size and quiet zone.
  • A versioned instruction manifest with the anchor, units, frame, step segmentation, bill of materials, contact events, and measured accuracy.
  • An animated 3D hand file (.glb) in the tag frame, ready for AR playback.
  • Hosting and a tag-id to manifest resolver, so the label stays static while the content can update.
Straight talk on accuracy

We ship measured accuracy, not a promise.

Every manifest ships with its own measured accuracy, split into free space motion (a few centimetres, which is fine) and tighter contact event accuracy relative to the object. A robot consuming the manifest does its own final centimetre sensing, while Axiom3 supplies the plan, the sequence, and the intent. We tell you exactly what the data is and isn't.

From the capture shown in the video above
1.4 cm
Hand pose vs. an independent measurement
median, better observed hand. 3.7 cm on the hand that faced away from the camera half the time.
0.9 cm
Hand-to-contact-point distance
RMS for that same hand against the same independent measurement, correlation 0.998
0.4 cm
Camera-to-tracker calibration
median over 207 poses, 0.73 px reprojection our tightest to date
Best measured to date, on an earlier capture with the same rig
1.0 cm
Per-fingertip error, held out from the solve
median over 96 samples, p90 3.5 cm, better observed hand. The other hand ran 1.5 cm median on that capture, with a longer tail from a tracker fault we have since fixed.
Per capture
Every manifest carries its own numbers
We do not publish one headline figure and apply it to your product. Yours gets measured, and you see it.

Read these as the gap your robot's own sensors need to close. The manifest puts the hand within a couple of centimetres of the product and tells it what to do next; the robot's cameras and force sensing handle the final contact.

Next step

Is your product robot-ready?

Send us your flagship product and we film an expert using it once. You get the manifest, the 3D playback file, and the printable anchor tag for your label.