Synchronized head-locked first-person video, world-view video, and IMU telemetry // perception training · robotics · driver monitoring · wearable AI
RT-Fusion delivers the human perspective data that no vehicle, robot, or fixed camera can generate — synchronized head-locked first-person video, world-view video, and 200Hz IMU telemetry. For autonomous driving, robotics, driver monitoring, and wearable AI. EU-native acquisition across 6 countries.
Operational Capacity: 4h+ continuous World-View (GoPro 5.3K) combined with POV-View clips (Ray-Ban Meta, up to 3 min each) — audio-synced.
BUILT FOR:
Whether you build for autonomous vehicles, robotics, driver monitoring, or wearable AI — your model needs human perspective data that no vehicle, robot, or fixed camera can generate. RT-Fusion captures this signal at human eye level, across real European environments, on-demand.
Bounding boxes don't show intent. What a pedestrian can see before a crossing decision is invisible to vehicle and robot sensors. RT-Fusion captures that first-person view and head orientation at human eye level.
A robot can record its own telemetry, but it can't record what a human body does while performing the same task. Head-locked first-person video and IMU from real human activity — the training data imitation learning needs.
IR cabin cameras track the driver's eyes, not the road the driver sees. RT-Fusion captures the driver's first-person view and head orientation, paired with the road scene, on real European roads.
GoPro-on-helmet footage has different perspectives, motion, and social context than smart glasses. RT-Fusion captures on the exact Ray-Ban Meta device wearable AI products ship on.
A pedestrian wears the rig in real traffic. The Ray-Ban Meta captures what they see and where their head turns before stepping off a kerb — head checks, orientation, social signaling. The GoPro captures synchronized world-view context with 200Hz IMU and GPS. Together, this produces the intent signal that no vehicle-mounted or robot-mounted sensor can generate: what the human was facing before they acted. Delivered as time-stamped MP4 + GPMF telemetry, loadable in PyTorch or OpenCV and convertible to ROS 2 bag.
A person wears the rig while performing physical tasks — climbing stairs, opening doors, navigating terrain, carrying objects. The Ray-Ban Meta captures first-person POV with natural head movement. The GoPro captures synchronized wider context with 200Hz IMU telemetry. This produces exactly what imitation learning, behavior cloning, and VLA architectures need: egocentric human demonstration video showing not just what the person did, but what they were facing while doing it. No robot fleet can generate this data — it requires a human performing the task.
The driver wears Ray-Ban Meta glasses while driving real routes. The glasses record the driver's first-person view and head orientation — head turns, shoulder checks, attention shifts at intersections and roundabouts. The GoPro on the dashboard captures the road scene at the same time. Together, this produces driver POV and head-orientation data paired with traffic context, on real European roads, at a fraction of instrumented vehicle rig cost. Head orientation only — no eye tracking.
A person wears the Ray-Ban Meta doing everyday activities — navigating city streets, shopping, commuting, socializing. The glasses capture first-person POV from the exact consumer device that Meta and its ecosystem partners are building for. This is not GoPro-on-helmet footage — the perspective, motion patterns, social context, and field of view match the actual user experience of smart glasses. The synchronized GoPro adds wider context with 200Hz IMU and GPS telemetry for spatial grounding.
Chest-mounted GoPro (world-view, continuous) + head-worn Ray-Ban Meta (POV-view, clips up to 3 min), running in parallel with audio-synced timestamps.
HARDWARE: GOPRO HERO 13 (CUSTOM ACQUISITION RIG)
Captures the "World Model." High dynamic range handles the "Tunnel Exit" blinding light problem. Rolling shutter stress-tests VIO pipelines against vibration artifacts.
HARDWARE: RAY-BAN META GEN 2
Captures the 'Agent Model.' Solves the High-Density VRU problem by recording the eye-contact negotiation and intent signaling that LiDAR cannot see.
7 scenario folders, each containing synchronized paired sensor output. Raw sensor data — no stabilization, no grading. Each folder is relevant to multiple buyer use cases.
RT-Fusion delivers structured, time-synchronized assets. Every frame is mapped to IMU telemetry and paired with head-locked POV footage, in standard formats for machine learning and robotics pipelines.
{
"timestamp_utc": "2026-02-11T09:14:22.045Z",
"frame_id": 4920,
"environment": {
"location": "NL_Amsterdam_Canal_District",
"weather": "overcast_diffuse",
"surface": "asphalt_bike_lane"
},
"telemetry": {
"imu_accel_x_y_z": [0.02, -0.81, 0.15],
"speed_mps": 5.8
},
"sensors": {
"world_cam_file": "GH010492.MP4",
"attention_cam_file": "RM010492.MP4",
"head_pose_proxy": true
}
}
All assets delivered as time-stamped MP4 + GPMF telemetry — standard formats, readable with open-source tooling and convertible to ROS 2 bag.
Works With Standard Engineering Pipelines — ROS 2 & Isaac Via Conversion
CREDENTIALS // METHODOLOGY
ARTY ZUEV
10+ years in professional media production — camera systems, color science, lighting, and post-production — across commercial, documentary, and marketing projects in the EU. When the industry shifted from language models to real-world perception, RT-Fusion identified a critical gap: companies building autonomous systems in Europe had few dedicated, on-demand sources for the human perspective data that no vehicle, robot, or fixed camera can generate. RT-Fusion was built to close that gap — applying professional acquisition methodology to capture synchronized head-locked first-person video, world-view video, and IMU telemetry across autonomous driving, robotics, driver monitoring, and wearable AI use cases.
FROM BRIEF TO PIPELINE-READY DATASET
You specify target scenarios, locations, and environmental conditions. Campaign scoped per acquisition day.
Dual-sensor rig deploys to target location. GoPro 5.3K World-View records continuously (4h+); Ray-Ban Meta POV-View captures clips of up to 3 minutes in parallel, audio-synced.
Time-stamped MP4 + GPMF telemetry, paired with JSON metadata per scene. All clips indexed by scenario category and sensor config.
Load MP4 directly into OpenCV or a PyTorch DataLoader. Parse GPMF telemetry with GoPro's open-source gpmf-parser or gopro2gpx. Convert to ROS 2 bag with a short script.
/// DIRECT ENGINEERING FEED
Direct line to Engineering. No sales agents.
Prefer async? [load email]
— or submit a full brief below:
CONNECTION: TLS-ENCRYPTED // PGP-4096 KEY AVAILABLE FOR SENSITIVE BRIEFS