Embodied AI research that ends in a robot.

Synthra Robotics Research is the laboratory program behind the U1 Series and Synthra’s custom robotic systems. It studies embodied intelligence as ten open questions, from vision-language-action models to multi-robot coordination, and every question has a route into a robot that has to work in a room full of people.

The research program on the Synthra intelligence loop. Each theme sits at the part of the loop it strengthens; embodied intelligence spans all of it. How Synthra robots see, understand, decide, act, and adapt

What does Synthra Robotics research?

Synthra researches embodied intelligence as ten open questions organized around the loop every Synthra robot runs: see, understand, decide, act, adapt. The themes are embodied intelligence, multimodal robotics, spatial intelligence, vision-language-action models, autonomous navigation, multi-robot coordination, robotic control, human-robot interaction, adaptive autonomy, and physical AI. Work that holds up moves into the Synthra intelligence platform.

Every theme below is an open question, not a published result. Findings and reports appear on this page when they hold.

What does a robot need to understand to act well in a place it has never been?

Embodied intelligence is intelligence that lives in a body: it perceives, remembers, and acts in one real place at a time.

Intelligence judged by what a body does in the world. We study how perception, memory, and action combine into competence that survives contact with real venues.

Theme
Embodied intelligence
Loop
Whole loop
Status
Open question
Figure
Illustrative
Arrive, orient, act. The same room at three moments: raw perception, formed context, decided action. Illustrative figure, not measured data.

Under investigation

  • How much of an unfamiliar venue a robot has to observe before it can begin service.
  • Which parts of competence carry from one venue to the next, and which have to be learned again in each room.
  • How perception, memory, and action can share one representation instead of passing messages down a pipeline.

Where it lands

Every U1 model runs the same loop, so what this work establishes reaches U1e, U1, and U1 Max together.

One intelligence platform, three physical configurations

Which signals should a robot trust when they disagree?

Multimodal robotics combines several kinds of sensing, such as sight, sound, and the robot’s own motion, into one understanding of the world.

Fusing vision, sound, and the robot’s own motion into one estimate of the world, with uncertainty that downstream decisions can use.

One estimate from three signalsThree sensor rows estimate where a guest is standing: vision, audio, and the robot’s own motion. When they agree, the fused estimate is narrow and the planner approaches at normal speed. When glare corrupts vision, its weight drops, the fused estimate widens, and the planner reduces speed and widens its margin.ESTIMATE · GUEST POSITIONWEIGHTVISIONGLAREAUDIOOWN MOTIONFUSEDPLANNER · APPROACH AT NORMAL SPEEDPLANNER · REDUCE SPEED, WIDEN MARGIN
  • Estimate and uncertainty
  • Corrupted signal
  • Planning decision
Theme
Multimodal robotics
Loop
See
Status
Open question
Figure
Illustrative
One estimate from three signals. Fusion that reports its own doubt. Switch to see a disagreement. Illustrative figure, not measured data.

Under investigation

  • How to weigh vision, sound, and the robot’s own motion when glare, noise, or clutter corrupts one of them.
  • How to carry uncertainty from perception into planning, so doubt changes how the robot moves rather than whether it moves.
  • How to tell a real change in the room from a sensor error.

Where it lands

Feeds the multimodal perception of the Synthra intelligence platform, which every U1 model uses to read people, tables, objects, and pathways as the scene changes.

Inside the Synthra intelligence platform

What is where, and what is it for?

Spatial intelligence is a robot’s understanding of a space as both geometry and purpose: where things are, and what each area is for.

Semantic maps that hold meaning as well as geometry: which surface is a table, which zone is service-only, where a robot should wait, and where it should never stop.

Geometry, then meaningA restaurant floor plan shown two ways. As geometry, it is a grid of occupied, free, and unknown cells. As meaning, the same plan is labelled with eight tables, the dining area, a service-only zone, the kitchen pass, a waiting spot, the dock, and two no-stopping zones at the entrance and the kitchen door.DININGTABLE 01TABLE 02TABLE 03TABLE 04TABLE 05TABLE 06TABLE 07TABLE 08SERVICE ONLYPASSWAIT HEREDOCKNEVER STOP · ENTRANCENEVER STOP · DOOR SWING
  • Occupied
  • Free
  • Unknown
Theme
Spatial intelligence
Loop
Understand
Status
Open question
Figure
Illustrative
Geometry, then meaning. A map that knows what things are for. Switch between geometry and meaning. Illustrative figure, not measured data.

Under investigation

  • How to attach meaning to a map, such as table, service-only zone, or waiting spot, from observation and the venue’s own configuration.
  • How a semantic map stays correct when furniture moves between services.
  • Where a robot should wait, and where it should never stop: entrances, kitchen door swings, and busy crossings.

Where it lands

Feeds the spatial understanding a U1 robot builds when it explores a venue during deployment, before service begins.

What a Synthra robot understands about a restaurant

How does “bring me the pasta” become a trajectory?

A vision-language-action model connects what a robot sees, what it is asked, and what it does: it grounds a request in the scene and turns it into actions the robot can carry out.

Models that ground language and vision in physical action, so a spoken request resolves into objects, places, and a plan the robot can execute.

From words to a pathThe spoken request “bring me the pasta” is split into two references. “Me” is grounded to the guest seated at table four and “the pasta” to an order waiting at the kitchen pass. A decided path runs from the pass to table four. A band beneath reads language, then objects and places, then action.GUEST“Bring me the pasta.”ORDER · PASTAGUEST · TABLE 04LANGUAGEOBJECTS AND PLACESACTION
Theme
Vision-language-action models
Loop
Understand
Status
Open question
Figure
Illustrative
From words to a path. Grounding: every word that matters is tied to something the robot can find. Illustrative figure, not measured data.

Under investigation

  • How references like “the pasta”, “that table”, or “the same again” resolve to specific objects and places.
  • How a request, the menu, and the table context combine into one task the robot can execute and check.
  • When to ask a short clarifying question instead of guessing.

Where it lands

Feeds conversational ordering on U1: a guest’s words become an order, a table, a destination, and an execution plan.

How a conversation becomes a service task

How do you move through a crowd without asking it to move for you?

Autonomous navigation is a robot’s ability to choose its own path to a destination and change it as the space changes, without a predefined route.

Planning in dense, changing, human spaces: paths that are efficient, predictable to the people around them, and re-decided the moment the scene changes.

A path people can readTop-down view of a busy aisle with seven people, four of them walking, each with a cone showing where they are likely to move. The shortest path to table six cuts between two people walking toward each other and is marked as a conflict. The decided path keeps to one side and follows behind a server, so it stays clear and easy to anticipate.TABLE 06CONFLICT
  • Predicted motion
  • Personal space
  • Shortest path
  • Decided path
Theme
Autonomous navigation
Loop
Decide
Status
Open question
Figure
Illustrative
A path people can read. Short is not the same as good. The decided path is the one people nearby can predict. Illustrative figure, not measured data.

Under investigation

  • How to predict where people are heading well enough to pass without making them adjust.
  • How to choose paths that the people nearby can anticipate, not only paths that are short.
  • How often to re-decide a route so the robot reacts quickly without looking hesitant.

Where it lands

Feeds contextual navigation in every U1 model: routes decided at the moment of service and re-decided when people, chairs, or traffic move.

Why a Synthra robot needs no predefined route

When does a group of robots become one system?

Multi-robot coordination plans several robots’ tasks, paths, and charging together, so a fleet shares one understanding of the venue instead of working in isolation.

Task allocation, shared environmental understanding, traffic management, and charging coordination across fleets that work the same floor.

Three robots, one planPlan of a large food hall with three robots. Service requests at several tables are assigned across them. Robot 01 has reserved a narrow passage, and robot 02 holds before it, yielding. Robot 03 heads to the dock to charge. An inset shows the shared map in which all three robots appear, beside a task board of their assignments.PASSAGE APASSAGE BT 12REQUEST · T 21REQUEST · T 14REQUEST · T 23RESERVED · ROBOT 01HOLD · YIELD TO 01TO DOCK · CHARGE WINDOWROBOT 01ROBOT 02ROBOT 03SHARED MAPTASK BOARD01 · T 2102 · T 14,THEN T 2303 · DOCK
  • Service request
  • Decided path and reservation
  • Hold
  • Charging
Theme
Multi-robot coordination
Loop
Decide
Status
Open question
Figure
Illustrative
Three robots, one plan. Allocation, traffic, and charging decided together. Illustrative figure, not measured data.

Under investigation

  • How to assign incoming requests across robots by position, workload, and remaining charge.
  • How robots keep one shared picture of the venue current as each of them perceives changes.
  • How to reserve narrow passages and charging time so robots don’t block each other during a rush.

Where it lands

Feeds fleet coordination on U1 Max: distributed task assignment, robot-to-robot communication, traffic management, and charging coordination.

Fleet coordination on U1 Max

How do you carry a full glass through a moving room?

Robotic control turns a planned motion into the velocities and forces the robot actually executes, smoothly and stably, whatever it is carrying.

Turning decisions into smooth, stable physical motion, with payload-aware acceleration, precise docking, and recovery from disturbances.

The same move, commanded two waysTop: two velocity profiles over time, a sharp step and a smooth shaped curve. Middle: a glass carried under each profile; with the step, the liquid tilts up to the spill line, and with the shaped curve it stays nearly level. Bottom: a robot approaching its dock, its sideways offset shrinking until it aligns and connects.VTSTEP COMMANDPAYLOAD-AWARESPILL LINESTEPSHAPEDDOCK APPROACHOFFSETCONNECTED
  • Step command
  • Payload-aware profile
  • Liquid
  • Connected
Theme
Robotic control
Loop
Act
Status
Open question
Figure
Illustrative
The same move, commanded two ways. Shape the motion to the load, and the glass stays level. Axes are unscaled. Illustrative figure, not measured data.

Under investigation

  • How to shape acceleration and turning to the load on board, from a single coffee to a full order.
  • How to recover from a nudge or a sudden stop while keeping the load steady.
  • How to align with a dock precisely enough to connect on the first approach.

Where it lands

Feeds how U1e carries drinks and coffee through narrow service areas, and how every U1 model returns to its dock and connects on its own.

U1e, the compact model for narrow service areas

How does a robot make its intent obvious to a guest who never read a manual?

Human-robot interaction studies how robots and people communicate and share space, through speech, motion, and timing.

Conversation, legible motion, and service etiquette: how robots announce, yield, wait, and hand over in spaces designed for people.

Theme
Human-robot interaction
Loop
Act
Status
Open question
Figure
Illustrative
Announce, yield, hand over. Intent shown through timing and motion. Illustrative figure, not measured data.

Under investigation

  • Which motion cues let guests predict a robot’s next move: slowing early, turning before moving, pausing where it can be seen.
  • How a robot should announce itself, yield, and wait at a table without interrupting a conversation.
  • How to hand over food and drinks so the guest knows what to take, and when.

Where it lands

Feeds how every U1 model approaches a table, talks with guests, and hands over an order.

U1, the general-purpose hospitality robot

What should change when the environment does?

Adaptive autonomy is a robot’s ability to keep re-evaluating its goals, plans, and constraints while it works, so a change in the room changes the plan instead of stopping it.

Continuous re-evaluation of goals, plans, and constraints, so autonomy does not stop when a chair moves, a door closes, or a floor fills up.

The plan follows the roomA floor plan and a list of constraints, before and after the room changes. Before, the robot heads to table seven through aisle B. After, a chair blocks aisle B, the kitchen door is closed, and aisle A is crowded. The goal stays the same, the route is re-decided through aisle C, and speed is reduced.TABLE 07AISLE AAISLE BAISLE CBLOCKEDGOALTABLE 07TABLE 07 ·KEPTROUTEAISLE BAISLE C ·RE-DECIDEDAISLE BOPENBLOCKEDKITCHEN DOOROPENCLOSEDFLOORQUIETBUSY ·SPEED REDUCED
  • Current plan
  • Previous plan
  • Changed constraint
Theme
Adaptive autonomy
Loop
Adapt
Status
Open question
Figure
Illustrative
The plan follows the room. Goal kept, route re-decided, constraints updated. Switch between before and after. Illustrative figure, not measured data.

Under investigation

  • Which changes should alter the goal, which only the path, and which nothing at all.
  • How to tell a lasting change, like a rearranged floor, from a passing one.
  • When a robot should ask staff for help instead of finding its own way around a problem.

Where it lands

Feeds how every U1 model keeps working when a chair moves, a door closes, or the floor fills up.

How the loop adapts as the room changes

What can only be learned by acting?

Physical AI is artificial intelligence that learns from acting in the physical world, rather than from text or images alone.

Learning from interaction with the physical world, and the data, simulation, and evaluation that such learning requires.

Theme
Physical AI
Loop
Adapt
Status
Open question
Figure
Illustrative
Act, record, vary, evaluate. Learning that starts and ends on a real floor. Illustrative figure, not measured data.

Under investigation

  • Which skills need real interaction data, and which can be learned in simulation first.
  • How to turn one real-world trace into many useful simulated variations.
  • How to test a new behavior before it reaches a live floor, with checks that predict how it will hold up there.

Where it lands

Feeds how the Synthra intelligence platform is trained and tested, which custom systems from Synthra Robotics Engineering build on as well.

Custom robotic systems from Synthra Robotics Engineering

Research is finished when a robot can do it.

How does Synthra’s research become a product?

Synthra’s research reaches its robots through one shared platform. Each question starts from a problem a robot will meet in a working venue and is answered with models, data, and evaluation. What holds up moves into the Synthra intelligence platform, which U1e, U1, U1 Max, and Synthra’s custom systems all share.

From question to robot

  1. Question
    A problem a robot will meet in a working venue, stated precisely enough to test.
  2. Method
    Models, data, and evaluation built to answer it.
  3. Platform
    What holds up becomes part of the Synthra intelligence platform.
  4. Robots
    U1e, U1, U1 Max, and custom systems share that platform, so they share the result.

Where each theme lands

Each research theme and the platform capabilities it feeds. Platform capabilities are shared by every U1 model; fleet coordination is specific to U1 Max.

Research themes and the Synthra platform capabilities they feed
ThemeMultimodal perceptionVision-language-actionSpatial understandingContextual navigationConversational serviceAutomatic dockingFleet coordination U1 Max
Embodied intelligence Whole loopFeedsFeedsFeedsFeedsFeedsFeedsNo
Multimodal robotics SeeFeedsNoNoNoFeedsNoNo
Spatial intelligence UnderstandNoNoFeedsFeedsNoFeedsNo
Vision-language-action models UnderstandNoFeedsNoNoFeedsNoNo
Autonomous navigation DecideNoNoNoFeedsNoFeedsNo
Multi-robot coordination DecideNoNoFeedsNoNoFeedsFeeds
Robotic control ActNoNoNoFeedsNoFeedsNo
Human-robot interaction ActNoNoNoFeedsFeedsNoNo
Adaptive autonomy AdaptNoNoFeedsFeedsNoNoNo
Physical AI AdaptFeedsFeedsNoFeedsNoNoNo

The rules the lab works by.

Real rooms set the bar.
Crowds, clutter, noise, and layouts that change between services are the conditions a method has to survive, so they are the conditions it is judged against.
Uncertainty is an output.
Every estimate should carry how sure it is, so planning can slow down or widen a margin instead of acting on a guess.
Legible beats clever.
People near the robot should be able to predict its next move without reading anything. A behavior that surprises guests is not finished.
Publish when it holds.
Results and technical reports are released when they are validated. Until then, this page lists open questions, not findings.

What makes a good research collaboration with Synthra?

A good collaboration starts from a question that needs a real robot in a real, changing environment to answer. Collaborators bring a sharp question and the expertise to pursue it. Synthra brings the U1 Series, the Synthra intelligence platform, and the engineering to move a method into a robot. Each collaboration is scoped together before either side commits.

What to include

  • The question you are working on, in a sentence or two.
  • Your group and its focus.
  • Why the question needs a physical robot, and what environment it needs.
  • What a collaboration would need: robots, environments, or shared work.

Inquiry · Research

Preview
Path
Research Collaboration
Institution
Your institution or lab
Area
Brief
The question you’re working on, and what a collaboration would need.
  1. Draft
  2. Queued
  3. En route
  4. Delivered
Propose a research collaboration

Opens the contact form with Research Collaboration selected. Nothing is sent until you submit it there.

Questions about Synthra research

Direct answers for researchers, partners, and anyone checking what Synthra works on. To propose work together, use the contact page.

Can research institutions collaborate with Synthra?

Yes. Synthra is open to collaboration with universities, labs, and research teams on problems where real robots in real environments are the bottleneck. Choose Research Collaboration on the contact page and describe the question you are working on and what a collaboration would need.

Does Synthra publish its research?

Synthra publishes results and technical reports when they are validated, and lists them on the research page as they are released. Until then, the page describes the program as open questions rather than findings, so nothing on it should be read as a published result or benchmark.

What is spatial intelligence in robotics?

Spatial intelligence is a robot’s understanding of a space as both geometry and purpose. Beyond knowing where walls, tables, and people are, the robot knows what each area is for: which surface is a table, which zone is service-only, where it should wait, and where it should never stop.

What is adaptive autonomy?

Adaptive autonomy is a robot’s ability to keep re-evaluating its goals, plans, and constraints while it works. When a chair moves, a door closes, or the floor fills up, an adaptive robot changes its route or its timing and carries on with the task instead of stopping until the original plan works again.

What is physical AI?

Physical AI is artificial intelligence that learns from acting in the physical world rather than from text or images alone. It needs its own infrastructure: data recorded from real robots, simulation that can vary real scenes, and evaluation that predicts how a behavior will hold up on a live floor.