Overhead long-exposure of an outdoor pool: the robotic fish shown at seven instants along the figure of eight it swam, with its full tracked path traced behind it, colored from amber at the start through to deep blue at the end, as it pursued a target moved by hand.

All You Need Is Low Fidelity

Zero-Shot Sim-to-Real of Learned Robotic Fish Control

Soft Robotics Lab, ETH Zurich  ·  *Equal contribution  ·  †Also ETH AI Center

Abstract

Complex tasks for underwater robots remain limited by the capabilities of their controllers. Learning a better one for a soft, underactuated robotic fish trades simulator cost against fidelity. We show that an intentionally low-fidelity simulator is enough: a stateless, quasi-steady fluid model with no wake and no added-mass history suffices to learn a general, closed-loop controller that transfers to hardware without tuning.

Our platform is a soft, single-motor, tendon-driven fish whose policy observes only what the hardware can measure. A staged pipeline grounds the simulator in two independent identifications, fixing the tail dynamics and a stateless fluid model; the policy then acts through a band-limited rhythmic trajectory generator rather than commanding the tail directly. Deployed unchanged in an outdoor pool, a single policy performs closed-loop target reaching, disturbance rejection, and out-of-distribution target acquisition and tracking. The transfer rests on the constraint rather than the fidelity: the generator cannot leave the band over which the fluid was identified. This raises the question of how much of the physics can reside in the controller rather than in the simulator.

Watch the Supplementary Video

Our Platform: An Underactuated Soft Tendon-Driven Fish

Cutaway of the hardware fish, labelled: the compliant PLA tail, the antagonistic tendon pair, the tendon pulley, the Dynamixel motor, the Adafruit IMU, the Raspberry Pi 5 and the AprilTag on its back. AprilTag Compliant PLA Tail Raspberry Pi 5 Antagonistic Tendon Pair Tendon Pulley Dynamixel Motor Adafruit IMU
The underactuated soft fish is largely FDM-printed in ordinary PLA. A thin printed tail is driven by a pair of antagonistic tendons routed down either side and pulled by a position-controlled Dynamixel motor. The tail is underactuated: one actuator drives a continuum whose shape emerges from the interaction of tendons, motor and fluid. A sealed oil-filled head carries the onboard compute, where a Raspberry Pi 5 runs the policy. A motor encoder, an Adafruit BNO055 IMU and a 3D-printed AprilTag provide its observations.

How It Works

Overview diagram: separate identification experiments fix the tail dynamics and a stateless fluid model; a policy sets the parameters of a band-limited trajectory generator that emits the motor setpoint; the same stack runs unchanged onboard the real fish, closed by hardware observations from an overhead camera. System Identification Tail dynamics Fluid model t y Reinforcement Learning Policy u* Trajectory generator Zero-Shot Deployment θ Onboard (50 Hz) Hardware observations
Separate experiments identify the tail dynamics and a stateless fluid model. The policy sets the parameters of a band-limited trajectory generator that emits the motor setpoint. The same policy and generator run unchanged onboard the real fish, closed by hardware observations.

A Constrained Action Space

Two motor-angle traces over three seconds, commanded against measured. Direct action collapses into a small fast flutter; the band-limited trajectory generator tracks a clean low-frequency stroke. Below them, the twin's transmission -- pulley, crossed tendon pair and five links on torsional springs -- driven by each in turn. 2026-10-08T18:33:36.054451 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ −1 0 1 DA — Direct action 0 1 2 3 Time (s) −1 0 1 TG — Trajectory generator Commanded Actual Motor angle (rad) TendonsRigid tail sections
Unconstrained learned direct action collapses into a small, fast flutter, and the tail cannot track what it is asked for. The band-limited trajectory generator holds a clean, low-frequency stroke instead: the policy sets its parameters rather than commanding the tail directly. The generator's frequency ceiling is the band the fluid was identified over, so the action space cannot leave the region the model was validated on. That constraint, not the simulator's fidelity, is what carries the policy to hardware.

Policy Trained in a Stateless Fluid Simulation

The training curriculum in three phases. Phase one introduces sensor noise: regular swimming, then encoder and IMU noise, then AprilTag noise. Phase two randomises the domain: initial velocity, body and tail, and the fluid coefficients. Phase three widens the target distribution from plus or minus sixty degrees to plus or minus eighty-five, shown as a polar wedge reaching 2.5 metres. Phase 1 Sensor Noise Phase 2 Domain Randomization Phase 3 Target Widening 1 LR 1× Regular Swimming 2 LR 0.3× Encoder & IMU Noise 3 LR 0.09× AprilTag Noise 4 LR 1× Initial Velocity DR 5 LR 0.3× Body & Tail DR 6 LR 0.09× Fluid Coefficient DR 7 LR 1× ±60° 8 LR 0.3× ±85° 9 LR 0.09× ±120° 1 m2 m
Training runs as a curriculum, all of it inside the low-fidelity simulator, a quasi-steady fluid with no wake and no added-mass history. The policy learns to swim first, on clean observations and from rest. Sensor noise is only introduced once it can, the encoder and IMU before the AprilTag it navigates by. The domain is randomized next, over initial velocity, the body and tail, and the fluid coefficients. The spread of targets is widened last, from ±60° out to ±120°, so the turns that demand the most of the generator are asked for only when everything else is already in place.

Zero-Shot Deployment

Polar plot of all thirty static-target deployment runs in the start frame of the fish, colored by outcome, with the trained region shaded. Targets behind the fish and beyond the trained range are still reached; the misses concentrate in the lateral-rear band. 2026-10-08T18:33:38.293705 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 0° 30° 60° 90° 120° 150° 180° −150° −120° −90° −60° −30° Reached (20/30) Undisturbed run Missed (10/30) Disturbed run 1 m 2 m 3 m
Every static-target run in the start frame of the fish. The shaded wedge is the trained region. Targets behind the fish and beyond the trained range are still reached; misses concentrate in the lateral-rear band, where the turn required is tighter than the generator can close in a pool this size. Not one parameter was retuned for hardware.

The Policy Rejects Untrained Disturbances

Disturbed Fish 00:00.00 1×
Moved Target 00:00.00 1×
Dynamic Target Tracking 00:00.00 3×
Introduced Initial Velocity 00:00.00 2×
None of these disturbances were trained for: the policy never saw turbulence, a target that moves, or a push off the line. Recovery comes free because the state it is given says where the target is now, not how it got there. A shove is simply a new starting state. Each clock runs on the recording, not the playback, so the sped-up runs still read in real seconds.

The Policy Stays Where the Simulator Is Honest

A quasi-steady fluid computes thrust from instantaneous velocity alone, so a slow large stroke and a fast small flutter at the same speed look identical to it. An unconstrained policy finds the flutter, and no real, unsteady fluid can supply it.

The trajectory generator makes that gait unreachable: its ceiling is the band the fluid was identified on. Nearly all of the generator's tail-velocity power falls inside that band; most of unconstrained direct action's falls outside.

2026-10-08T18:33:42.926768 image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 0.5 1 2 5 10 20 Tail frequency (Hz) 0.0 0.2 0.4 0.6 0.8 1.0 Cumulative tail velocity power fraction Identified band Identification (92%) Trajectory generator (97%) Deployed hardware (62%) Direct action (44%)

BibTeX

@misc{maloney2026lowfidelity,
  title         = {All You Need Is Low Fidelity: Zero-Shot Sim-to-Real of
                   Learned Robotic Fish Control},
  author        = {Maloney, Liam and Ramchandani, Simon and Michelis, Mike Y.
                   and Hinchet, Ronan and Katzschmann, Robert K.},
  year          = {2026},
  eprint        = {2609.36993},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  note          = {Under review at the IEEE International Conference on
                   Robotics and Automation (ICRA) 2027}
}