top of page
cheat_pages1-2-4-1_800dpi.png

Bipedal Robot

Writer: Maysarah Sukkar
Maysarah Sukkar
Nov 3, 2025
5 min read

Updated: Sep 26

During the fall 2025 semester at Columbia University, I began designing and building a bipedal robot inspired by the form of a walking UFO for MECE 4611: Robotics Studio. The goal of this ongoing project is to develop a compact, stable biped capable of dynamic motion and balance, using a mix of mechanical design, embedded control, simulation, and iterative testing.

The robot has two four-degree-of-freedom legs, each with control over hip adduction, hip pitch, knee, and ankle. It is driven by LX-16A bus servos and powered by a Raspberry Pi, with an IMU and foot contact sensors for feedback. Alongside the hardware, I built a full physics model of the robot in MuJoCo and trained a walking policy with reinforcement learning. The UFO-inspired body gives the robot a distinct visual identity while also housing its electronics and power.



Design and Modeling

All mechanical design and modeling were done in Onshape, where I created the full robot assembly, joint geometries, and internal cable routing features. The structure is held together with M3 heat-set inserts and bolts, and includes cutouts for ventilation, wiring, and component access.

  • Modular legs: Each leg was designed for strength and modularity, with removable servos, a replaceable foot module, and belt-driven knees that keep the center of gravity low.

  • TPU feet: Textured TPU soles improve ground grip and absorb impact during motion tests.

  • Sensor dome: A translucent dome houses a Raspberry Pi Camera and NeoPixel lighting, providing both visual flair and onboard sensing.


Prototyping and Printing

The prototype is entirely 3D-printed from a combination of PLA and TPU. Early print iterations focused on refining fit tolerances, bearing housings, and servo mounting brackets.

  • Hip bearings: The bearings weren't seating properly in the hip adductor assembly. I fixed this by adjusting clearances and reprinting the parts at finer layer heights.

  • Knee tensioner: Deepening the gear teeth on the knee servo tensioner eliminated slipping and improved torque transfer.

  • Finishing: Every component was post-processed, sanded, cleaned, and trimmed before assembly.

  • Cable management: Internal slots and channels keep wiring secure and prevent tangling during movement. The battery and Pi mount to the underside of the chassis, with power and data routed through the core body.



Testing and Motion Development

The robot is fully assembled and can perform joint motion tests across all four degrees of freedom per leg. The validated motion ranges are:

  • Hip adductor: 60°

  • Hip raise: 90°

  • Knee: 100°

  • Foot: 140°


These ranges were confirmed through manual jogging and position testing to ensure no components collide or overextend. Early stability tests showed the most balanced pose uses slightly flexed knees, which distributes load evenly across the servos. This finding later shaped the standing pose used in simulation.

  • Boot and shutdown: On startup, the robot moves into a stable standing position. On shutdown, it slowly lowers into a resting crouch to prevent sudden load drops.

  • Homing: Homing uses current-based detection. Each joint moves until it hits a physical limit, records that point, and defines a safe operating range around it. This prevents overtravel damage and simplifies setup for gait experiments.


Foot Contact Sensing

To support closed-loop walking, I added contact sensors to each foot module to detect stance and swing phases. Both the gait controller and the RL policy use this signal to recognize real steps.

  • Hardware: [sensor type, mounting in the TPU foot, how the Pi reads them]

  • Status: The sensors are installed, but their readings are currently unreliable on hardware.

  • Plan: Either repair the sensors, or estimate contact from joint and IMU data and retrain the policy without explicit contact inputs.

Simulation and Reinforcement Learning

Building the Simulation Model

I exported each link from Onshape into a MuJoCo model with matching mass and inertia properties, joint limits, and kinematic structure.

  • Actuators: Position-controlled, with torque limited to ±1.67 N·m to match the servos' 17 kg·cm rating.

  • Contact: Simplified box collision geometry on the feet, with friction tuned to approximate TPU on a hard floor.

  • Timing: 500 Hz physics with a 125 Hz control rate.

Training Setup

  • Algorithm: PPO (Stable-Baselines3). Early runs used 4 parallel environments; the final setup scaled to 384, which made rapid reward iteration practical.

  • Observations (31): Pelvis position, orientation, and linear and angular velocity; 8 joint positions and velocities; 2 foot contact flags.

  • Actions (8): One target per joint.

  • Curriculum: Training advances from balancing to slow walking to moderate-speed walking as reward thresholds are met.

How the Reward Evolved

Most of the lessons came from the policy finding loopholes.

  • Falling with style: A velocity-only reward produced "fast" episodes that ended in about half a second. The policy had learned to lunge forward and fall.

  • Rewarding real steps: I added terms for alternating foot contacts, single-leg support, and staying upright. Forward velocity now only counts while the robot is actually stepping.

  • Too much of a good thing: Patching each new failure mode grew the reward to around 17 terms, including harsh anti-crawling terminations. Reward variance exploded, episodes ended before the policy could explore, and PPO stopped learning.

  • Simplify and smooth: I cut back to a few balanced terms: stepping-gated forward progress, upright posture and height, hip and knee activity favored over ankle shuffling, and a mild penalty for non-foot ground contact. Smooth falloffs replaced the hard thresholds so gradients stay informative.

Fixing the Crouch

Early policies walked in a deep crouch. Fixing it meant tracing problems through the whole pipeline:

  • Joint zero isn't straight. The CAD zero pose has heavily bent knees. A tall stance actually needs about 41° knee, −25° hip pitch, and −15° ankle. I saved this as a MuJoCo keyframe with the pelvis at 0.40 m.

  • Actions relative to the stance. With absolute joint targets, an untrained policy instantly collapsed the robot. I switched the actions to offsets from the standing pose, so "do nothing" means "hold the stance."

  • Train/eval mismatch. The checkpoint viewer was still using the old action mapping, so it showed a crouch the policy never learned. Aligning the viewer with the training environment fixed it. Simple bug, many hours.

Results

The final policy walks upright for the full 8-second evaluation window without falling, holding the pelvis near its 0.40 m target. Over 282 evaluation episodes it averaged about 66 cm/s. That speed is very high for a robot this size and beyond what the servos can deliver, suggesting the policy exploits idealized simulation physics, such as instant actuator response and perfect contact.

Sim-to-Real

Transferring the policy to hardware was the hardest part of the project.

  • Actuator mismatch: The simulated servos respond instantly and ideally. The real LX-16As have bus latency, limited speed, deadband, and backlash.

  • Unrealistic sim speed: The policy relies on dynamics the hardware can't reproduce.

  • Sensing gaps: The policy expects clean pelvis velocity and foot contact data. On the robot, velocity has to be estimated from the IMU, and the contact sensors are unreliable.

  • Interim solution: To get the hardware walking, I wrote a scripted, open-loop side-stepping gait that the robot executes on the bench.

Next Steps

  • Model servo lag, speed limits, and backlash in MuJoCo.

  • Add domain randomization for mass, friction, latency, and sensor noise.

  • Cap commanded speed to what the hardware can realistically achieve.

  • Resolve foot contact sensing, or retrain using estimated contact.

  • Export the retrained policy to ONNX and deploy it on the Pi.

  • Use the scripted gait as a baseline, or as a reference motion for residual RL.

  • Add IMU-based balance feedback, and use motion capture to analyze real walking dynamics.

Reflection

While the project is still a work in progress, it has been an invaluable hands-on experience combining mechanical design, embedded systems, simulation, and learning-based control. Just as valuable were the lessons about reward design and the sim-to-real gap, taking a characterful biped from concept to a working prototype with a trained walking brain in simulation.

Comments


bottom of page