- Robot: 5.5 kg / 50 cm, 6 DOF (hip roll, hip pitch and knee per leg); Robstride RS01 ×4 plus EL05 ×2. All six motors sit on one plane at the top to cut leg inertia; the knee motors are mounted on the hip chassis and drive the knee remotely down the thigh through a PU steel-core timing belt. Raspberry Pi 5, 100 Hz control loop, Yesense IMU. PETG chassis, PLA legs, TPU point feet, all 3D printed.
- Decisions, mine: no mechanical or robotics background, so I researched with AI and then made and owned every technical call: Isaac Lab + rsl_rl PPO; the 6-DOF point-foot configuration; a PD position-target action space; a Python 100 Hz control loop on the Raspberry Pi rather than a migration to C++ or ROS; a DC motor model with 500 Hz physics; a stand-up curriculum; the budget and scale; and the selection and purchase of every motor, IMU, bearing, belt and shaft. The stack held all the way to standing; nothing was thrown out.
- Mechanical and hardware, mine: designed in one pass, and the first architecture is the one that stood. Battery tray fixed to the back of the hip roll motors, V-groove bearings taking the radial load off the flange; hip pitch and knee motors stacked back to back; a ground shaft and bearings as the knee axle, with its axial travel doubling as belt tensioning. All assembly, wiring, motor and CAN bring-up, bench testing and hardware testing. Every key hypothesis came from watching the machine: the IMU bias hypothesis; locating the “knee motor at the hip on hardware, at the knee in simulation” coupling; spotting belt tooth-skip by eye where the encoder could not see it; rejecting a 740-line bench plan in favour of a one-day trim sweep, which led directly to the first stand; and the training direction of “stack every harmless robustness measure.”
- Training, my judgement, agent execution: stepping and gait reward design. The dual-4090 workstation meant every run was an A/B pair in parallel; I chose what to compare and pushed for the bolder variant (the agent leaned toward small conservative ablations; P13, “stack every harmless robustness measure,” came out of that). The agent ran training and tracked TensorBoard; I read the curves throughout, judged stalls and collapses, and decided when to change a parameter, continue or kill a run. When P9 stalled, I found the swing-zero criterion across runs and reproduced it three times, which broke the plateau.
- Execution, initiated by me, implemented by the agent, reviewed by me: the control stack and SocketCAN driver; the CAN link migration (round trip 2.99 → 0.57 ms); derivation and implementation of the actuator and transmission model (rotor inertia decomposition, the off-diagonal hip-knee inertia coupling); the training-to-deployment consistency audit, which located four defects visible only in the code path; and data analysis. I overturned the agent's conclusions with physical observations three times. After calibration the original policy fell 512 out of 512 trials; the retrained policy survived 99%.
- Result: first autonomous balancing and standing on the real robot on 2026-09-03, and it stayed up through a belt tooth-skip. All compute from the self-built dual-4090 workstation, plus 3D printing and about $800 in agent tokens. The retrospective gives the item-by-item attribution and what working with an AI agent on a physical robot taught me.