Spec to silicon, hands off. A horse is strong, fast, and willing — and useless for heavy loads until you harness it. The harness is not a part of the horse and not a part of the cart: it is the coupling that turns raw strength into pulled weight.
$ npx -y skills add SensorsIot/Embedded-AI-Harness --agent claude-code
What's inside
Spec to silicon, hands off.
A horse is strong, fast, and willing — and useless for heavy loads until you harness it. The harness is not a part of the horse and not a part of the cart: it is the coupling that turns raw strength into pulled weight.
An AI is the same. It can write firmware all day — but it can't flash a board, can't see it boot, can't know whether its fix actually worked on real hardware. Unharnessed, it generates code and hopes. This repository is the harness: strap the AI in, and it pulls — writes the code, compiles it, flashes it onto a real ESP32, tests it against real WiFi, MQTT, BLE and RF, reads the failures, corrects itself, and goes again — until the tests run clean.
Today's AI coding is open-loop: prompt → code → hope. No feedback, so errors accumulate uncorrected — which is exactly why people don't trust AI-written firmware. AI Closed-Loop Programming (AICLP) closes the loop with reality:
FSD ────── the setpoint: what "done" means
│
▼
┌──── code → build → flash ────┐ forward path
│ ▼
│ real hardware
│ │
└── correct ◄── tests ◄────────┘ feedback path
the loop exits when the error signal is zero: tests green
Every embedded engineer knows this diagram — it's a control loop. The spec (FSD) is the setpoint, the firmware on the chip is the plant, the tests are the sensor, failing tests are the error signal, and the AI is the controller that corrects until the error reaches zero.
True TDD, enabled by AI. For twenty-five years, developers drove and tests advised — written after the code, skipped under deadline, tuned until they passed. Here the tests drive for the first time: derived from the spec, run on real silicon, and the only way the AI gets to stop.
| Phase | You do | You get | Gate |
|---|---|---|---|
0 · Definition — /define | Describe the product; answer an interview, one question at a time | An FSD where every requirement already says how it will be proven | Load defined |
1 · Harness — /harness | One command; answer the two questions only you can | The project strapped in: docs, test plan, firmware hooks, CI, runner | AI harnessed |
2 · Commissioning — /commission | Plug the board into a slot, wire any peers | Your board and peers proven working here — a failing test now means the code | DUT ready |
3 · Build — /build | Start sessions; approve the occasional spec question | Requirements turning green, one by one, on real hardware | Ready for shipment |
⚑ Shipment — git tag | Push the version tag — the one act that stays human | A release built in a pinned container and verified on the testbench: the journey runs once more on the exact bytes users download | Shipped |
Each phase ends at a gate, derived from project state and never declared — nobody types "phase complete". A gate is not a marker you pass: if anything it requires is unmet, the work loops back to the step that owns it and the whole check runs again. And after shipment the same journey repeats in miniature for every new feature: describe it in a sentence, the loop refuses to code anything no requirement covers, the spec absorbs the delta, and the phases collapse to minutes. No code without a clause is what keeps the spec true for the product's whole life.
/define (the
FSD: atomic, falsifiable requirements, each with its verification
contract), /harness (one-time setup), /commission and /build (the
loop's driver: test design, the plan, audit, what's next).| You need | For | Where |
|---|---|---|
| A Raspberry Pi testbench (Pi 3/4/5, or Zero 2 W + USB hub + Ethernet adapter) with an ESP32 board in a slot | The loop's hands and eyes — flash, reset, observe on real hardware | Build it: Quick Start below |
A GitHub account, git + gh authenticated | This is where the firmware is built — see below | github.com |
| Claude Code with this repo's skills | The AI that pulls; the skills are the method | npm i -g @anthropic-ai/claude-code, then copy .claude/skills/ from this repo into your project |
| Same LAN | Your dev machine or devcontainer must reach the bench | curl http://$BENCH:8080/api/devices answers |
Throughout this README, $BENCH is the bench's IP address — never a
.local name, which does not resolve from inside a container. Find it once:
BENCH=$(python3 .claude/skills/esp-idf-handling/discover-testbench.py | jq -r .ip)
Nothing else — and in particular no local ESP-IDF or PlatformIO installation. The forward path runs through GitHub Actions:
git push ─► CI builds in a pinned container ─► gh run download ─► flash to a slot ─► observe
The loop flashes the artefact CI produced, never a binary that happens to
be sitting in a local build/ directory. That is not a convenience, it is
what makes the evidence mean anything: the bytes on the chip are the bytes in
the run you can point at, built by a toolchain whose version is pinned in the
workflow rather than whatever is installed on somebody's laptop. The same
property is what lets a tag ship — the release-verify job flashes the
released artefact to the bench and reruns the journey, so a red journey is a
release that does not happen.
The toolchain skills (esp-idf-handling, esp-pio-handling) drive that
pipeline and know how to build locally too, which is useful when you are
iterating by hand. The loop does not depend on it.
The FSD, tests, firmware, CI and documentation are what the loop produces, not what you bring.
Working on an ESP32 normally means being physically attached to it — and an AI can't hold a USB cable. The testbench puts the boards on a Raspberry Pi and turns everything into HTTP:
LAN (192.168.0.x)
|
| eth0 (wired)
v
Raspberry Pi ---- wlan0 (WiFi test AP: 192.168.4.x)
testbench_7e71 hci0 (Bluetooth LE)
| UDP :5555 (log receiver)
| USB hub (internal on Pi 3/4/5, external on Zero)
|
+----+----+----+----+
| | | |
:4001 :4002 :4003 :4004 <- auto-assigned (4001 + slot index)
SLOT1 SLOT2 SLOT3 SLOT4 <- one per detected hub port
Plug in a board → it's ready. Auto-detected in seconds and mapped to a
fixed port by which USB connector it's in — same connector, same port,
always. That's slot-based identity: a slot is a physical hole in the
hub, so scripts and platformio.ini never go stale when boards swap or the
kernel renames /dev/ttyACM0.
Serial over the network at rfc2217://$BENCH:4001 — esptool,
PlatformIO, ESP-IDF and anything on pyserial speak it natively.
Flash three ways — over the network, locally on the Pi, or over the air.
Debugging out of the box — OpenOCD starts itself for USB-JTAG chips and
GDB connects on that slot's own port (3333 + slot index). Slots are
independent: two boards debug at once, each selected by its USB port path,
because every ESP32 with built-in JTAG enumerates as 303a:1001 and
VID:PID alone cannot say which board is which.
The Pi is the test equipment. Its WiFi becomes the access point your board joins, its Bluetooth scans and connects, optional SDR and Si5351 hardware receive and transmit on 433 MHz, and boards log to it over UDP when USB is busy.
It presses the buttons. GPIO wired to reset and boot forces download mode and rescues boot-looping boards with nobody in the room.
Claude drives all of it through 70 MCP tools or the bundled skills. The newest endpoints — serial write, bench reset, the slot access manager — are HTTP and skills only; the MCP surface has not caught up yet.
Every slot is recorded, always. A reader runs whether or not anyone is
watching, so a boot banner is in the buffer before you think to ask for it,
and tcp_port + 1000 is a read-only fan-out of the same bytes for as many
watchers as you like.
It has its own test partner. test-firmware/ is an ESP32 image the
bench builds, versions and flashes itself. It is not a device under test —
it is the counterpart the bench measures itself against: it joins the test
AP, hosts one of its own, answers HTTP, and replies on a serial console. The
bench has one radio and cannot be the access point its own station tests
join, so the partner is how half these requirements are provable at all —
and why the suite never depends on a project's firmware.
Honest limits: one writing serial client per board — RFC2217 gives one session, though any number can read the fan-out — the SDR is one dongle, one user, and the API has no authentication: keep the bench on a network you trust.
You need a Raspberry Pi with onboard WiFi and Bluetooth running Raspberry Pi OS Lite (64-bit). A Pi Zero 2 W also needs a USB hub and a USB Ethernet adapter, since wlan0 is reserved for testing; a Pi 3/4/5 has both built in. An RTL-SDR dongle, an Si5351 + PE4302, and jumper wires to the board's EN/BOOT pins are all optional.
git clone https://github.com/SensorsIot/Embedded-AI-Harness.git
cd Embedded-AI-Harness/pi
sudo bash install.sh
That installs every dependency (pyserial, hostapd, dnsmasq, bleak, esptool, OpenOCD, rtl-sdr/rtl_433, mosquitto), sets up the udev hotplug rules, and starts the portal as a systemd service. Plug in a board and check:
curl http://$BENCH:8080/api/devices | jq
Slots are auto-detected — no config file needed. Create
/etc/rfc2217/testbench.json only to rename slots, pin ports, declare GPIO
pins, or register an ESP-Prog probe; sudo rfc2217-learn-slots prints one
for you.
On a Pi Zero 2 W, do the memory hardening first. With 512 MB the board OOM-crashes under load, and hard crashes corrupt the SD card. See User Manual §2.2.
Watch a board boot — no client library, just HTTP:
curl -X POST http://$BENCH:8080/api/serial/reset \
-H 'Content-Type: application/json' -d '{"slot":"SLOT1"}'
Point your existing tools at it. PlatformIO needs one line
(upload_port = rfc2217://$BENCH:4001); esptool takes the same URL,
and the binaries stay on your machine:
esptool --port rfc2217://$BENCH:4001 --chip esp32c3 \
write-flash 0x10000 firmware.bin
Write a test that uses the whole bench — reset the board, give it a network to join, wait for it to appear, then talk to it:
FAQ
embedded-ai-harness is a Claude Code plugin with 18 hand-picked skills for development work, indexed on Flowy. Install it with the command on its page. It includes build, commission, define. Its skills do not fire on their own yet. Request auto-invocation to have Flowy route them as you prompt. Free and open source.
Is this plugin yours?
Claim it with GitHubSubmit a pluginPromote it