Security model¶
Automation lets an AI agent run experiments on real instruments. These rules keep a person in charge of the hardware:
- The agent can only propose a protocol: which devices, which methods, the bounds on every number, call and rate caps, a duration and a safe state. A person approves the exact protocol (by its sha256) or rejects it.
- Only an approved protocol yields a grant. During it, every device call goes through the Control Center's gateway, which refuses anything outside the protocol. The agent's scripts never hold a device or a key.
- Ending a grant, for any reason, applies the safe state and releases the devices. If the Control Center dies mid-grant, the safe-state watchdog does it.
- Independently of all this, each Lab PC enforces its drivers' published limits and its own site limits on every remote call, from any client.
- The runner runs the agent in locked-down containers that reach only the Control Center and the model API; see The Automation runner.
Protocols and approval¶
Before an agent may touch hardware, it writes a protocol: which devices,
which methods, the bounds of every argument, call and rate caps, how long,
and the safe state. A person reviews it and approves or rejects it
(automation_protocol.py). While a run is active, its phase says where it
is:
planning ──(submits a protocol)──► awaiting_approval ──approve──► executing ──(grant ends)──► analyzing
▲ │
└──────────reject + note─────────────┘
A run that never needs devices goes from planning straight to its finish.
The files. For Tasks/Task1.md the agent writes
Protocols/Task1.protocol.json and, in its own words,
Protocols/Task1.protocol.md. The runner uploads both, then calls
POST /api/automation/runner/runs/<run>/protocol {}. The Control Center
reads the JSON from the project itself:
{
"schema_version": 1,
"task": "Tasks/Task1.md",
"title": "AOM efficiency vs RF frequency",
"duration_minutes": 20,
"devices": [
{
"device_id": "dev01-aom",
"device_type": "dummy_aom",
"methods": {
"set_frequency": {"args": {"value": {"min": 70, "max": 90, "max_step": 2}}, "max_calls": 200, "max_rate_hz": 5},
"set_rf_enabled": {"args": {"enabled": {"choices": [true, false]}}, "max_calls": 10, "max_rate_hz": 1},
"read_diffraction_efficiency": {"max_calls": 500, "max_rate_hz": 10},
"stream_diffraction_efficiency": {
"args": {"channel": {"choices": ["CH1"]}},
"stream": {"max_sample_rate_hz": 20, "max_samples": 1000},
"max_calls": 50, "max_rate_hz": 1
}
}
}
],
"safe_state": [{"device_id": "dev01-aom", "method": "set_rf_enabled", "args": {"enabled": false}}]
}
The checks. Everything in the file is untrusted, so the checks are strict.
Each refusal names the place (devices[0].methods.set_frequency.args.value.max)
and says what to change:
- The file: UTF-8 JSON, at most 64 KiB. Unknown keys are refused at
every level, and so are repeated keys, NaN and Infinity.
- Top level: schema_version 1. task is the run's own task. title
has at most 120 characters (text never has control or invisible characters,
nor more than 2 combining marks on one character, which could draw over the
review). duration_minutes is from 1 to
LABNET_AUTOMATION_GRANT_MAX_MINUTES (120). There are 1 to 16 devices,
each listed once.
- Devices: each one must be on a canvas linked to the project, allowed
for automation in the project's settings, and on the network now.
device_type must be its live type, and the linked experiment must agree.
- Methods: each one must be in that type's contract, with max_calls
(a whole number, 1 … 100 000) and max_rate_hz (above 0, at most 100).
A stream also needs stream: {max_sample_rate_hz (above 0 … 200),
max_samples (1 … 100 000)}; nothing else may have one.
- Arguments: args bounds every argument the method sends. An optional
one may be left out, and then it may never be sent; one that is bounded is
sent in every call (otherwise leaving out a channel limited to CH2
would drive the driver's default, CH1). A write's unit is never allowed.
| Argument type | Bounds |
|---|---|
number (float, int) |
{"min", "max", "max_step"?}: finite, min ≤ max, max_step > 0; whole numbers for an int |
bool |
{"choices": [true, false]} (one or both) |
str |
{"choices": [...]}: up to 32, each up to 64 characters |
true is never a number, and 1 is never true.
- max_step needs a reading. Steps are measured from where the device
is, so max_step is allowed only on a set_<x> that has a scalar
read_<x> taking the same choice arguments (set_amplitude(value,
channel) → read_amplitude(channel); automation_protocol.step_reader).
Today that is every setter of dummy_aom, dummy_awg, dummy_eom and
dummy_laser; not dummy_device.write_value, the simulated_tunable_laser methods
or simulated_wavemeter.set_exposure, where the check says to drop
max_step (any value in range is then allowed) or narrow the range.
- Published limits are hard limits. A device publishes the driver's
hints narrowed by its Lab PC's site limits. [min, max]
must lie within its limits["<method>.<argument>"]. Choices (text or
true/false) must be among its choices, and a channel among its
channels, whenever it publishes them. Each refusal says whether the limit
is the site's or the driver's.
- No limit, no automation. Every number a protocol bounds needs a
published limit. Without one the protocol is refused (a problem, not a
warning): "dev05-laser publishes no limit for
go_to_wavelength_nm.wavelength_nm; set one in the Lab PC inventory with
site_limits(...) before automating it." A person sets that limit in the
inventory and restarts the Lab PC. Text arguments without published choices
still need explicit choices in the protocol.
- The safe state: up to 16 calls, each one the protocol itself allows:
a listed method (not a stream), every argument it sends and every argument
the protocol bounds (e.g. its channel), and values within the bounds.
max_step, call counts and rates don't apply to it. A whole-number
argument stays an integer (0, not 0.0), so the gateway accepts it.
A protocol that passes is normalized: defaults filled in, bounds and
rates as floats, counts as integers. It is serialized canonically
(json.dumps(sort_keys=True, separators=(",", ":"), ensure_ascii=False)).
That text is stored as state/<project>/protocols/<run>-<n>.json, with the
prose beside it as <run>-<n>.md. Its sha256 is what a person approves.
Submitting.
- 200 {"phase": "awaiting_approval", "sha256", "submission"}: the run
now waits for a person.
- 422 {"detail", "problems": [{"path", "message"}]}: the run stays
planning. A refused file doesn't count as a submission; the agent fixes it
and submits again.
- 409 when the run isn't planning, or has already sent 3 protocols for
approval.
- 503 while Central can't be reached: a protocol is never checked against
a stale picture of the network.
Reviewing. Clicking a run that waits for approval (Review in the runs
list) shows:
- The approval card: the protocol's sha256 prefix and the requested
duration. The duration box can shorten the time but never lengthen it.
Approve asks to confirm. Reject needs a note for the agent.
- The review: built by the server from the normalized copy and the
device contracts, never from the agent's words. Each device shows its
type and live state. Each method shows every argument's range or choices,
its largest step (the first measured from the device's own reading) and
the device's published limit with its source (limit, limit_source:
site or driver; choices as offered), that each bounded argument is always sent, which
optional arguments may never be sent, its call and rate caps, and its
stream caps. Agent-written text (the title, choices) is clipped to its own
line, so it can't draw over the rest. Then come the
safe state and warnings: no safe state, no max_step, or a limit no
longer published (the device went offline; approving checks again).
- The agent's explanation: the .md, in a box labelled as the agent's,
shown as plain text only.
Once approved, the page shows the grant with a countdown and Stop now.
A run whose safe state failed ends ended_unsafe, with a red banner.
Approving sends the sha256 the page showed. The approver is the
signed-in person (with logins) or the name typed, and is
required. The Control Center then:
1. refuses a different sha256 (409), a run not waiting for approval (409),
or a longer or unusable duration (422, also for an integer too large for a
float);
2. checks the stored protocol again against the project and the live
network. A device no longer allowed, gone or of another type answers 409
with the problems, and the run keeps waiting: reject it so the agent can
write a new one;
3. creates the grant, taking every lease.
A busy device answers 409, and nothing changes;
4. records the approval (reviewed_by, approved_minutes). The run is then
executing.
One approval runs at a time. If the run ends while its grant is being created, the grant is ended again.
Editing before approving (Edit… on the approval card). The approver
may change the protocol and approve their own version instead of rejecting
it:
- What can be edited, with plain fields: each argument's min, max
and max_step (empty: any step) or its choices; max_calls,
max_rate_hz and the stream caps; removing a method or a whole device; the
duration it asks for; and the safe state (change a value, remove a call,
move it earlier, or add a call to a method the protocol lists).
Advanced: edit as JSON shows the whole edited protocol for anything else.
- The same checks as the agent's. The approver is trusted to widen or
narrow anything, but only within the device's and the lab's limits:
validate_protocol runs on the edit against the project and the live
network, and a removed method still used by the safe state is refused (the
page drops those calls itself and lists them). The duration may be raised
up to LABNET_AUTOMATION_GRANT_MAX_MINUTES; the grant may then be as long
as the edit asks.
- Review changes… sends the edit to POST …/protocol/preview
{"sha256", "protocol"}, which checks it without approving and answers
{"normalized", "sha256", "problems", "changes", "review"}. Problems are
shown beside the fields. Otherwise a confirmation lists the changes the
server found, one line each (dev01-aom.set_frequency.value.max: 90 → 85,
dev01-aom.read_diffraction_efficiency: removed (may not be called),
safe_state: added dev01-aom.set_frequency(value=75)).
- Approving sends {"sha256", "duration_minutes", "protocol",
"edited_sha256"}: sha256 is the submission the edit started from (409 if
another one is waiting now), protocol the previewed version and
edited_sha256 its sha256 (409 if the server's normalized copy differs).
The server checks the edit again: one that no longer passes answers 422
with problems, and nothing changes. An edit that changes nothing is a
plain approval.
- What is kept. The edit is normalized and stored as
state/<project>/protocols/<run>-<n>-edited.json (its canonical text, so
the file's sha256 is its own); the agent's <run>-<n>.json stays as it
was. The grant is created from the edit, under the edit's sha256, so
grant/token hands the runner the edited protocol. The run's protocol
record gains approved_sha256 (the edit's), edited_by and changes (at
most 60 lines, each clipped to 240 characters); the log says approved …
EDITED with the changes.
- The agent is told. On resume the runner's prompt says the protocol was
approved WITH EDITS, quotes the changes from the run record, and gives the
edited bounds (the runner checks the protocol it got against
approved_sha256). grant_status shows them too.
- Afterwards the run's page shows Edited by … before approval with the
changes, the approved version's review, and the agent's submission under
As the agent submitted it.
Rejecting ({"note"}, up to 2000 characters) returns the run to
planning. The note is in the run record (protocol.note), where the
runner reads it, and the agent may submit again, up to 3 protocols per run.
Waiting. A protocol not approved or rejected within
LABNET_AUTOMATION_APPROVAL_HOURS (24) fails the run. Meanwhile, the
runner keeps heartbeating; the 10-minute rule still applies.
The run record gains phase and protocol: {submissions, status
(submitted | approved | rejected), sha256, title, duration_minutes,
submitted_at, reviewed_by, reviewed_at, approved_minutes, approved_sha256,
edited_by, changes, note} for the latest submission. approved_sha256 is
the version the grant enforces: sha256 unless the approver edited it. It also gains grant, written by the GrantManager: {id,
status (active | ending | ended | ended_unsafe | interrupted), approved_by,
starts_at, ends_at, devices, calls, refused_calls, ended_reason, ended_at}.
- A run whose grant ended ended_unsafe ends ended_unsafe however it
finishes (done, failed, stopped or cancelled), and keeps its
summary.
- Run records from before Phase 3 load with no phase.
| Request | Does |
|---|---|
POST /api/automation/runner/runs/<run>/protocol {} |
Checks Protocols/<task>.protocol.json and sends it for approval (runner token) |
GET /api/automation/projects/<id>/runs/<run>/protocol |
{"protocol" (normalized), "sha256", "review", "prose" (the agent's, untrusted), "status", "submission", "phase", "run_status", "max_minutes", "approved_sha256", "edited_by", "changes"}; with an edit also "approved_protocol", "approved_review" |
POST …/runs/<run>/protocol/preview {"sha256", "protocol"} |
Checks an approver's edit without approving: {"normalized", "sha256", "problems", "changes", "review"} |
POST …/runs/<run>/protocol/approve {"sha256", "duration_minutes", "approved_by"?, "protocol"?, "edited_sha256"?} |
Approves (the edit, if sent); starts the grant; {"run", "grant"} |
POST …/runs/<run>/protocol/reject {"note", "sha256"?, "rejected_by"?} |
Rejects; the run is planning again |
Device grants and the gateway¶
When a person approves a run's protocol, the Control Center creates a
grant (automation_grants.py). A grant holds the protocol's bounds, a
time limit (at most what the protocol asked for), and the leases on every
device it names. It is the only way an agent's script reaches hardware, and
every call is checked against it.
Creating a grant. Every lease is taken up front, and each device's live
type must match the protocol's device_type. If one device is busy or
mismatched, the leases already taken are released and approval answers 409.
If the Control Center starts shutting down while the leases are taken, they
are released and no grant is made. The grant's token (lng_…) exists only
in memory until the runner takes it once (POST runs/<run>/grant/token),
together with the normalized protocol the person approved, from which the
runner tells the agent its bounds. state/<project>/grants/<id>.json keeps
only its sha256.
Every call goes to POST /api/automation/runner/gateway/calls with
Authorization: Bearer lng_…:
- The body is at most 16 KiB of application/json:
{"device_id", "method", "args": {…}, "sample_rate_hz"?, "sample_count"?}.
- A browser's request (Sec-Fetch-Site) is refused. The path is under the
runner prefix, so it needs no Caddy login, and the grant token is its only
credential.
- An allowed call answers {"result", "calls_left", "seconds_left"}. A
stream's result is its list of values.
- Anything else answers 4xx {"detail", "code"}:
| Code | HTTP | Meaning |
|---|---|---|
refused |
422 (413/415 for the body, 403 for a browser) | Outside the approved protocol. Final: fix the script, don't retry. |
busy |
409 | A call to that device is still running (one at a time per device). |
expired |
410 | The grant's time is up. |
ended |
410 (401 for an unknown token) | Stopped, ended by the agent, too many violations, or the run ended. |
device_error |
424 | The device failed. The message is redacted: no lease ids, keys or tokens. |
What the gateway checks. Anything it does not expect is refused, and a
malformed protocol refuses rather than allows.
- The grant is active and has time left.
- The device is in the grant, and the method is listed for it.
- args is a JSON object of exactly the method's keyword arguments:
- every required one is there;
- every one the protocol bounds is there (a bounded channel left out
would drive the driver's default channel);
- an optional one only if the protocol bounds it;
- never unit, and nothing else.
- Numbers:
- they are JSON numbers (true is not a number), finite, and within
[min, max];
- a whole-number argument takes only integers;
- with max_step, a number is at most that far from where that setting
may be now. "That setting" is the device, method, argument and the
call's choice arguments (dev02-awg.set_amplitude.value|channel=CH2),
so alternating channels can't double a step. Before the first change the
gateway reads it with the matching reader and the same choices
(read_amplitude(channel="CH2")): with the grant's lease, not counted
against the protocol's calls, audited as a reading line. A reading that
fails (or isn't a number) refuses the change (device_error, the change
not sent). A call that succeeded sets the position to its value; one
whose outcome is uncertain (a device error after sending, a timeout, the
runner hanging up) widens it to [min(old, sent), max(old, sent)], and
the next value must be within max_step of both ends.
- Booleans and text must be one of the approved choices. A boolean must be
a real JSON true/false.
- A stream sends sample_rate_hz (above 0, at most max_sample_rate_hz)
and an integer sample_count (1 … max_samples), and
sample_count / sample_rate_hz must fit in the time left. Other methods
send neither.
- Each method keeps to its max_calls and to max_rate_hz: at least
1/max_rate_hz seconds since that method's previous call.
- Each call is cut off when the grant's time runs out.
Audit and violations. Every call is one line in
state/<project>/audit/<run>/grant-<id>.jsonl, whether it was allowed or
not, including a body refused before it could be read (not JSON, too large,
the wrong media type, a number too long), which also counts as a refused
call once the token is known. A line has the arguments, the outcome and the
duration (and the result, or the number of samples). Ten refused calls in a
row end the grant. A busy reply or a device error during a call doesn't
count toward that; a change refused because its reading failed does.
Lapsed leases. The leases renew themselves while the grant holds them. Central's idle timeout (15 minutes without a command) can still end one. Then the gateway takes the lease again once. It retries the call only when the call never left the Control Center, or when the call only reads. A write or command the Lab PC refused is not sent twice. If the lease can't be taken again, the grant ends.
Ending. A grant ends when:
- its time runs out;
- someone presses Stop;
- ten calls in a row are refused;
- the run is cancelled, fails, goes quiet or finishes;
- the agent ends it (POST runs/<run>/grant/end {"reason"});
- the Control Center shuts down.
It always ends the same way:
1. New calls are refused at once (ended/expired).
2. Calls in flight are cancelled.
3. The protocol's safe_state calls are made. Each gets up to three tries,
and a lost lease is taken again once.
4. The leases are released, and the end is written to the audit file.
5. The grant becomes ended. If any safe-state call failed, it becomes
ended_unsafe, and the run is marked ended_unsafe with a red banner:
check those devices by hand.
6. The run moves on to analyzing, unless it has already ended.
If the Control Center dies or hangs mid-grant, it can't do any of this
itself: its leases lapse about a minute later. The
safe-state watchdog, a separate process, then
applies the safe state. At the Control Center's next start, grants left
active are marked interrupted, unless the watchdog has already finished
with them; then its ended/ended_unsafe stands.
While a grant holds a device:
- The Network and Devices pages show Automation:
| Request | Does |
|---|---|
POST /api/automation/runner/runs/<run>/grant/token {} |
The grant's token, grant and the approved protocol, once (runner token; 409 after that) |
POST /api/automation/runner/runs/<run>/grant/end {"reason"} |
Ends the grant now (safe state, release); {"grant"} |
POST /api/automation/runner/gateway/calls |
One device call, with the grant token (above) |
POST /api/automation/projects/<id>/runs/<run>/grant/stop {} |
Stop from the project page (anyone; same-origin JSON) |
POST /api/automation/grants/<grant_id>/stop {} |
Stop from the Network/Devices page (anyone; same-origin JSON) |
The safe-state watchdog¶
labnet-watchdog is a small separate process on
central_server_pc (labnet_control_center/watchdog/). If the Control Center
dies, or its event loop hangs, while a grant holds devices, the watchdog
applies that grant's safe state. It never fights a live Control Center:
Central lets only one of them hold a device's lease at a time.
What the Control Center writes for it. While a grant holds its leases
(active or ending), a task on the Control Center's event loop rewrites
state/<project>/grants/<id>.json every 5 s with a fresh alive_at, so a
hung loop stops the heartbeat. The record holds everything the watchdog
needs: the approved normalized protocol (devices, types, safe-state calls),
its sha256, ends_at and the status. If the heartbeat can't be written
three times in a row, the Control Center ends the grant itself, with its safe
state.
When the watchdog acts. It reads every grant record every 2 s:
- A grant is taken over when its heartbeat hasn't changed for 20 s (timed on
the watchdog's own clock) and either:
- the Control Center's /health/live hasn't answered for 20 s either (it
died); or
- it answers but the heartbeat stays stale for 40 s (its loop hung).
- A grant the Control Center marked interrupted at a restart is taken over
at once: the process that held it is gone.
- Never:
- an ended or ended_unsafe grant;
- one it has already finished;
- one with no heartbeat for more than an hour, which is only logged
(LABNET_WATCHDOG_MAX_AGE_MINUTES; check those devices by hand).
- Expired grants (time ran out while the Control Center was dead) are safed
like the others.
Taking over. The watchdog writes only its own claim file,
<id>.watchdog.json next to the record, and lines of "kind": "watchdog"
in the grant's audit file. Two processes never write the same file.
1. The claim becomes claimed.
2. It asks Central for each device the safe state needs, every 3 s for up to
120 s. Central refuses while the Control Center's lease lives.
3. If the Control Center's heartbeat comes back before a lease is taken, or
it ends the grant itself, the watchdog stands down (stood_down).
The same happens when nothing was leased in 120 s and the Control Center
still answers.
4. With the first lease, the claim becomes safing. From then on the grant
is the watchdog's, and a Control Center that comes back hands it over: it
refuses calls, cancels calls in flight and releases its leases without a
safe state of its own.
5. The protocol's sha256 is checked against the approved one, and each
safe-state call is checked against its bounds as the gateway checks it.
Each call gets up to three tries, re-taking a lost lease once. Then every
lease is released.
6. The claim ends ended (applied by the watchdog: the Control Center stopped
answering) or ended_unsafe (a call failed, or a device couldn't be
leased in time: check it by hand).
7. The Control Center copies the outcome into the grant's record
(safe_state_by: "watchdog") and into the run, whose red banner shows an
ended_unsafe. If it is down, its next start does this.
A watchdog stopped mid-takeover releases what it holds. If its claim was
safing, the next watchdog finishes it.
Set it up on central_server_pc:
1. Its own Central user (role user, never the Control Center's key):
$admin = "C:\ProgramData\LabNet\central-venv\Scripts\labnet-central-admin.exe"
& $admin --env-file C:\ProgramData\LabNet\central.env setup `
--user automation-watchdog --write-env C:\LabNet-Keys --server-url http://127.0.0.1:8008
- Move
C:\LabNet-Keys\automation-watchdog.envtoC:\ProgramData\LabNet\watchdog.env, then deleteC:\LabNet-Keys. Add the lines fromsites/pqt/watchdog.env.example: at leastLABNET_AUTOMATION_DIR, the Control Center's own (defaultC:\ProgramData\LabNet\automation), andLABNET_CA_FILEifcontrol-center.envsets one. - Run it as the same Windows account as the Control Center, which owns that folder. The command is installed with the Control Center:
C:\ProgramData\LabNet\control-center-venv\Scripts\labnet-watchdog.exe --env-file C:\ProgramData\LabNet\watchdog.env
Invoke-RestMethod http://127.0.0.1:8081/health/ready
Start it by hand in its own window, after the Control Center (see Starting the lab by hand). To start it automatically at logon instead, register a Task Scheduler task. It has no window, so it writes a log file:
New-Item -ItemType Directory -Force C:\ProgramData\LabNet\logs | Out-Null
$exe = "C:\ProgramData\LabNet\control-center-venv\Scripts\labnet-watchdog.exe"
$action = New-ScheduledTaskAction -Execute "cmd.exe" `
-Argument "/c $exe --env-file C:\ProgramData\LabNet\watchdog.env >> C:\ProgramData\LabNet\logs\watchdog.log 2>&1"
$trigger = New-ScheduledTaskTrigger -AtLogOn -User $env:USERNAME
$settings = New-ScheduledTaskSettingsSet -ExecutionTimeLimit ([TimeSpan]::Zero) `
-RestartCount 3 -RestartInterval (New-TimeSpan -Minutes 1)
Register-ScheduledTask -TaskName "LabNet safe-state watchdog" -Action $action -Trigger $trigger -Settings $settings
Start-ScheduledTask "LabNet safe-state watchdog" starts it now, and
Stop-ScheduledTask stops it.
- Only one watchdog runs per computer: a second one can't take port 8081
and exits.
- It doesn't matter whether it starts before or after the Control Center.
Stop it before an update, like the Control Center.
In development, labnet dev seed creates the user and
.devnet/automation-watchdog.env. labnet dev up starts the watchdog
after the Control Center (log: .devnet/logs/watchdog.log,
health: http://127.0.0.1:8081/health/ready). labnet dev down stops it first,
labnet dev status shows it, and labnet dev restart watchdog restarts it. A
network seeded before the watchdog existed gets its user with labnet dev seed
--watchdog, which deletes nothing.