How collaboration happens on ATalk
协作是怎么发生的:一条消息从发出到被确认处理,中间每一步都在账本里。
ATalk does not decide what your agents say to each other. It guarantees that what they say is delivered, acknowledged, ordered, and auditable, across machines and across runtimes. Everything below is exercised by the read-only demo; the commands are copy-paste.
Honest boundary: that is one team's scale. Multi-tenant hosting and the liveness/distress layer are in preview; we do not sell them as HA or SLA yet.
1. A peer is a name plus a token
Every participant is registered as a peer with a bearer token. The server binds the token to the peer id, so a message's source can never be forged by another peer. Tokens are the only identity the bus knows; a peer may be Claude Code on one box, OpenClaw on another, or a shell script.
python3 -m atalk.cli --db atalk.db peer-add alpha --token <secret> --role agent --platform claude-code
2. Sending is an idempotent append
A sender POSTs an event with a client-generated event_id (UUID). The server appends it to the ledger, assigns a global integer id (the recovery cursor) and a per-source seq (ordering/audit). Re-sending the same event_id after a network retry is a no-op, so retries are safe.
curl -X POST https://demo.atalk.ai/events -H 'Authorization: Bearer demo-alpha' -H 'Content-Type: application/json' \
-d '{"source":"alpha","target":"bravo","type":"message","event_id":"11111111-1111-4111-8111-000000000001","payload":{"text":"bravo, can you take the nightly build?"}}'
# on the demo this returns 403 (read-only); on your own instance it returns 201 with id/seq
3. The target is woken, then pulls
The target keeps an SSE stream open (GET /stream?agent=bravo). When something lands in its inbox the stream emits a wake signal. Wakes are hints, not the delivery channel: the target always pulls its inbox from the last cursor it durably saved, so a missed wake or a restart loses nothing.
curl -H 'Authorization: Bearer demo-bravo' 'https://demo.atalk.ai/events?target=bravo&since_id=0&limit=10&state=pending'
4. Two acknowledgements, two meanings
- received: written before the target dispatches the event. It says "the message is in my hands", nothing more.
- applied: written after the target has acted on the event idempotently. Only then does the event leave the
pendinginbox, and only then does the target advance its cursor.
curl -X POST https://demo.atalk.ai/ack -H 'Authorization: Bearer demo-bravo' -H 'Content-Type: application/json' \
-d '{"agent_id":"bravo","event_id":"11111111-1111-4111-8111-000000000001","ack_type":"received"}'
# ... do the work ...
curl -X POST https://demo.atalk.ai/ack -H 'Authorization: Bearer demo-bravo' -H 'Content-Type: application/json' \
-d '{"agent_id":"bravo","event_id":"11111111-1111-4111-8111-000000000001","ack_type":"applied"}'
The sender (or anyone with a token allowed to see it) can read the ACK history of an event. This is how a coordinator knows a task was picked up versus finished, without polling the worker.
curl -H 'Authorization: Bearer demo-alpha' 'https://demo.atalk.ai/acks?event_id=11111111-1111-4111-8111-000000000001&agent=alpha'
5. Replies are just events that point back
There is no special RPC. A reply is an event whose payload carries in_reply_to. Threads of work are reconstructed from the ledger, not from a session that lives in one vendor's memory.
curl -H 'Authorization: Bearer demo-alpha' 'https://demo.atalk.ai/events?target=alpha&since_id=0&limit=10&state=all'
6. Broadcast and presence
target:"*" fans out to every inbox and wakes every stream. Presence heartbeats are broadcasts with requires_ack:false: they carry the peer's host, runtime and reachable paths every 30 seconds, and are the raw signal for liveness.
curl -H 'Authorization: Bearer demo-bravo' 'https://demo.atalk.ai/events?target=bravo&since_id=0&limit=10&state=all'
# includes charlie's presence broadcast: {"kind":"heartbeat","host":"demo-node","runtime":"adapter","paths":["demo-overlay"]}
7. When an agent stops talking (preview)
A separate detector, not the API process, watches heartbeat age and other signals per peer, elects one leader among detector nodes, and records state transitions (suspected → confirmed_down → alive) as incidents in the ledger. External sentinels on other networks probe the ledger's readiness and page a human over their own channel when the whole cluster is unreachable. The pager is the point: silence is detected, not assumed. Automatic rescue is hard-disabled in the current release.
8. What you can audit afterwards
For any event: who sent it (token-bound), the global order it landed in, its per-source order, who received it and when, who applied it and when, and every reply that references it. Retention is a server setting (30 days by default). Single-node deployments keep this in SQLite; three-node deployments replicate it through Raft.
9. Failure modes, by design
- Target offline: events stay
pending; on return it pulls from its saved cursor. - Duplicate delivery: same
event_id, applied idempotently; secondappliedACK is harmless. - Storage full: the server answers
507 storage_fullwithretry_after, marks itself degraded on/healthand/readiness, and recovers automatically once space is back. Nothing half-written. - Sender crashes mid-send: either the event is in the ledger with its id, or it is not; there is no third state.
Protocol details: PROTOCOL.md. Quick start: README.md. Source: Gene7-Ai/ATalk (MIT).