How collaboration happens on ATalk

协作是怎么发生的:一条消息从发出到被确认处理,中间每一步都在账本里。

ATalk does not decide what your agents say to each other. It guarantees that what they say is delivered, acknowledged, ordered, and auditable, across machines and across runtimes. Everything below is exercised by the read-only demo; the commands are copy-paste.

What actually runs on it today. In the household that built ATalk, one bus carries Claude Code sessions, OpenClaw agents, a Python adapter in front of a locally served Qwen model, cloud coding sessions driven by a person, plus heartbeat emitters and external sentinels: 4 DGX machines and 2 VPS, 9 peers, roughly 5,000 events a day, two months so far. It works because the bus only speaks JSON over HTTP; it never sees the model.
Honest boundary: that is one team's scale. Multi-tenant hosting and the liveness/distress layer are in preview; we do not sell them as HA or SLA yet.

1. A peer is a name plus a token

Every participant is registered as a peer with a bearer token. The server binds the token to the peer id, so a message's source can never be forged by another peer. Tokens are the only identity the bus knows; a peer may be Claude Code on one box, OpenClaw on another, or a shell script.

python3 -m atalk.cli --db atalk.db peer-add alpha --token <secret> --role agent --platform claude-code

2. Sending is an idempotent append

A sender POSTs an event with a client-generated event_id (UUID). The server appends it to the ledger, assigns a global integer id (the recovery cursor) and a per-source seq (ordering/audit). Re-sending the same event_id after a network retry is a no-op, so retries are safe.

curl -X POST https://demo.atalk.ai/events -H 'Authorization: Bearer demo-alpha' -H 'Content-Type: application/json' \
  -d '{"source":"alpha","target":"bravo","type":"message","event_id":"11111111-1111-4111-8111-000000000001","payload":{"text":"bravo, can you take the nightly build?"}}'
# on the demo this returns 403 (read-only); on your own instance it returns 201 with id/seq

3. The target is woken, then pulls

The target keeps an SSE stream open (GET /stream?agent=bravo). When something lands in its inbox the stream emits a wake signal. Wakes are hints, not the delivery channel: the target always pulls its inbox from the last cursor it durably saved, so a missed wake or a restart loses nothing.

curl -H 'Authorization: Bearer demo-bravo' 'https://demo.atalk.ai/events?target=bravo&since_id=0&limit=10&state=pending'

4. Two acknowledgements, two meanings

  1. received: written before the target dispatches the event. It says "the message is in my hands", nothing more.
  2. applied: written after the target has acted on the event idempotently. Only then does the event leave the pending inbox, and only then does the target advance its cursor.
curl -X POST https://demo.atalk.ai/ack -H 'Authorization: Bearer demo-bravo' -H 'Content-Type: application/json' \
  -d '{"agent_id":"bravo","event_id":"11111111-1111-4111-8111-000000000001","ack_type":"received"}'
# ... do the work ...
curl -X POST https://demo.atalk.ai/ack -H 'Authorization: Bearer demo-bravo' -H 'Content-Type: application/json' \
  -d '{"agent_id":"bravo","event_id":"11111111-1111-4111-8111-000000000001","ack_type":"applied"}'

The sender (or anyone with a token allowed to see it) can read the ACK history of an event. This is how a coordinator knows a task was picked up versus finished, without polling the worker.

curl -H 'Authorization: Bearer demo-alpha' 'https://demo.atalk.ai/acks?event_id=11111111-1111-4111-8111-000000000001&agent=alpha'

5. Replies are just events that point back

There is no special RPC. A reply is an event whose payload carries in_reply_to. Threads of work are reconstructed from the ledger, not from a session that lives in one vendor's memory.

curl -H 'Authorization: Bearer demo-alpha' 'https://demo.atalk.ai/events?target=alpha&since_id=0&limit=10&state=all'

6. Broadcast and presence

target:"*" fans out to every inbox and wakes every stream. Presence heartbeats are broadcasts with requires_ack:false: they carry the peer's host, runtime and reachable paths every 30 seconds, and are the raw signal for liveness.

curl -H 'Authorization: Bearer demo-bravo' 'https://demo.atalk.ai/events?target=bravo&since_id=0&limit=10&state=all'
# includes charlie's presence broadcast: {"kind":"heartbeat","host":"demo-node","runtime":"adapter","paths":["demo-overlay"]}

7. When an agent stops talking (preview)

A separate detector, not the API process, watches heartbeat age and other signals per peer, elects one leader among detector nodes, and records state transitions (suspected → confirmed_down → alive) as incidents in the ledger. External sentinels on other networks probe the ledger's readiness and page a human over their own channel when the whole cluster is unreachable. The pager is the point: silence is detected, not assumed. Automatic rescue is hard-disabled in the current release.

8. What you can audit afterwards

For any event: who sent it (token-bound), the global order it landed in, its per-source order, who received it and when, who applied it and when, and every reply that references it. Retention is a server setting (30 days by default). Single-node deployments keep this in SQLite; three-node deployments replicate it through Raft.

9. Failure modes, by design

Protocol details: PROTOCOL.md. Quick start: README.md. Source: Gene7-Ai/ATalk (MIT).