Reframe README around the real bottleneck (human coordination, not engineering ability) with the "five-minute meeting" cost story; rewrite the demo script to match. Update gemini/livekit/mongodb/digitalocean specs to reflect shipped code (Gemini Live agent, TTS, embeddings, Lyria via Interactions API; current API routes; hermes_jobs collections). Add hermes.md. Rename graph.md -> cont_learning.md. Remove outdated/dead docs (agent-learning scaffolding, graph-discovery, handoffs, superpowers, idea/plan/demo-setup) and the bundled LiveKit starter under examples/. Rewire CLAUDE.md doc-first gate off the deleted PLAN.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
3.6 KiB
LiveKit Integration Spec
Status: active / matches code.
LiveKit is the real-time backbone for PodMan. It carries the screen-share perception input, the intervention data channel, and all room audio (Gemini TTS escalations, the Lyria score, and the live conversation agent). It is load-bearing, not decorative.
Room structure
- One LiveKit room per pod:
room = podId. - Engineers join as named participants (e.g.
alice,bob). - PodMan runs multiple agent identities in/around a room:
podman-hermes— the main vision + intervention agent (@livekit/rtc-node).podman-live-conversation— the Gemini Live voice agent (Python).- short-lived
podman-hermes-job-*publishers for async job events.
- A fixed identity matters: a second
podman-hermesevicts the first and they flap, dropping interventions. systemd keeps exactly one alive in production.
Engineer side (PWA)
Joining:
- PWA calls
POST /api/tokenwith{ podId, identity }→{ token, url }. - LiveKit client connects with the token.
- PWA publishes the screen track via
getDisplayMedia. - PWA enables mic for ambient presence (used by the conversation agent).
Receiving:
- Subscribes to remote agent audio tracks (TTS, Lyria, conversation) and attaches them to a hidden audio sink.
- Browser autoplay restrictions apply: the PWA calls
room.startAudio()from a user gesture (Enable audio,Test PodMan voice,Share screen, first room click). - Listens on the data channel for cards and
VOICE_CUEfallback text.
Data channel listener (PWA):
room.on(RoomEvent.DataReceived, (payload, participant) => {
if (!participant?.identity.startsWith('podman-')) return;
const msg = JSON.parse(new TextDecoder().decode(payload));
// msg.type: COLLISION | ACK | GIT_REPORT | VOICE_CUE | HERMES_JOB_EVENT
appendInterventionToFeed(msg);
});
All data messages share the podman.intervention topic (DATA_TOPIC).
Agent side (podman-hermes)
Framework: @livekit/rtc-node. Code: backend/src/agent/podman.ts,
backend/src/action/hermes.ts, backend/src/voice/live.ts.
- Subscribes to engineers' screen-share tracks and samples frames for Gemini Vision.
- Detects collisions, gates them through the learning policy, then publishes a card/message on the data channel.
- For critical collisions, generates Gemini TTS audio and publishes it as a microphone-source audio track, held for the audio duration plus a tail/hold window so subscribers finish playout. Voice publishing logs frame count, estimated duration, queued playout, and hold time for diagnostics.
Live conversation agent (podman-live-conversation)
Framework: LiveKit Agents for Python (AgentSession, function_tool,
google.realtime.RealtimeModel). Code:
agents/podman-live-conversation/agent.py.
Joins the pod room on demand (POST /api/pods/:id/live-conversation/start),
streams speech-to-speech with Gemini Live, and answers using repo/git/memory
function tools. It can delegate long tasks to the async Hermes job runner and
narrate progress. See docs/hermes.md.
Token endpoint
POST /api/token mints room tokens for engineers and agents alike. Grants:
roomJoin: truecanPublish: true(audio + screen)canPublishData: true(data channel)canSubscribe: true
Short-lived job publishers use canSubscribe: false.
What LiveKit does NOT do in PodMan
- No video tracks published by agents.
- No mic transcription outside the live conversation agent.
- No custom SFU mixing — standard room behavior is sufficient.