68 Commits

Author SHA1 Message Date
sb-iam c0fe424780 docs: add visual HTML demo runbook (ascii architecture + RSI-loop diagrams)
Self-contained one-pager: how-it-works ASCII diagrams (data flow + continual-
learning loop), the 3-min demo timeline, sponsor coverage, last-hour caveats
(suppression disabled, don't redeploy prod), and recovery. Demo-branch only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 13:11:27 -05:00
sb-iam 1a592abb39 docs: add final demo runbook (demo-branch only; not for main)
Reflects the current build: suppression disabled (Ramis), so the demo uses the
'Seen before.' escalate beat + learned_from edge, not the dead 'stays quiet' beat.
Kept on the demo branch so main stays clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 13:04:35 -05:00
Yahya Alhinai 454fdab13b chore(box): snapshot in-progress auth + user-learning WIP before deploy
Captures uncommitted work present on the live droplet so origin/main can be integrated and the new control toolbar deployed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 17:51:21 +00:00
Kartikeya 63a411490c feat(ui): compact, tooltip-driven pod control toolbar
Collapse the five full-width control buttons (audio, mic, background
music, test voice, share) into a compact icon pill with hover/focus
tooltips, plus a labelled Share primary CTA. Row now flex-wraps instead
of a fixed 2-col grid, so it stays on one line on desktop and wraps
cleanly on narrow screens. State is encoded visually (filled = engaged,
ghost = idle, primary = needs action). Background music gets a distinct
Music2 icon instead of reusing the audio Volume2 icon. Handlers are
unchanged — presentation only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 10:44:07 -07:00
Kartikeya faa9769470 fix(intervention): clear stale voice cue/message on new card and dismiss
voiceCue and hermes were never reset, so a new intervention card could show the previous one's spoken line. Clear both when a COLLISION arrives and on respond(); safe against the reliable publish order COLLISION -> HERMES_MESSAGE -> VOICE_CUE.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 10:34:31 -07:00
Kartikeya 79620bccaa fix(collision): treat case/whitespace-variant identities as one person
Vision reports a display name ("Karti") while the git watcher reports a handle ("karti"); the detectors compared engineerId case-sensitively, so the same human collided with themselves. Canonicalize identity at the decision point in both the file and research detectors, keeping the display-cased name for the card.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 10:34:31 -07:00
Ramis d8eaf8d2e2 feat(voice): speak every intervention, not just critical escalations
Voice was gated to severity === 'critical', so most conflict cards were
silent. Now every intervention is spoken: priority is derived from
severity in publishHermesIntervention — critical jumps the queue,
the rest serialize so concurrent alerts don't garble each other.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
2026-06-28 10:31:35 -07:00
Ramis e6e7a42d90 fix(agent): re-alert persistent + new-pair conflicts (README always surfaces)
The edge-trigger single-shot keyed conflicts by file basename only and
re-armed solely when the conflict left a detection cycle. README is dirty
for someone in the pod the whole session, so its key never re-armed: it
alerted once (ram+karti at restart) and stayed muted even as the colliding
pair changed to karti+Ramis Hasanli. Transient files (research, new files)
churned out and back, so only those kept firing.

- conflictKey now includes the sorted engineer pair, so a different pair on
  the same file re-alerts immediately.
- activeConflicts is now Map<key, lastAlertedAt>; a persistent conflict
  re-alerts every CONFLICT_REALERT_MS (default 20s) instead of going silent
  forever, without spamming once-per-frame.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
2026-06-28 10:25:10 -07:00
sb-iam 9658d700fa fix(graph): harden suppressed-repeat activity (Codex deep review) (#44)
Four demo-safety fixes on top of #43, none touching demo-pod data:

- De-dup at source: the suppression path returned before activeConflicts.add(),
  so the same unresolved dismissed collision wrote a SuppressionDoc every frame
  (~194 spam rows observed on prod). Now mark the conflict handled first, so it
  records ONCE per recurrence and re-arms via the onScreenFrame resolution sweep.
- Await the write: recordSuppression is the visible learning proof, so await it
  (like recordCollision/recordIntervention) instead of void ...catch().
- Display de-dup: collapse suppressed beats to one per file (most recent) with a
  stable per-file id, so any pre-fix duplicate rows never render as spam.
- Live-graph gate: suppression beats now satisfy the "has activity" check in
  materializePodGraph, so a clean pod with preserved suppressions (but no
  collision/file nodes) no longer falls back to the demo graph and hides the proof.
- Ops: /api/memory/stats counts `suppressions`; docs/mongodb.md documents the
  collection and flags it "preserve in DB cleanup" (visible learning evidence).

Verified end-to-end on a throwaway pod (not demo-pod): suppression-only pod
materializes; 2 dupe rows render as 1 beat. backend+frontend typecheck + eslint pass.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 10:18:10 -07:00
Ramis 06098f3c4e fix(policy): disable intervention suppression for demo
Dismissing one collision permanently silenced all future ones: the
dismissal gate (priorOutcome.accepted === false) plus loose file/vector
recall meant one README discard muted every later README clash. The
per-pod cooldown additionally hid consecutive conflicts for 3 min.

shouldIntervene now only filters `info` severity — Accept/Dismiss still
record outcomes but no longer gate whether a collision surfaces. The
same-collision single-shot dedupe (activeConflicts in PodmanAgent.handle)
still prevents repeat spam, so removing the cooldown is safe.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
2026-06-28 10:12:51 -07:00
sb-iam 65764f0a1c feat(graph): record + surface suppressed-repeat activity at repeat time (Feature A) (#43)
Codex review fix. The earlier approach synthesized a 'suppressed' beat from every
dismissed OUTCOME, which is wrong: a dismissal is not a later suppressed repeat,
and the old dismissal timestamp sank below the 12-row activity cap (invisible).

Now the negative-feedback loop is recorded WHEN IT HAPPENS. When shouldIntervene()
returns false specifically because a signature was dismissed before
(agent/podman.ts), the agent writes a durable SuppressionDoc to a new
`suppressions` collection, timestamped at the repeat. graph/live.ts materializes
those into kind:'suppressed' activity (recent -> surfaces at the top). The
suppressed collision is never written to `collisions` (agent returns before
recordCollision), so this is its own record.

Verified end-to-end: a transient demo-pod suppression renders at the TOP of the
activity stream; cleaned up after. shared build + backend/frontend typecheck +
eslint pass.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 09:35:56 -07:00
Ramis b6447bb12c docs: surface Gemini in vision agent node
Make it explicit in the infra diagram that the vision agent worker uses
Gemini to read each screen frame.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
2026-06-28 08:18:53 -07:00
Ramis f7521fb5df docs: group engineer-browser edges in infra diagram
Fold the cards/alerts return path into the bidirectional Dev<->LiveKit
edge so the browser shows two grouped routes (app+token, combined live
channel) instead of three.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
2026-06-28 08:18:09 -07:00
Ramis 91403abe9f docs: simplify infra diagram to value-focused paths
Combine Caddy+API and Hermes+detector into single nodes, collapse the
LiveKit room internals, and relabel every edge with the value it carries
(pre-commit signal, nudge before clash, learn from outcomes) instead of
listing every transport hop.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
2026-06-28 08:16:15 -07:00
Ramis f4ad3bec63 docs: split architecture graph into infra + Gemini diagrams
Replace the single all-in-one architecture flowchart with two focused
diagrams: (1) how LiveKit, Hermes, MongoDB, and DigitalOcean wire
together, and (2) Gemini's five surfaces with their triggers, models,
and call sites.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
2026-06-28 08:11:43 -07:00
Ramis 9fb252ec6a feat: add coordination-ROI band to member Work History
Surfaces "rework saved" at the top of the teammate Work History dialog:
a transparent heuristic over collisions Hermes caught early (eligible =
gitOverlap or critical, with an intervention), credit split across
involved engineers. Shows hard counts (clashes caught, conflict-free
files) plus an info-tooltip breakdown of the estimate.

- shared: optional MemberWorkHistoryRoi field (back-compat, self-zeroes)
- backend: query collisions + interventions in getMemberWorkHistory,
  computeRoi helper
- frontend: RoiBand + RoiTooltip components, hidden when no clashes

Implements docs/plans/work-history-roi.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
2026-06-28 07:51:40 -07:00
Ramis 8d17663734 docs: add Work History coordination-ROI build plan
Build-ready spec for adding a "rework saved" ROI band to the member
Work History dialog: transparent heuristic over caught collisions,
additive shared field + backend queries + presentational component.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
2026-06-28 07:46:47 -07:00
Yahya Alhinai 3cb66e3f27 Simplify frontend room surfaces 2026-06-28 14:24:04 +00:00
Yahya Alhinai dec6557cad Add research overlap nudges 2026-06-28 14:21:29 +00:00
Ramis 50cc4a2900 docs: add research-overlap build spec; switch git policy to push-to-main
- docs/plans/research-overlap.md: self-contained build spec for code-edit
  vs research cross-channel overlap nudge (vision classifies editing vs
  research, semantic embedding match, collaboration nudge). Ready for an
  implementing agent.
- CLAUDE.md: replace branch-first habit with push-directly-to-main policy
  plus mandatory pull --rebase before push for concurrent teammates.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
2026-06-28 07:06:24 -07:00
Ramis H. d1570a44d6 Remove hackathon theme from README
Removed the hackathon theme section from the README.
2026-06-28 06:43:32 -07:00
Ramis 9459203ffc Merge docs/cleanup-and-positioning: reposition + spec sync + doc cleanup
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
2026-06-28 06:41:44 -07:00
Ramis 0077e0d7c3 docs: fix mangled README title
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
2026-06-28 06:41:34 -07:00
Ramis 12dbf69431 docs: reposition around team-coordination, sync specs to code, prune stale docs
Reframe README around the real bottleneck (human coordination, not
engineering ability) with the "five-minute meeting" cost story; rewrite the
demo script to match. Update gemini/livekit/mongodb/digitalocean specs to
reflect shipped code (Gemini Live agent, TTS, embeddings, Lyria via
Interactions API; current API routes; hermes_jobs collections). Add hermes.md.
Rename graph.md -> cont_learning.md. Remove outdated/dead docs (agent-learning
scaffolding, graph-discovery, handoffs, superpowers, idea/plan/demo-setup) and
the bundled LiveKit starter under examples/. Rewire CLAUDE.md doc-first gate
off the deleted PLAN.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LuV8W8oNYRsDWKoqK8Mkqc
2026-06-28 06:41:09 -07:00
sb-iam 6f01571edb Merge pull request #41 from karti-ai/feat/rsi-step3-wasreal-verifier
feat(continual-learning): derive wasRealCollision from git overlap (RSI step 3)
2026-06-28 06:07:10 -07:00
sb-iam cf9cca933c Merge pull request #40 from karti-ai/feat/rsi-steps-1-2-negative-loop
feat(continual-learning): activate negative-feedback loop (RSI steps 1-2)
2026-06-28 06:06:11 -07:00
sb-iam 1901010996 fix(continual-learning): harden wasRealCollision verifier (Codex review)
Addresses the P1 brittleness in the Step 3 verifier so learned_from edges
(graph/live.ts:413) can't be silently zeroed on stage.

- Capture overlap AT detection time as Collision.gitOverlap (podman.ts), while
  engineer_states are still fresh, instead of re-deriving from possibly-stale
  state when the user clicks. deriveWasRealCollision() now prefers this stored
  evidence and only falls back to a live re-derivation for older collisions.
- Canonicalize engineer names (trim + lowercase) on both the capture and the
  fallback path, so "Karti" vs "karti" no longer misses the git state.
- conflictKey now reuses the shared comparableBasename() helper (dedupe).

shared/src/collision.ts gains optional `gitOverlap?: boolean` (additive).
shared rebuilt; backend+frontend typecheck + eslint pass. PLAN.md rung 3 updated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 07:57:38 -05:00
sb-iam d9776452b2 feat(continual-learning): derive wasRealCollision from git overlap (RSI step 3)
Backend becomes authoritative for the verifier signal. recordOutcome now
overrides the client-supplied wasRealCollision with deriveWasRealCollision():
a flagged collision counts as REAL only if BOTH named engineers currently have
the collided file in their git changedFiles (getGitStates, 120s freshness TTL).
Conservative false when the collision is orphaned/missing or git state is
stale. Restores the (accepted x wasReal) 2x2 the spec assumes instead of the
hardcoded 107/107 true.

frontend/useInterventions.ts stops sending a hardcoded `true` (now a
backend-overridden placeholder). Spec: continual-learning/spec.md:98-108,
policy.md:35-42. backend+frontend typecheck + eslint pass. PLAN.md P0.5 rung 3
added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 07:47:53 -05:00
sb-iam 0d69f82139 feat(continual-learning): activate negative-feedback loop (RSI steps 1-2)
Step 1 — policy.ts shouldIntervene: suppress on a prior dismissal alone.
The former `&& !priorOutcome.wasRealCollision` term was dead code (outcomes
record wasRealCollision hardcoded true), so the 85 real dismissals in Atlas
were never used. Now a prior accepted===false suppresses the next identical
nudge. Spec: continual-learning/policy.md:41, spec.md:163.

Step 2 — podman.ts handle: only escalate severity to 'critical' (the spoken
alert trigger) when the recalled prior was an accepted *real* collision,
instead of blanket-escalating every recall. Surfaces learned routing in
preferredAction; stops dismissed/false priors over-escalating to voice.
Spec: continual-learning/policy.md:62-63, plan.md:66.

Documentation-first: adds PLAN.md section 8 "P0.5 - RSI negative-feedback
activation" with both rungs + follow-ups. No schema change. backend typecheck
passes. Independent of the MongoDB-cleanup handoff (Codex).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 07:42:42 -05:00
sb-iam 7f922c0332 Merge pull request #39 from karti-ai/docs/team-memory-handoff
docs(handoff): Team memory graph and MongoDB cleanup handoffs
2026-06-28 04:01:34 -07:00
sb-iam 9f509e5f34 docs: add Codex handoff for DB cleanup 2026-06-28 04:00:24 -07:00
sb-iam e7acab2c56 docs(handoff): Team-memory dynamic graph redesign + deploy reconciliation
Fresh-session handoff covering the dynamic force-directed graph, honest de-noised
metrics, the click-to-explain Flow pane, the divergence vs main's parallel impl,
the best-of-both reconciliation (PR #38), build/verify quirks, and gotchas.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 03:17:10 -07:00
sb-iam 1d097b13af Merge pull request #38 from karti-ai/feat/team-memory-deploy
feat(graph): dynamic Team-memory graph + honest metrics (best-of-both on main)
2026-06-28 03:02:02 -07:00
Kartikeya ab8ea07c12 feat: pod-wide Lyria background music (replaces test-audio drums)
The 'Test audio' button becomes 'Background Music': each pod gets a calm looping track from Gemini Lyria 3 that sings the pod name once up front then stays instrumental. Backend GET /api/pods/:id/music generates via the Gemini interactions endpoint and caches the MP3 per pod in Mongo (pod_music); the frontend fetches and loops it via Web Audio, published pod-wide on the existing podman-beat track.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 02:55:19 -07:00
sb-iam b8b6b16531 feat(graph): dynamic force-directed Team-memory graph + honest metrics (on main's pipeline)
Lands the dynamic Team-memory redesign on top of main's continual-learning
pipeline. main already computes loop/activity in the API but its GraphView never
rendered them and kept a static-column graph; this swaps in the dynamic graph and
surfaces the rails, reusing main's richer PodLearningLoop / PodGraphActivity types.

Frontend (new frontend/src/components/graph/*, composed into GraphView.tsx):
- forceSim.ts: dependency-free force layout (charge, link springs, centroid
  recenter, 2-pass collision, bounds clamp, alpha anneal) — no new deps/lockfile churn.
- GraphCanvas.tsx: SVG render from the sim — draggable + pinnable nodes, curved
  edges, weight-sized shapes, fade-in, animated learned_from dash, risk-path
  lighting / rest dimmed, label collision-avoidance.
- MetricsRail / LearningLoop / ActivityStream / SelectedNodePanel / encoding.ts:
  the mock's rails + stream + detail panel in light shadcn. LearningLoop consumes
  main's PodLearningLoop (steps + activeStep); ActivityStream consumes
  PodGraphActivity (title + detail). SelectedNodePanel adds a Flow section that
  narrates the path through a clicked node (flowNarrative); default copy is
  mode-aware. Edge legend rounded out (editing/touches).
- GraphView polls every 5s and diffs (positions preserved), + ws /api/events nudge.
  Stale selection (node gone across a poll) dropped so the canvas can't dim entirely.

Backend (surgical — main's materializer + buildLoop kept):
- live.ts: headline metric cards derived from the FINAL de-noised graph (Open risk
  paths = distinct collision files; Learned owners = distinct owner engineers)
  instead of raw collision-signature / accepted-outcome counts that inflate with
  test churn (50 -> 4 risk files, 16 -> 1 owner on live). buildLoop untouched.
- demo.ts: metrics realigned to the demo graph (3 owners / 1 risk path / 100%).

Verified: lint + -r typecheck + -r build pass; Playwright confirmed the dynamic
graph (ticks on load, draggable), the loop rail (5 steps, ADAPT active) and
activity stream rendering main's shapes, honest metrics (3/1/100%), and the flow
narrative per node kind — zero page errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 02:52:26 -07:00
Kartikeya 604bf9d5ac feat(live-conversation): add git-history tools to repo intelligence
Adds repo_recent_commits (history/authorship, optional path or author) and
repo_find_commits (by commit message or pickaxe code search) so the voice
agent can answer who-changed-what and which-commit-introduced-X over the
local karti-ai/podman checkout.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 02:17:32 -07:00
Kartikeya 87c271995c feat(live-conversation): add search_repo tool over karti-ai/podman
Voice agent can now search the repo (code, symbols, files) via ripgrep over
the live local checkout of main. GitHub code-search index is empty for this
repo, so back it with the always-current clone instead.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 02:05:28 -07:00
Yahya Alhinai 277cee8a7e feat(hermes): add async Gemini handoff jobs 2026-06-28 08:33:52 +00:00
Kartikeya 5ac66c0a02 fix(voice): edge-trigger conflict alerts so Hermes speaks once, not on a loop
Both dedup gates (shouldIntervene per-pod cooldown + hasRecentInterventionForCollision per-file window) were purely time-based, so an unresolved conflict re-alerted every NUDGE_COOLDOWN_MS (30s on prod). Track active conflicts in-memory in PodMan and alert only on the transition into conflict; re-arm when a detection cycle no longer sees it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 01:28:46 -07:00
Yahya Alhinai a80a737ea7 fix(live-conversation): harden agent runtime setup 2026-06-28 08:09:48 +00:00
Yahya Alhinai 97e5af4487 Add member work history and learning docs 2026-06-28 08:06:21 +00:00
Ramis 1adb8413b5 MongoDB plan 2026-06-28 00:41:05 -07:00
Yahya Alhinai 6083f45f05 feat(livekit): upgrade Gemini starter 2026-06-28 07:38:09 +00:00
Yahya Alhinai c4c65f8802 fix(voice): keep LiveKit TTS tracks alive 2026-06-28 07:20:54 +00:00
Kartikeya 46d277d203 feat(frontend): always-visible Enable audio + Enable mic controls
Surface audio-playback unlock and mic publish as persistent buttons in the pod control row (alongside Test audio / Test PodMan voice / Share screen), replacing the conditional 'Sound is off' banner that only appeared after a refresh. Enable mic calls setMicrophoneEnabled, which triggers the browser permission prompt so users can verify their mic is set up right.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 23:55:22 -07:00
Yahya Alhinai aaf1dc5ede fix(voice): throttle repeated pod alerts 2026-06-28 06:49:04 +00:00
Yahya Alhinai c1ac687dcf feat(voice): on-demand PodMan voice test + Hermes notify + TTS tuning
Snapshot of in-progress voice work, committed to unblock concurrent edits: speakInRoom + POST voice-test endpoint, Test PodMan voice button (PodView/api), Hermes notify action + scripts/hermes-notify.mjs, agent exits on LiveKit disconnect for auto-restart, TTS playback tuning (subscriber-ready delay, preroll/tail silence, fallback line, microphone source), verify script updates.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-28 06:45:37 +00:00
Kartikeya 7269098ea3 feat(frontend): shared pod-wide test audio
Make the "Test audio" connectivity check pod-wide instead of local to the
publisher. Shared on/off state is derived from the presence of the podman-beat
track across participants (self-syncing on join/leave). Any participant can
stop it: the owner unpublishes directly; non-owners send a new additive
BEAT_STOP data message that the owner honors. Drives the button label, status
line, and status waveform for everyone, and guards against rapid-double-click
double-publish and mid-start unmount leaks.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 23:39:34 -07:00
Kartikeya e0757088ef fix(frontend): remove PWA service worker to stop demo reloads
Drop vite-plugin-pwa and add an index.html snippet that unregisters any leftover service worker and clears its caches on load. Continuous deploys during the event plus a self-destroying SW were forcing open tabs to reload mid-demo. Browsers self-heal on next load; re-add a precaching PWA post-event.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 23:15:51 -07:00
Kartikeya 1c491866ba fix(frontend): collapse pod join into a single name→Join action
Replace the confusing Join / Add-and-join / +Add trio on PodCard with one flow: type your name, hit Join → added to the roster (deduped server-side) and connected in one step. Removes the join-as-members[0] impersonation footgun where a second person could kick the first.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 23:13:17 -07:00
Yahya Alhinai 562bc1979a fix: stabilize livekit gemini voice playback 2026-06-28 05:59:21 +00:00
Yahya Alhinai 721fe2b8d9 feat: use gemini tts for livekit voice 2026-06-28 05:43:38 +00:00
Yahya Alhinai c771e1594c feat: show screen thumbnails in pod streams 2026-06-28 05:33:00 +00:00
Yahya Alhinai 1080bfd85f feat: label screen observations in pod streams 2026-06-28 05:24:35 +00:00
sb-iam 95a9bde5ee Merge pull request #7 from karti-ai/feat/live-graph-glue
fix(graph): readable Team-memory graph (short labels, clean collision labels, file cap)
2026-06-27 22:17:02 -07:00
Yahya Alhinai dc937a783d feat: fold pod controls into workspace body 2026-06-28 05:14:08 +00:00
Ramis aeb75c310e feat: split pod streams into Signals vs Reasoning sections
My/Team streams were one flat lane; source provenance was buried as a
gray badge by the filename. Group each lane into Signals (vision/git
inputs) and Reasoning & decisions (collision/intervention/outcome), and
promote source to a color-coded provenance chip with a readable kind
label. Presentation-only; no shared-type or backend changes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VpySDjxpM1usXMEngJRWqG
2026-06-27 22:08:53 -07:00
sb-iam 4427e18bb9 Merge pull request #6 from karti-ai/feat/live-graph-glue
feat(graph): live Team memory graph from real Mongo data
2026-06-27 22:02:05 -07:00
Yahya Alhinai 866558e100 feat: stabilize pod sidebars layout 2026-06-28 04:35:24 +00:00
Ramis 4f012aa54a feat: speak collision voice cues via browser Web Speech API
Agent-published LiveKit audio is unreliable for voice: the track is
short-lived (publish, speak ~3s, unpublish) and blocked by browser
autoplay, so participants heard nothing even though the server published
fine. The VOICE_CUE text already arrives over the data channel, so speak
it in the browser via speechSynthesis instead — instant and reliable.

Prime speechSynthesis from user gestures (Enable sound, Share screen) so
later cues are allowed to play.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AaCFWMkYQmTcuPsxaaACft
2026-06-27 21:24:06 -07:00
Ramis cfdbbb3ad7 docs: add production deployment & ops section to CLAUDE.md
Document the box topology so future sessions (and teammates' agents) don't
rediscover it the hard way: systemd-managed services (never launch the agent
manually — duplicate podman-hermes identities evict each other and silently
drop interventions), Caddy serving the static frontend from /var/www/podman,
the deploy procedure, and the shared-package build order.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AaCFWMkYQmTcuPsxaaACft
2026-06-27 21:17:26 -07:00
Ramis 4142876b23 feat: unlock voice audio + terse demo-style collision alerts
Voice generation, critical escalation, and audio publish all worked, but
participants heard nothing: browsers block autoplay of the agent's audio
track until a user gesture. Add the canonical LiveKit unlock — listen for
AudioPlaybackStatusChanged and show an "Enable sound" button that calls
room.startAudio() so PodMan's voice cues are actually heard.

Also make collision alerts short and direct instead of chatty AI prose:
  card:  "Conflict: ram + yahya both on README.md (unpushed). Seen before."
  voice: "Conflict. ram + yahya, both on README.md."

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AaCFWMkYQmTcuPsxaaACft
2026-06-27 21:08:41 -07:00
Yahya Alhinai 16e959072b feat: route pods by URL slug 2026-06-28 04:07:10 +00:00
Yahya Alhinai 563bc6f8fd feat: add collapsible pod stream sidebars 2026-06-28 04:03:25 +00:00
Yahya Alhinai 59f92f18df feat: add realtime pod activity streams 2026-06-28 03:50:56 +00:00
Ramis e474fef228 fix: detect same-file collisions reliably via basename + git overlap
Live observations showed two engineers both dirty on README.md but zero
collisions firing. Two root causes:

1. normalize() only lowercased and prepended src/, so the same file at
   different path depths ("agent.ts" vs "backend/src/agent.ts") never
   matched. Replace with lowercased-basename matching.
2. Git ground truth (engineer_states.changedFiles) was only used to set
   the unpushed flag, never to match the file. The strongest signal was
   wasted.

Detector now fuses vision currentFile AND git changedFiles into one
file->engineers map, so a collision fires when two people share a dirty
file even if both screens aren't on it at the same instant.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AaCFWMkYQmTcuPsxaaACft
2026-06-27 20:44:44 -07:00
Yahya Alhinai dd99050be8 fix: prevent agent frame backpressure 2026-06-28 03:42:08 +00:00
Yahya Alhinai d9e6e8cef4 fix: harden agent worker crash handling 2026-06-28 03:38:50 +00:00
112 changed files with 14537 additions and 4572 deletions
+6
View File
@@ -11,6 +11,7 @@ GEMINI_API_KEY=
# canonical deployment secret name used by DigitalOcean and docs.
GEMINI_VISION_MODEL=gemini-2.0-flash
GEMINI_LIVE_MODEL=gemini-3.1-flash-tts-preview
GEMINI_TTS_VOICE=Charon
GEMINI_EMBEDDING_MODEL=gemini-embedding-001
# --- GitHub (repo state + sync PR artifacts) ---
@@ -25,13 +26,18 @@ VOYAGE_EMBEDDING_MODEL=voyage-4-lite
# --- Backend server ---
PORT=8787
POD_ROOM=demo-pod
CLERK_SECRET_KEY=
# --- Nudge cooldown (ms) — set to 0 during demo if needed ---
NUDGE_COOLDOWN_MS=180000
RESEARCH_OVERLAP_THRESHOLD=0.6
# --- Frontend (Vite — must be VITE_ prefixed to reach the client) ---
VITE_LIVEKIT_URL=wss://your-project.livekit.cloud
VITE_BACKEND_URL=http://localhost:8787
VITE_CLERK_PUBLISHABLE_KEY=
# Keep off by default so users hear Gemini audio delivered through LiveKit.
VITE_ENABLE_BROWSER_TTS_FALLBACK=false
# --- Deployment verification ---
# Optional override when the deployed SPA and API use different origins.
+3
View File
@@ -35,8 +35,11 @@ Thumbs.db
coverage/
.cache/
.turbo/
__pycache__/
*.py[cod]
# Ramis
.remember/
.claude/
.hermes/
.playwright-mcp/
+1
View File
@@ -1,3 +1,4 @@
link-workspace-packages=true
prefer-workspace-packages=true
auto-install-peers=true
prefix=/home/ramis/.npm-global
+5
View File
@@ -1,7 +1,12 @@
node_modules
.venv
**/.venv
.pytest_cache
**/.pytest_cache
dist
build
pnpm-lock.yaml
.omc
*.log
.agents
examples/livekit-gemini-hacker-starter
+77 -15
View File
@@ -318,30 +318,31 @@ Everything should serve that outcome.
## Documentation-first enforcement — HARD RULE
**Every line of code must trace back to a task in `docs/PLAN.md` or a spec in `docs/`.**
**Every line of code must trace back to a spec in `docs/`.**
This is not a guideline. This is a gate.
This is not a guideline. This is a gate. The canonical specs are
`docs/gemini.md`, `docs/livekit.md`, `docs/mongodb.md`, `docs/cont_learning.md`,
`docs/hermes.md`, and `docs/digitalocean.md`. `docs/demo.md` is the demo script.
### Before writing any code, verify:
1. **Is this task in `docs/PLAN.md`?** Find the exact task number. If it's not there, stop.
2. **Is the approach consistent with the relevant spec?** Check `docs/gemini.md`, `docs/livekit.md`, `docs/mongodb.md`, `docs/digitalocean.md` as applicable.
3. **Do the file names and API shapes match what's documented?** If the plan says `backend/src/db/states.ts`, do not create `backend/src/database/engineStates.ts` without updating the spec first.
1. **Is the approach consistent with the relevant spec?** Check `docs/gemini.md`, `docs/livekit.md`, `docs/mongodb.md`, `docs/cont_learning.md`, `docs/hermes.md`, `docs/digitalocean.md` as applicable.
2. **Do the file names and API shapes match what's documented?** If a spec says `backend/src/memory/store.ts`, do not create `backend/src/database/engineStates.ts` without updating the spec first.
### If a developer asks for something not in the plan:
### If a developer asks for something not in the specs:
**Do not write the code.** Instead:
1. Say explicitly: _"This isn't in the current plan. Let me understand what you're trying to do."_
1. Say explicitly: _"This isn't in the current specs. Let me understand what you're trying to do."_
2. Ask what problem they're solving and whether it's required for the demo path.
3. Evaluate whether it fits within scope or replaces something planned.
4. If it's valid: **update `docs/PLAN.md` and the relevant spec first**, then proceed to code.
5. If it's scope creep: say so directly and recommend the nearest in-plan alternative.
3. Evaluate whether it fits within scope or replaces something documented.
4. If it's valid: **update the relevant spec first**, then proceed to code.
5. If it's scope creep: say so directly and recommend the nearest in-spec alternative.
### Signs a request is off-plan (stop and consult):
### Signs a request is off-spec (stop and consult):
- Introducing a new file not mentioned in any task's **Files** list
- Changing an API signature documented in a spec (`/ingest`, `/health`, `/pods/:podId/token`, `/pods/:podId/state`)
- Introducing a new file or API route not described in any spec
- Changing a documented API signature (`/health`, `POST /api/token`, `POST /api/outcome`, `GET /api/pods/:id/...`)
- Adding a dependency not in the existing `package.json` files without a clear spec reason
- Building a feature in the **Cut immediately** list
- Touching another engineer's ownership area without explicit cross-team coordination
@@ -359,7 +360,68 @@ This repo is actively used by **4 engineers at the same time**. Claude sessions
### What this means for how you help
- **Assume other files are actively being edited.** Never refactor code outside the immediate task scope without explicit coordination from the user.
- **Treat integration points as contracts.** The shared types in `shared/src/` and the API shapes of `POST /ingest`, `GET /pods/:podId/token`, and `GET /pods/:podId/state` are the interfaces between all teammates — do not change their signatures unilaterally.
- **Treat integration points as contracts.** The shared types in `shared/src/` and the API shapes of `POST /api/token`, `POST /api/outcome`, and the `GET /api/pods/:id/*` routes are the interfaces between all teammates — do not change their signatures unilaterally.
- **Flag merge risk explicitly** before editing a shared file (e.g., `backend/src/index.ts`, `frontend/src/App.tsx`). Say so, then proceed only if the user confirms.
- **Prefer additive changes** — new files, new functions — over modifying existing ones. This minimizes merge conflicts in a concurrent team.
- **When proposing new files**, verify they match the file names listed in the relevant task in `docs/PLAN.md`. Do not invent new paths.
- **When proposing new files**, verify they match the file names and paths described in the relevant spec in `docs/`. Do not invent new paths.
### Git workflow — push directly to `main`, stay in sync
This repo **does not use feature branches**. Commit straight to `main` and push.
There is no branch-first step. Because several people push concurrently:
- **Always sync before pushing:** `git pull --rebase origin main` immediately
before `git push`. Never force-push `main`.
- **Keep commits small and additive** so rebases stay clean — prefer new files
and new functions over editing shared hot files (`backend/src/agent/podman.ts`,
`backend/src/server.ts`, `frontend/src/App.tsx`).
- **If a rebase conflicts**, resolve it locally and re-run the pull-rebase before
pushing; do not overwrite a teammate's commit.
- Commit/push only when the user asks (overrides any default branch-first habit).
---
## Production deployment & ops — READ BEFORE TOUCHING THE SERVER
The live system runs on a DigitalOcean droplet at `165.22.129.249`
(public: `https://165-22-129-249.sslip.io/` and `podman.live`). Repo on box:
`/root/podman`. SSH is `root@165.22.129.249` (password auth; ask the team for
the password — it is **not** stored in the repo).
### HARD RULE: manage processes with systemd, never manual `node`
The backend API and the LiveKit agent run as **systemd services** with
`Restart=always`:
- `podman-platform-api.service``node backend/dist/server.js` (cwd `/root/podman`)
- `podman-platform-agent.service``node dist/agent.js` (cwd `/root/podman/backend`,
`Environment=POD_ROOM=demo-pod`, `EnvironmentFile=/root/podman/backend/.env`)
**Do not start the agent or server by hand** (`node ...`, `nohup`, `setsid`,
`tsx`). The agent joins LiveKit with a fixed identity (`podman-hermes`); a
second instance with the same identity **evicts the first from the room**, and
they flap forever — silently dropping every intervention/voice update. systemd
keeps exactly one of each alive. If you launched a manual process, kill it and
let systemd own the singleton.
### Frontend is static, served by Caddy
The frontend is a Vite build served by **Caddy** from `/var/www/podman`
(`/etc/caddy/Caddyfile`). Caddy reverse-proxies `/api/*` and `/health` to
`127.0.0.1:8787`. The LiveKit URL reaches the browser via the backend
`/api/token` response, **not** `VITE_LIVEKIT_URL` (intentionally empty in
`frontend/.env`).
### Deploy procedure (run on the box)
```bash
cd /root/podman && git pull && pnpm -r build
rm -rf /var/www/podman/* && cp -r frontend/dist/* /var/www/podman/
systemctl restart podman-platform-api podman-platform-agent
systemctl status podman-platform-agent --no-pager # verify it came up
```
`pnpm -r build` order matters: `@podman/shared` builds first, or backend/frontend
typecheck fails with "Cannot find module '@podman/shared'". MongoDB is
**mandatory** — both services ping Mongo at boot and exit loudly if it is
unreachable (intentional; fix the `.env` creds, do not re-add silent fallbacks).
+480 -179
View File
@@ -1,90 +1,263 @@
# PodMan - Real-time AI Team Coordination Agent
# PodMan
[![TypeScript](https://img.shields.io/badge/TypeScript-6.x-3178C6?logo=typescript&logoColor=white)](https://www.typescriptlang.org/)
[![React](https://img.shields.io/badge/React-19-61DAFB?logo=react&logoColor=111)](https://react.dev/)
[![LiveKit](https://img.shields.io/badge/LiveKit-realtime-000000?logo=livekit&logoColor=white)](https://livekit.io/)
[![MongoDB](https://img.shields.io/badge/MongoDB-memory-47A248?logo=mongodb&logoColor=white)](https://www.mongodb.com/)
[![Gemini](https://img.shields.io/badge/Gemini-vision-8E75B2?logo=googlegemini&logoColor=white)](https://ai.google.dev/)
[![DigitalOcean](https://img.shields.io/badge/DigitalOcean-deploy-0080FF?logo=digitalocean&logoColor=white)](https://www.digitalocean.com/)
**An ambient pair programmer for engineering teams.**
**2026 AI Engineer World's Fair Hackathon** - Track: **Continual Learning**
PodMan watches the work happening inside a shared LiveKit room, understands what
each engineer is doing, remembers which interventions helped, and nudges the
team before duplicated work, merge collisions, or missed handoffs slow everyone
down.
PodMan is a non-intrusive AI teammate for active coding. It watches consented
LiveKit screen-share context, combines it with local git truth and shared team
memory, and coordinates teammates before a problem becomes a GitHub problem.
[LiveKit](https://livekit.io/) · [MongoDB](https://www.mongodb.com/) ·
[Gemini](https://ai.google.dev/) · [Hermes](https://hermes-agent.nousresearch.com/) ·
[vLLM](https://vllm.ai/) · [DigitalOcean](https://www.digitalocean.com/)
> GitHub sees pushed work. PodMan sees work while it is still happening.
## Read This First
PodMan is not a dashboard and not a raw screenshot analyzer. Its job is to
notice useful coordination moments, remember what helped before, and route the
least intrusive intervention: a small card first, a Hermes message when teammates
need coordination, and voice only for urgent escalation.
| Need | Use this |
| ----------------------- | --------------------------------------------------------------------------------------------- |
| Open the product | `https://podman.live` |
| Check the app | `curl https://podman.live/health` |
| Check pods and memory | `curl https://podman.live/api/pods && curl https://podman.live/api/memory/stats` |
| Use the LLM externally | Base URL `https://llm.alhinai.dev/v1`, model `gemma-4-31B-it` |
| Test Hermes | `hermes -z 'Reply with exactly: working' --provider gemma4-31b-vllm --model gemma-4-31B-it` |
| Start local development | API, vision agent, and frontend commands are in [Local Development](#local-development) |
| Debug production | Public checks first, then systemd services in [Production Operations](#production-operations) |
---
## Architecture
## Current Live System
PodMan is split into a browser PWA, an HTTP API service, a LiveKit agent worker,
and a persistence/action layer. The screen signal flows through LiveKit, not a
manual screenshot upload endpoint.
| Surface | Running now | Purpose |
| ------------- | ---------------------------- | -------------------------------------- |
| Product | `https://podman.live` | Team room, screen context, cards |
| API | `https://podman.live/api/*` | Pods, tokens, outcomes, memory |
| Health | `https://podman.live/health` | Backend readiness |
| Local API | `127.0.0.1:8787` | Express service behind Caddy |
| Reasoning LLM | `https://llm.alhinai.dev/v1` | OpenAI-compatible Gemma/vLLM endpoint |
| API key | `not-needed` | Placeholder key for OpenAI clients |
| Hermes model | `gemma-4-31B-it` | 262K-context tool-using agent |
| Hermes config | `gemma4-31b-vllm` | Custom provider used by Hermes locally |
```mermaid
flowchart TB
Browser["Engineer Browser<br/>React + Vite PWA"]
Room["LiveKit Room<br/>screen share + audio + data messages"]
subgraph Droplet["DigitalOcean PodMan Droplet"]
Caddy["Caddy<br/>static app + /api proxy"]
API["Express API<br/>127.0.0.1:8787"]
Vision["Vision Agent<br/>screen frame observer"]
Voice["Live Conversation Agent<br/>Python + Gemini Live"]
Ops["Hermes Ops Timers<br/>watchdog + sync deploy"]
end
subgraph Memory["MongoDB Atlas"]
Observations["observations"]
State["engineer_states"]
Outcomes["interventions + outcomes"]
end
subgraph Google["Google Gemini APIs"]
GeminiVision["Vision"]
GeminiVoice["Live voice + TTS"]
GeminiEmbed["Embeddings fallback"]
Lyria["Music"]
end
subgraph Reasoning["External Reasoning Endpoint"]
Tunnel["Cloudflare Tunnel<br/>llm.alhinai.dev"]
VLLM["vLLM OpenAI Server<br/>gemma-4-31B-it<br/>262144 context"]
Hermes["Hermes Agent<br/>provider: gemma4-31b-vllm"]
end
Browser --> Caddy --> API
Browser <-->|screen, audio, cards| Room
API --> Room
API --> Observations
API --> State
API --> Outcomes
Vision <-->|screen tracks| Room
Vision --> GeminiVision
Vision --> Observations
Vision --> Outcomes
Voice <-->|conversation| Room
Voice --> GeminiVoice
Vision --> GeminiEmbed
Voice --> GeminiEmbed
Ops --> Hermes --> Tunnel --> VLLM
API --> Hermes
Lyria --> Voice
classDef user fill:#e8f1ff,stroke:#3366cc,color:#0b1f44;
classDef app fill:#eef8ee,stroke:#2f8a3a,color:#123915;
classDef data fill:#fff6df,stroke:#c47f00,color:#3d2b00;
classDef ai fill:#f4edff,stroke:#805ad5,color:#2d1857;
class Browser,Room user;
class Caddy,API,Vision,Voice,Ops app;
class Observations,State,Outcomes data;
class GeminiVision,GeminiVoice,GeminiEmbed,Lyria,Tunnel,VLLM,Hermes ai;
```
## What It Does
```mermaid
flowchart LR
subgraph Laptop["Engineer laptop"]
PWA["React PWA"]
Screen["Screen share track"]
Git["Git watcher<br/>scripts/podman-agent.mjs"]
end
subgraph Realtime["LiveKit room"]
Room["Pod room"]
Data["Data topic<br/>podman.intervention"]
end
subgraph Backend["PodMan backend"]
API["API service<br/>/api/token /api/pods /api/outcome"]
Agent["Agent worker<br/>@livekit/rtc-node"]
Vision["Gemini Vision<br/>structured JSON"]
Detector["Coordination detector<br/>collisions, blockers, dead ends"]
end
subgraph Memory["Memory and actions"]
Mongo["MongoDB<br/>observations, outcomes, pods"]
GitHub["GitHub<br/>repo state + sync PR artifact"]
Hermes["Hermes action layer<br/>cards, messages, urgent voice"]
end
PWA -->|"POST /api/token"| API
API -->|"LiveKit JWT"| PWA
PWA --> Screen
Screen --> Room
Room --> Agent
Agent --> Vision
Vision --> Detector
Git --> Mongo
Detector --> Mongo
Detector --> GitHub
Detector --> Hermes
Hermes --> Data
Data --> PWA
PWA -->|"POST /api/outcome"| API
API --> Mongo
A["Engineer shares screen"] --> B["Gemini extracts work context"]
B --> C["PodMan detects overlap<br/>files, symbols, research, unpushed work"]
C --> D["MongoDB recalls<br/>similar prior events"]
D --> E{"Policy gate"}
E -->|"seen false alarm"| F["stay quiet"]
E -->|"seen real collision"| G["raise urgency"]
E -->|"new useful signal"| H["show card or message"]
G --> I["Hermes / voice escalation"]
H --> J["teammate accepts or dismisses"]
I --> J
J --> K["outcome becomes future memory"]
K --> D
```
### Runtime shape
PodMan removes the reason to interrupt. It gives the team a live picture of work
in progress, then improves from every accepted or dismissed intervention.
> GitHub sees pushed work. PodMan sees work while it is still happening.
---
## The Model Stack
PodMan uses multiple AI surfaces. They are intentionally split by job.
```mermaid
flowchart LR
Work["Screen + room activity"] --> VisionRoute["Perception route"]
VoiceInput["Engineer speech"] --> ConversationRoute["Conversation route"]
OpsNeed["Ops / autonomous task"] --> HermesRoute["Reasoning route"]
MemoryNeed["Similarity search"] --> EmbedRoute["Memory route"]
Ambient["Session atmosphere"] --> MusicRoute["Audio route"]
VisionRoute --> Gemini20["gemini-2.0-flash<br/>screen understanding"]
ConversationRoute --> GeminiLive["gemini-3.1-flash-live-preview<br/>live Q&A"]
ConversationRoute --> GeminiTTS["gemini-3.1-flash-tts-preview<br/>urgent speech"]
HermesRoute --> Gemma["gemma-4-31B-it<br/>Hermes + vLLM + tools"]
EmbedRoute --> Voyage["voyage-4-lite<br/>primary embeddings"]
EmbedRoute --> GeminiEmbed["gemini-embedding-001<br/>fallback embeddings"]
MusicRoute --> Lyria["lyria-3-clip-preview<br/>background music"]
classDef signal fill:#e8f1ff,stroke:#3366cc,color:#0b1f44;
classDef route fill:#eef8ee,stroke:#2f8a3a,color:#123915;
classDef model fill:#f4edff,stroke:#805ad5,color:#2d1857;
class Work,VoiceInput,OpsNeed,MemoryNeed,Ambient signal;
class VisionRoute,ConversationRoute,HermesRoute,EmbedRoute,MusicRoute route;
class Gemini20,GeminiLive,GeminiTTS,Gemma,Voyage,GeminiEmbed,Lyria model;
```
### Gemma 4 31B via vLLM
Hermes uses the external OpenAI-compatible endpoint:
```text
Base URL: https://llm.alhinai.dev/v1
API key: not-needed
Model: gemma-4-31B-it
Context: 262144 tokens
```
The vLLM server is configured for Hermes-style tool use:
```text
--max-model-len 262144
--max-num-seqs 1
--max-num-batched-tokens 16384
--gpu-memory-utilization 0.85
--kv-cache-dtype fp8
--enable-auto-tool-choice
--tool-call-parser gemma4
--enable-chunked-prefill
```
Hermes should point at that endpoint with this provider shape:
```yaml
model:
default: gemma-4-31B-it
provider: gemma4-31b-vllm
providers:
gemma4-31b-vllm:
name: Gemma 4 31B vLLM (256K)
api: https://llm.alhinai.dev/v1
api_key: not-needed
transport: chat_completions
default_model: gemma-4-31B-it
discover_models: true
models:
gemma-4-31B-it:
context_length: 262144
agent:
tool_use_enforcement: auto
```
Why this matters: Hermes sends OpenAI tool schemas and `tool_choice: "auto"`.
Without `--enable-auto-tool-choice` and `--tool-call-parser gemma4`, vLLM returns
HTTP 400 before Hermes can initialize an agent.
### Gemini Surfaces
Gemini remains the realtime perception and voice layer inside PodMan.
| Use | Model | Code |
| -------------------------- | ------------------------------- | ------------------------------------------ |
| Screen understanding | `gemini-2.0-flash` | `backend/src/vision/gemini.ts` |
| Urgent spoken alerts | `gemini-3.1-flash-tts-preview` | `backend/src/voice/live.ts` |
| Live room conversation | `gemini-3.1-flash-live-preview` | `agents/podman-live-conversation/agent.py` |
| Memory embeddings fallback | `gemini-embedding-001` | `backend/src/memory/vectors.ts` |
| Ambient background music | `lyria-3-clip-preview` | `backend/src/voice/music.ts` |
Voyage embeddings can be used first when `VOYAGE_API_KEY` is set. Gemini
embeddings remain the fallback. If no embedding provider is available, PodMan
falls back to exact-signature matching.
---
## How The Learning Loop Works
The learning loop is the product. A teammate only has to accept or dismiss an
intervention; the rest is captured automatically.
```text
observe -> detect -> recall prior outcomes -> policy gate -> act -> record outcome
^ |
+-------------------------- next recall -----------------------------+
```
| Stage | What happens | Code |
| ------- | --------------------------------------------------------------------------- | ----------------------------------- |
| Observe | Gemini Vision turns sampled screen frames into structured work context. | `backend/src/vision/gemini.ts` |
| Detect | PodMan detects overlapping files, symbols, research, and unpushed work. | `backend/src/collision/detector.ts` |
| Recall | MongoDB Atlas recalls similar prior events and outcomes. | `backend/src/memory/vectors.ts` |
| Gate | Policy suppresses dismissed false alarms and escalates recurring real ones. | `backend/src/memory/policy.ts` |
| Act | PodMan publishes a card, Hermes message, or urgent voice cue. | `backend/src/action/hermes.ts` |
| Learn | Accept/dismiss feedback is written back to memory. | `backend/src/memory/store.ts` |
---
## Runtime Components
| Layer | Runtime | Responsibility |
| ------------ | ------------------- | --------------------------------------------------------------------------------------------------- |
| Frontend PWA | React + Vite | Join pods, publish screen share, show live room state, render interventions |
| Backend API | Express | Mint LiveKit tokens, manage pods, record outcomes, expose memory stats, create sync PR artifacts |
| Agent worker | `@livekit/rtc-node` | Join the room as PodMan, subscribe to screen-share tracks, sample frames, publish intervention data |
| Vision loop | Gemini | Convert sampled IDE frames into structured work context |
| Team memory | MongoDB | Store observations, collisions, interventions, outcomes, pods, and git watcher state |
| Git watcher | Node script | Poll each laptop's local git state so dirty/unpushed work is not guessed from vision alone |
| Action layer | Hermes concept | Route cards, teammate messages, optional research summaries, and urgent voice escalation |
| Deployment | DigitalOcean | Static site for frontend, HTTP service for API, worker for the LiveKit agent |
| ------------------ | ----------------------------------- | -------------------------------------------------------------------- |
| Frontend | React + Vite | Pod rooms, screen share, cards, voice controls, member state |
| Backend API | Express on `:8787` | LiveKit tokens, pod CRUD, outcomes, memory stats, sync PRs |
| Vision agent | Node + `@livekit/rtc-node` | Subscribes to screen tracks, samples frames, publishes interventions |
| Live voice agent | Python LiveKit Agents + Gemini Live | Real-time voice Q&A in a pod room |
| Memory | MongoDB Atlas | Observations, engineer state, collisions, interventions, outcomes |
| Reasoning agent | Hermes + Gemma vLLM | Tool-using autonomous assistant and ops layer |
| Realtime transport | LiveKit Cloud | Screen tracks, audio tracks, data messages |
| Hosting | DigitalOcean + Caddy + systemd | Static frontend, API, workers, watchdog timers |
### Data flow
---
## Data Flow
```mermaid
sequenceDiagram
@@ -92,150 +265,106 @@ sequenceDiagram
participant Dev as Engineer PWA
participant API as Backend API
participant LK as LiveKit Room
participant Agent as PodMan Agent
participant Gemini as Gemini Vision
participant Mongo as MongoDB Memory
participant Hermes as Hermes / Action Layer
participant GH as GitHub
participant Agent as Vision Agent
participant Gemini as Gemini APIs
participant Mongo as MongoDB Atlas
participant Hermes as Hermes/Gemma
Dev->>API: POST /api/token
API-->>Dev: LiveKit URL + JWT
Dev->>LK: Join pod room
Dev->>LK: Publish screen-share track
Dev->>LK: Join pod room + publish screen share
Agent->>LK: Subscribe to screen-share video
Agent->>Gemini: Sampled JPEG frame
Agent->>Gemini: Sampled frame
Gemini-->>Agent: Structured work context
Agent->>Mongo: Record observation
Agent->>GH: Read public repo state
Agent->>Mongo: Recall prior patterns
Agent->>Hermes: Create intervention
Hermes->>LK: Publish small data packet
LK-->>Dev: Render card / message / urgent voice cue
Agent->>Mongo: Store observation
Agent->>Mongo: Recall similar prior events
Mongo-->>Agent: Prior outcome + policy hints
Agent->>Hermes: Escalate when autonomous help is useful
Agent->>LK: Publish card / message / voice cue
Dev->>API: POST /api/outcome
API->>Mongo: Store learning signal
API->>Mongo: Store accept/dismiss feedback
```
### Why this architecture matters
- **LiveKit is the realtime spine.** Screens and intervention data move through a
shared room, so PodMan can react before code is pushed.
- **Gemini is the perception layer.** The agent samples frames and asks Gemini
for structured JSON such as current file, symbol, activity, unpushed hints,
and confidence.
- **MongoDB is the learning loop.** Outcomes and repeated patterns make later
interventions quieter and more useful.
- **Local git is the truth source.** The watcher reports dirty files and branch
state directly from each laptop, which avoids relying on vision for facts
GitHub cannot see.
- **Hermes keeps it non-intrusive.** Most events are cards. Team messages and
voice are escalation paths, not the default.
---
## How it works
1. Engineers open the PWA and join a pod room.
2. The backend API mints a LiveKit token via `POST /api/token`.
3. The PWA publishes screen share into the pod room when the engineer chooses
"Share my screen".
4. The PodMan agent worker joins the same room and subscribes to screen-share
tracks.
5. The agent samples frames, sends them to Gemini Vision, and records structured
observations in MongoDB.
6. Each engineer runs the git watcher so PodMan has deterministic dirty/unpushed
state.
7. The detector combines live screen context, git truth, GitHub state, and team
memory.
8. PodMan sends the smallest useful intervention: card first, Hermes message for
coordination, voice only when urgent.
9. The user's response is saved as an outcome, closing the continual-learning
loop.
---
## Public interfaces
## Public Interfaces
| Interface | Purpose |
| ----------------------------------------------------------- | ---------------------------------------------- |
| `GET /health` | API health check |
| `POST /api/token` | Mint LiveKit room tokens |
| `POST /api/sync-pr` | Create a visible sync PR artifact |
| `POST /api/outcome` | Store accepted/dismissed intervention outcomes |
| `GET /api/memory/stats` | Show memory collection counts |
| `GET /api/pods` | List pods |
| `GET/POST/PATCH/DELETE /api/pods` | Pod CRUD |
| `POST/DELETE /api/pods/:id/members` | Pod membership |
| `GET /api/pods/:id/members/:name/history` | Recent member work history |
| `POST /api/outcome` | Store accepted/dismissed intervention outcomes |
| `GET /api/memory/stats` | Live memory collection counts |
| `POST /api/sync-pr` | Create a visible sync PR artifact |
| LiveKit topic `podman.intervention` | Intervention data channel |
| Wire messages `COLLISION`, `ACK`, `GIT_REPORT`, `VOICE_CUE` | Shared agent/PWA message contract |
| Wire messages `COLLISION`, `ACK`, `GIT_REPORT`, `VOICE_CUE` | Agent/PWA contract |
---
## Monorepo layout
## Monorepo Layout
| Folder | What |
| ----------- | -------------------------------------------------------------------------------- |
| `frontend/` | React + Vite PWA for pods, LiveKit room UI, screen share, and intervention cards |
| `backend/` | Express API plus separate LiveKit agent worker |
| `shared/` | Shared TypeScript types and LiveKit data message contracts |
| Folder | Purpose |
| ----------- | ----------------------------------------------------------------------- |
| `frontend/` | React + Vite PWA |
| `backend/` | Express API, vision agent, memory, collision detection, Hermes job APIs |
| `agents/` | Python LiveKit conversation agent |
| `shared/` | Shared TypeScript contracts |
| `database/` | MongoDB setup and seed utilities |
| `infra/` | DigitalOcean App Platform specs and Dockerfile |
| `scripts/` | Local git watcher for demo laptops |
| `docs/` | Canonical plan and deeper sponsor/integration notes |
| `infra/` | Caddy, Docker, DigitalOcean, systemd units |
| `scripts/` | Git watcher, deploy doctor, watchdog, verification tooling |
| `docs/` | Demo, deployment, learning, graph, and architecture notes |
---
## Docs
## Local Development
| File | What |
| ---------------------------------------------- | ----------------------------------------- |
| [`docs/PLAN.md`](docs/PLAN.md) | Canonical master plan and source of truth |
| [`docs/idea.md`](docs/idea.md) | Product concept and demo framing |
| [`docs/livekit.md`](docs/livekit.md) | LiveKit notes and room model |
| [`docs/gemini.md`](docs/gemini.md) | Gemini vision and voice notes |
| [`docs/mongodb.md`](docs/mongodb.md) | MongoDB memory design |
| [`docs/digitalocean.md`](docs/digitalocean.md) | Deployment notes |
| [`docs/demo-setup.md`](docs/demo-setup.md) | Demo laptop and stage checklist |
Run the core app in three terminals:
---
```mermaid
flowchart LR
Env["1. Configure .env"] --> Install["2. pnpm install"]
Install --> API["Terminal A<br/>backend API :8787"]
Install --> Agent["Terminal B<br/>vision agent"]
Install --> UI["Terminal C<br/>frontend :5173"]
Install --> Voice["Optional<br/>conversation agent"]
API --> Browser["Open local PWA"]
Agent --> Browser
UI --> Browser
Voice --> Browser
## Prizes targeted
- **Best Gemini:** structured vision over live IDE context, with voice as an
optional escalation path.
- **Best LiveKit:** realtime screen-share tracks, presence, data packets, and
eventual voice in one pod room.
- **Best DigitalOcean:** frontend static site, API service, and LiveKit agent
worker deployment.
- **MongoDB + Voyage story:** persistent memory first, vector recall once exact
signature recall is proven.
---
## Quick start
classDef step fill:#eef8ee,stroke:#2f8a3a,color:#123915;
classDef run fill:#e8f1ff,stroke:#3366cc,color:#0b1f44;
class Env,Install step;
class API,Agent,UI,Voice,Browser run;
```
```bash
cp .env.example .env
# fill in LIVEKIT_*, GEMINI_*, GITHUB_*, and MONGODB_URI
# Fill LIVEKIT_*, GEMINI_*, GITHUB_*, MONGODB_URI.
pnpm install
pnpm --filter @podman/backend dev # API on :8787
pnpm --filter @podman/backend dev:agent # PodMan LiveKit agent
pnpm --filter @podman/backend dev:agent # LiveKit vision agent
pnpm --filter @podman/frontend dev # PWA on :5173
```
---
## Git watcher - run this on every demo laptop
Each engineer runs this in a terminal before the demo. It polls the local git
working tree every 15 seconds and writes git state to MongoDB so PodMan has
deterministic dirty/unpushed truth that vision alone cannot reliably infer.
Run the Python live conversation agent:
```bash
# from the repo root
node scripts/podman-agent.mjs --name <yourname> --pod <podId>
pnpm livekit:conversation:agent
```
**Demo setup:**
Run the local git watcher on each demo laptop:
```bash
node scripts/podman-agent.mjs --name <engineer-name> --pod <pod-id>
```
Demo identities:
```bash
node scripts/podman-agent.mjs --name alice --pod demo-pod
@@ -243,12 +372,184 @@ node scripts/podman-agent.mjs --name bob --pod demo-pod
node scripts/podman-agent.mjs --name carol --pod demo-pod
```
The script logs one line per cycle: branch, changed file count, and latest
commit. Leave it running in a background terminal tab throughout the session.
Stop with `Ctrl+C`.
---
**Requirements:**
## Production Operations
- `MONGODB_URI` must be exported in the shell or present in `backend/.env`.
- Run `pnpm install` first so workspace dependencies are available.
- Run from the repo root.
The production droplet is systemd-supervised. Caddy serves the built frontend
and proxies `/api/*` to the backend on `127.0.0.1:8787`.
```mermaid
flowchart LR
Public["https://podman.live"] --> Caddy["caddy.service"]
Caddy --> Static["/var/www/podman"]
Caddy --> API["podman-platform-api.service<br/>:8787"]
API --> Agent["podman-platform-agent.service"]
API --> Voice["podman-live-conversation-agent.service"]
Watchdog["podman-hermes-watchdog.timer"] --> Public
Sync["podman-hermes-sync-deploy.timer"] --> API
classDef public fill:#e8f1ff,stroke:#3366cc,color:#0b1f44;
classDef service fill:#eef8ee,stroke:#2f8a3a,color:#123915;
classDef timer fill:#fff6df,stroke:#c47f00,color:#3d2b00;
class Public public;
class Caddy,Static,API,Agent,Voice service;
class Watchdog,Sync timer;
```
| Service / timer | Purpose |
| ---------------------------------------- | -------------------------------------------------------------- |
| `podman-platform-api.service` | Built backend API on port `8787` |
| `podman-platform-agent.service` | Node LiveKit vision agent |
| `podman-live-conversation-agent.service` | Python Gemini Live conversation agent |
| `podman-hermes-watchdog.timer` | Periodic public health and remediation |
| `podman-hermes-sync-deploy.timer` | Clean-tree fast-forward deploy loop |
| `caddy.service` | Serves `/var/www/podman`, proxies `/api/*` to `127.0.0.1:8787` |
Check the app from the outside first:
```bash
curl https://podman.live/
curl https://podman.live/health
curl https://podman.live/api/pods
curl https://podman.live/api/presence
curl https://podman.live/api/memory/stats
```
Then check the droplet services:
```bash
systemctl is-active podman-platform-api podman-platform-agent
systemctl is-active podman-live-conversation-agent
systemctl is-active podman-hermes-watchdog.timer podman-hermes-sync-deploy.timer
```
Hermes operations scripts:
```bash
pnpm hermes:watchdog
pnpm hermes:watchdog:strict
pnpm hermes:sync-deploy
pnpm deploy:doctor:strict
```
Gemma/vLLM endpoint checks:
```bash
curl https://llm.alhinai.dev/v1/models \
-H "Authorization: Bearer not-needed"
curl https://llm.alhinai.dev/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer not-needed" \
-d '{
"model": "gemma-4-31B-it",
"messages": [{"role": "user", "content": "Reply with exactly: working"}],
"temperature": 0,
"max_tokens": 512
}'
```
---
## Required Environment
```bash
LIVEKIT_URL=wss://your-livekit-server.livekit.cloud
LIVEKIT_API_KEY=...
LIVEKIT_API_SECRET=...
LIVEKIT_CONVERSATION_AGENT_NAME=podman-live-conversation
GEMINI_API_KEY=...
GEMINI_VISION_MODEL=gemini-2.0-flash
GEMINI_LIVE_MODEL=gemini-3.1-flash-tts-preview
GEMINI_CONVERSATION_MODEL=gemini-3.1-flash-live-preview
GEMINI_EMBEDDING_MODEL=gemini-embedding-001
GEMINI_TTS_VOICE=Charon
GITHUB_TOKEN=...
GITHUB_REPO=karti-ai/podman
MONGODB_URI=mongodb+srv://...
VOYAGE_API_KEY=...
VOYAGE_EMBEDDING_MODEL=voyage-4-lite
PORT=8787
POD_ROOM=demo-pod
```
Gemma/Hermes provider values live in Hermes config, not PodMan `.env`:
```text
provider: gemma4-31b-vllm
model: gemma-4-31B-it
base_url: https://llm.alhinai.dev/v1
api_key: not-needed
```
---
## Verification
Before calling a deployment healthy:
```bash
pnpm verify
pnpm verify:infra
pnpm deploy:doctor:strict
pnpm hermes:watchdog:strict
```
For the public site:
```bash
curl https://podman.live/
curl https://podman.live/health
curl https://podman.live/api/pods
curl https://podman.live/api/presence
curl https://podman.live/api/memory/stats
```
For Hermes/Gemma:
```bash
hermes -z 'Reply with exactly: working' \
--provider gemma4-31b-vllm \
--model gemma-4-31B-it
```
Expected output:
```text
working
```
---
## Troubleshooting
| Symptom | Likely cause | Check |
| ---------------------------------------- | ---------------------------------------------- | ------------------------------------------------------------ |
| Frontend loads but API fails | Backend or Caddy proxy issue | `systemctl status podman-platform-api caddy` |
| `/api/*` returns 502 | API not listening on `8787` | `ss -ltnp`, `curl http://127.0.0.1:8787/health` |
| No screen observations | Vision agent not in LiveKit room | `journalctl -u podman-platform-agent -n 80` |
| Live voice missing | Python conversation agent down | `journalctl -u podman-live-conversation-agent -n 80` |
| Memory empty | MongoDB unavailable or env missing | `curl /api/memory/stats`, backend logs |
| Hermes says tool auto-choice is disabled | vLLM missing tool flags | verify `--enable-auto-tool-choice --tool-call-parser gemma4` |
| `llm.alhinai.dev` returns 502 | Gemma vLLM still loading or tunnel target down | `curl /v1/models`, vLLM logs on Gemma host |
---
## Why It Gets Better
PodMan is not a static alert system. It remembers what actually helped.
- A dismissed false alarm lowers future urgency.
- An accepted real collision raises future urgency for similar work.
- Exact-signature recall catches repeats even without vector search.
- MongoDB outcomes become the policy signal for the next session.
- Hermes and Gemma give the system a tool-using agent when the coordination
problem needs active investigation instead of a passive card.
The goal is simple: fewer interruptions, fewer duplicate branches, and a team
that can move fast without constantly asking what everyone else is doing.
@@ -0,0 +1,8 @@
LIVEKIT_URL=wss://stackauthnov28-kt4gd6fq.livekit.cloud
LIVEKIT_API_KEY=
LIVEKIT_API_SECRET=
GOOGLE_API_KEY=
GEMINI_CONVERSATION_MODEL=gemini-3.1-flash-live-preview
GEMINI_CONVERSATION_VOICE=Aoede
PODMAN_BACKEND_URL=http://127.0.0.1:8787
INTERNAL_AGENT_TOKEN=
+22
View File
@@ -0,0 +1,22 @@
# PodMan Live Conversation Agent
Private 1:1 LiveKit Agent worker for PodMan Live Conversation.
Run locally:
```bash
cd agents/podman-live-conversation
uv sync --extra test
cp .env.example .env.local
uv run agent.py dev
```
Required env:
- `LIVEKIT_URL`
- `LIVEKIT_API_KEY`
- `LIVEKIT_API_SECRET`
- `GOOGLE_API_KEY` or `GEMINI_API_KEY`
- `PODMAN_BACKEND_URL`
- `INTERNAL_AGENT_TOKEN`
+455
View File
@@ -0,0 +1,455 @@
import asyncio
import json
import logging
import os
import subprocess
import time
from typing import Any
from urllib import error, request
from dotenv import load_dotenv
from livekit import agents
from livekit.agents import Agent, AgentServer, AgentSession, RunContext, function_tool
from livekit.plugins import google
load_dotenv(".env.local")
if not os.getenv("GOOGLE_API_KEY") and os.getenv("GEMINI_API_KEY"):
os.environ["GOOGLE_API_KEY"] = os.environ["GEMINI_API_KEY"]
logger = logging.getLogger("podman-live-conversation")
AGENT_NAME = "podman-live-conversation"
MODEL = os.getenv("GEMINI_CONVERSATION_MODEL", "gemini-3.1-flash-live-preview")
VOICE = os.getenv("GEMINI_CONVERSATION_VOICE", "Aoede")
BACKEND_URL = os.getenv("PODMAN_BACKEND_URL", "http://127.0.0.1:8787").rstrip("/")
INTERNAL_AGENT_TOKEN = os.getenv("INTERNAL_AGENT_TOKEN", "")
REPO_SLUG = os.getenv("PODMAN_REPO_SLUG", "karti-ai/podman")
def _resolve_repo_root() -> str:
override = os.getenv("PODMAN_REPO_ROOT")
if override:
return override
here = os.path.dirname(os.path.abspath(__file__))
try:
out = subprocess.run(
["git", "-C", here, "rev-parse", "--show-toplevel"],
capture_output=True,
text=True,
timeout=5,
)
if out.returncode == 0 and out.stdout.strip():
return out.stdout.strip()
except Exception:
pass
return os.path.abspath(os.path.join(here, "..", ".."))
REPO_ROOT = _resolve_repo_root()
INSTRUCTIONS = """You are PodMan, a concise real-time engineering teammate.
You are in a private 1:1 voice conversation with one developer.
Use PodMan tools before making claims about current work, git state, collisions, blockers,
team memory, or recent decisions. Keep spoken answers short. Prefer one useful next step.
For questions about a person's style, goals, past work, collaboration habits, personal context,
or what they know across pods/sessions, call get_user_learning_profile before answering.
If a critical collision event arrives, stop the current turn and state the alert immediately.
To find code, files, symbols, or how something is implemented in the repository, call search_repo.
For git commit history, authorship, recent changes, or which commit introduced something, call
repo_recent_commits or repo_find_commits.
For complex repository, terminal, GitHub, MongoDB, build, install, deploy, or multi-step tasks,
call delegate_to_hermes. Do not run those actions directly. If the user says stop, wait, cancel,
or change of plans while Hermes is running, call abort_active_hermes_job immediately.
Do not reveal raw secrets, API keys, private tokens, or another teammate's private notes."""
def parse_metadata(raw: str | None) -> dict[str, str]:
try:
data = json.loads(raw or "{}")
except json.JSONDecodeError:
return {}
return {str(k): str(v) for k, v in data.items() if v is not None}
def request_json(path: str, *, method: str = "GET", body: dict[str, Any] | None = None) -> Any:
if not INTERNAL_AGENT_TOKEN:
raise RuntimeError("INTERNAL_AGENT_TOKEN is not configured")
data = None if body is None else json.dumps(body).encode("utf-8")
req = request.Request(
f"{BACKEND_URL}{path}",
data=data,
method=method,
headers={
"authorization": f"Bearer {INTERNAL_AGENT_TOKEN}",
"content-type": "application/json",
},
)
try:
with request.urlopen(req, timeout=5) as res:
payload = res.read().decode("utf-8")
return json.loads(payload) if payload else {}
except error.HTTPError as exc:
detail = exc.read().decode("utf-8", "replace")
raise RuntimeError(f"PodMan backend returned {exc.code}: {detail}") from exc
async def _run_git(args: list[str], timeout: float = 15.0) -> tuple[int, str, str]:
"""Run a read-only git command inside the repo checkout and capture its output."""
try:
proc = await asyncio.create_subprocess_exec(
"git",
"-C",
REPO_ROOT,
*args,
stdout=asyncio.subprocess.PIPE,
stderr=asyncio.subprocess.PIPE,
)
out, err = await asyncio.wait_for(proc.communicate(), timeout=timeout)
except asyncio.TimeoutError:
return 124, "", "git command timed out"
except FileNotFoundError:
return 127, "", "git is not available on this host"
return proc.returncode, out.decode("utf-8", "replace"), err.decode("utf-8", "replace")
class PodManLiveAgent(Agent):
def __init__(self, pod_id: str, identity: str, session_id: str, conversation_room: str) -> None:
super().__init__(instructions=INSTRUCTIONS)
self.pod_id = pod_id
self.identity = identity
self.session_id = session_id
self.conversation_room = conversation_room
self.active_hermes_job_id: str | None = None
self.last_spoken_progress_at = 0.0
@function_tool()
async def get_active_pod_context(self, context: RunContext) -> str:
"""Get the current PodMan context for this developer and pod."""
data = await asyncio.to_thread(
request_json,
f"/api/internal/pods/{self.pod_id}/live-context?identity={self.identity}",
)
return json.dumps(data, ensure_ascii=True)[:12000]
@function_tool()
async def get_user_learning_profile(self, context: RunContext) -> str:
"""Get persistent cross-pod, cross-session knowledge about this developer:
collaboration style, goals, known work, recent activity, and Hermes history.
"""
data = await asyncio.to_thread(
request_json,
f"/api/internal/pods/{self.pod_id}/live-context?identity={self.identity}",
)
profile = data.get("userLearningProfile")
if not profile:
return "No persistent user learning profile has been built for this developer yet."
return json.dumps(profile, ensure_ascii=True)[:10000]
@function_tool()
async def record_conversation_note(self, context: RunContext, note: str, kind: str = "summary") -> str:
"""Store a useful decision, outcome, or preference learned during this conversation."""
await asyncio.to_thread(
request_json,
f"/api/internal/pods/{self.pod_id}/live-conversation/{self.session_id}/note",
method="POST",
body={"identity": self.identity, "kind": kind, "note": note},
)
return "Saved to PodMan memory."
@function_tool()
async def get_recent_changes(self, context: RunContext) -> str:
"""Get recent local git and activity signals for this developer."""
data = await asyncio.to_thread(
request_json,
f"/api/internal/pods/{self.pod_id}/live-context?identity={self.identity}",
)
focused = {
"identity": data.get("identity"),
"currentGitState": data.get("currentGitState"),
"memberHistory": data.get("memberHistory"),
"recentCollisions": data.get("recentCollisions"),
}
return json.dumps(focused, ensure_ascii=True)[:8000]
@function_tool()
async def search_team_memory(self, context: RunContext, query: str) -> str:
"""Search current compact team memory for information relevant to a query."""
data = await asyncio.to_thread(
request_json,
f"/api/internal/pods/{self.pod_id}/live-context?identity={self.identity}",
)
haystack = json.dumps(data, ensure_ascii=True)
query_terms = [term.lower() for term in query.split() if len(term) > 2]
if not query_terms:
return haystack[:6000]
snippets = []
lower = haystack.lower()
for term in query_terms[:8]:
idx = lower.find(term)
if idx >= 0:
snippets.append(haystack[max(0, idx - 400) : idx + 1200])
return "\n---\n".join(snippets)[:8000] or haystack[:6000]
@function_tool()
async def search_repo(self, context: RunContext, query: str, max_results: int = 12) -> str:
"""Search the team's code repository (github.com/karti-ai/podman) for code, symbols,
filenames, config, or any text. Use this to find where something is implemented or which
files mention a term before answering questions about the codebase. Searches the live
local checkout of the main branch, so results are always current.
"""
cleaned = " ".join(query.split()).strip()
if not cleaned:
return "Provide a non-empty search query."
limit = max(1, min(int(max_results or 12), 40))
cmd = [
"rg",
"--line-number",
"--no-heading",
"--color",
"never",
"--smart-case",
"--max-count",
"3",
"--max-columns",
"240",
"-g",
"!*.lock",
"-g",
"!pnpm-lock.yaml",
"-g",
"!uv.lock",
"-g",
"!*.min.*",
"--",
cleaned,
REPO_ROOT,
]
try:
proc = await asyncio.create_subprocess_exec(
*cmd,
stdout=asyncio.subprocess.PIPE,
stderr=asyncio.subprocess.PIPE,
)
stdout, stderr = await asyncio.wait_for(proc.communicate(), timeout=15)
except asyncio.TimeoutError:
return "Repo search timed out. Try a more specific query."
except FileNotFoundError:
return "Repo search is unavailable on this host (ripgrep is not installed)."
if proc.returncode not in (0, 1): # rg: 0=match, 1=no match, 2=error
return f"Repo search failed: {stderr.decode('utf-8', 'replace')[:300]}"
prefix = REPO_ROOT + os.sep
lines: list[str] = []
for line in stdout.decode("utf-8", "replace").splitlines():
lines.append(line[len(prefix) :] if line.startswith(prefix) else line)
if len(lines) >= limit:
break
if not lines:
return f'No matches for "{cleaned}" in {REPO_SLUG}.'
body = "\n".join(lines)
return f'Matches for "{cleaned}" in {REPO_SLUG} (path:line):\n{body}'[:7000]
@function_tool()
async def repo_recent_commits(
self, context: RunContext, path: str = "", author: str = "", limit: int = 15
) -> str:
"""Show recent git commit history for github.com/karti-ai/podman: who committed what and when.
Optionally scope to a file or folder (path) or filter by author name/email (author).
Use this for questions about recent changes, authorship, or a specific file's history.
"""
n = max(1, min(int(limit or 15), 50))
args = [
"log",
f"--max-count={n}",
"--no-color",
"--date=short",
"--pretty=format:%h | %an | %ad | %s",
]
if author.strip():
args.append(f"--author={author.strip()}")
if path.strip():
args += ["--", path.strip()]
code, out, err = await _run_git(args)
if code != 0:
return f"Git history lookup failed: {err.strip()[:300] or 'unknown error'}"
out = out.strip()
if not out:
scope = f" for {path.strip()}" if path.strip() else ""
who = f" by {author.strip()}" if author.strip() else ""
return f"No commits found{scope}{who}."
return f"Recent commits in {REPO_SLUG} (hash | author | date | subject):\n{out}"[:7000]
@function_tool()
async def repo_find_commits(
self, context: RunContext, query: str, by: str = "message", limit: int = 15
) -> str:
"""Find commits in github.com/karti-ai/podman. by='message' searches commit messages;
by='code' finds commits that added or removed the query text in the code (pickaxe).
Use by='code' for "which commit introduced X"; use by='message' for "commits about X".
"""
cleaned = " ".join(query.split()).strip()
if not cleaned:
return "Provide a non-empty query."
n = max(1, min(int(limit or 15), 50))
args = [
"log",
f"--max-count={n}",
"--no-color",
"--date=short",
"--pretty=format:%h | %an | %ad | %s",
]
mode = by.strip().lower()
if mode == "code":
args.append(f"-S{cleaned}")
else:
mode = "message"
args += ["-i", f"--grep={cleaned}"]
code, out, err = await _run_git(args)
if code != 0:
return f"Commit search failed: {err.strip()[:300] or 'unknown error'}"
out = out.strip()
if not out:
return f'No commits found matching "{cleaned}" (by {mode}).'
return (
f'Commits in {REPO_SLUG} matching "{cleaned}" (by {mode}) — hash | author | date | subject:\n{out}'[
:7000
]
)
@function_tool()
async def delegate_to_hermes(
self,
context: RunContext,
prompt: str,
context_scope: str = "current_repo",
target_repository: str = "",
risk_level: str = "read_only",
requires_confirmation: bool = False,
success_criteria: list[str] | None = None,
) -> str:
"""Hand off a complex engineering task to Hermes, PodMan's autonomous backend execution engine.
Use this for filesystem, terminal, GitHub, MongoDB, build, install, deploy, test,
or multi-step repository tasks. Do not use this for simple conversational answers.
"""
body = {
"prompt": prompt,
"contextScope": context_scope,
"targetRepository": target_repository or "karti-ai/podman",
"riskLevel": risk_level,
"requiresConfirmation": requires_confirmation,
"successCriteria": success_criteria or ["Hermes completes the requested inspection."],
"podId": self.pod_id,
"identity": self.identity,
"sessionId": self.session_id,
"conversationRoom": self.conversation_room,
}
job = await asyncio.to_thread(request_json, "/api/internal/hermes/jobs", method="POST", body=body)
self.active_hermes_job_id = str(job["id"])
return json.dumps(
{
"status": "accepted",
"job_id": self.active_hermes_job_id,
"spoken_ack": "Hermes is starting that now. I will keep you posted.",
},
ensure_ascii=True,
)
@function_tool()
async def abort_active_hermes_job(self, context: RunContext, reason: str = "User changed plans") -> str:
"""Abort the currently running Hermes job immediately."""
if not self.active_hermes_job_id:
return "No active Hermes job is running."
job = await asyncio.to_thread(
request_json,
f"/api/internal/hermes/jobs/{self.active_hermes_job_id}/abort",
method="POST",
body={"reason": reason},
)
return json.dumps(
{
"status": job.get("status", "aborting"),
"job_id": self.active_hermes_job_id,
"spoken_ack": "Stopped. Hermes is aborting the job before making further changes.",
},
ensure_ascii=True,
)
def should_speak_progress(self, event: dict[str, Any]) -> bool:
event_type = event.get("type")
if event_type in {"completed", "failed", "aborted", "needs_confirmation"}:
return True
if event_type not in {"heartbeat", "step_started", "step_completed"}:
return False
monotonic = time.monotonic()
if monotonic - self.last_spoken_progress_at < 8:
return False
self.last_spoken_progress_at = monotonic
return True
server = AgentServer()
@server.rtc_session(agent_name=AGENT_NAME)
async def entrypoint(ctx: agents.JobContext):
metadata = parse_metadata(getattr(ctx.job, "metadata", None))
pod_id = metadata.get("podId", "demo-pod")
identity = metadata.get("identity", "developer")
session_id = metadata.get("sessionId", "unknown")
session = AgentSession(
llm=google.realtime.RealtimeModel(
model=MODEL,
voice=VOICE,
),
)
agent = PodManLiveAgent(
pod_id=pod_id,
identity=identity,
session_id=session_id,
conversation_room=ctx.room.name,
)
def on_data_received(*args: Any):
payload = args[0] if args else b""
if isinstance(payload, str):
raw = payload
else:
raw = bytes(payload).decode("utf-8", "replace")
try:
msg = json.loads(raw)
except json.JSONDecodeError:
return
msg_type = msg.get("type")
if msg_type == "HERMES_JOB_EVENT":
event = msg.get("event") or {}
summary = str(event.get("message") or "").strip()
if str(event.get("type")) in {"completed", "failed", "aborted"}:
agent.active_hermes_job_id = None
elif msg_type == "LIVE_CONVERSATION_EVENT":
event = msg.get("event") or {}
summary = str(event.get("summary") or "").strip()
else:
return
if not summary:
return
async def interrupt_and_say() -> None:
try:
if msg_type == "LIVE_CONVERSATION_EVENT":
await session.interrupt(force=True)
except Exception as exc:
logger.warning("interrupt failed: %s", exc)
if msg_type == "LIVE_CONVERSATION_EVENT" or agent.should_speak_progress(event):
await session.say(summary, allow_interruptions=True, add_to_chat_ctx=True)
asyncio.create_task(interrupt_and_say())
ctx.room.on("data_received", on_data_received)
await session.start(room=ctx.room, agent=agent)
await ctx.connect()
if __name__ == "__main__":
agents.cli.run_app(server)
@@ -0,0 +1,19 @@
[project]
name = "podman-live-conversation"
version = "0.1.0"
description = "PodMan private LiveKit/Gemini live conversation agent"
requires-python = ">=3.10,<3.14"
dependencies = [
"livekit-agents[google]>=1.6.4,<1.7",
"python-dotenv>=1.0.0",
]
[project.optional-dependencies]
test = ["pytest>=8.0.0"]
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
[tool.hatch.build.targets.wheel]
packages = ["."]
@@ -0,0 +1,26 @@
from agent import PodManLiveAgent, parse_metadata
def test_parse_metadata_accepts_valid_json():
assert parse_metadata('{"podId":"demo-pod","identity":"yahya","sessionId":"s1"}') == {
"podId": "demo-pod",
"identity": "yahya",
"sessionId": "s1",
}
def test_parse_metadata_handles_bad_json():
assert parse_metadata("not json") == {}
def test_hermes_terminal_events_always_speak():
agent = PodManLiveAgent("demo-pod", "yahya", "s1", "room")
assert agent.should_speak_progress({"type": "completed"}) is True
assert agent.should_speak_progress({"type": "failed"}) is True
assert agent.should_speak_progress({"type": "aborted"}) is True
def test_hermes_progress_is_throttled():
agent = PodManLiveAgent("demo-pod", "yahya", "s1", "room")
assert agent.should_speak_progress({"type": "heartbeat"}) is True
assert agent.should_speak_progress({"type": "step_started"}) is False
File diff suppressed because it is too large Load Diff
+1
View File
@@ -17,6 +17,7 @@
"typecheck": "tsc -p tsconfig.json --noEmit"
},
"dependencies": {
"@clerk/express": "^2.1.32",
"@google/genai": "^2.10.0",
"@livekit/rtc-node": "^0.13.29",
"@podman/shared": "workspace:*",
+60
View File
@@ -1,6 +1,11 @@
import type { Room } from '@livekit/rtc-node';
import { Room as LiveKitRoom } from '@livekit/rtc-node';
import { AccessToken } from 'livekit-server-sdk';
import type { Collision, DataMessage, HermesMessage, Intervention } from '@podman/shared';
import { DATA_TOPIC } from '@podman/shared';
import { env } from '../env.js';
import { speak } from '../voice/live.js';
import { notifyCriticalLiveConversations } from '../live-conversation/sessions.js';
const encoder = new TextEncoder();
@@ -37,3 +42,58 @@ export async function publishHermesMessage(
topic: DATA_TOPIC,
});
}
export async function publishHermesIntervention(
room: Room,
collision: Collision,
intervention: Intervention,
voiceLine?: string,
): Promise<void> {
const data: DataMessage = { type: 'COLLISION', collision, intervention };
await room.localParticipant?.publishData(encoder.encode(JSON.stringify(data)), {
reliable: true,
topic: DATA_TOPIC,
});
await publishHermesMessage(room, collision, intervention);
void notifyCriticalLiveConversations(collision, intervention, voiceLine).catch((err) =>
console.warn(`[live-conversation] critical notify failed: ${(err as Error).message}`),
);
if (voiceLine)
await speak(room, voiceLine, {
priority: collision.severity === 'critical' ? 'critical' : 'normal',
});
}
async function hermesToken(roomName: string): Promise<string> {
const at = new AccessToken(env.LIVEKIT_API_KEY, env.LIVEKIT_API_SECRET, {
identity: `podman-hermes-${Date.now()}`,
name: 'PodMan Hermes',
ttl: '10m',
});
at.addGrant({
roomJoin: true,
room: roomName,
canPublish: true,
canSubscribe: true,
canPublishData: true,
});
return at.toJwt();
}
export async function notifyHermesInterventionInRoom(
roomName: string,
collision: Collision,
intervention: Intervention,
voiceLine?: string,
): Promise<void> {
const room = new LiveKitRoom();
try {
await room.connect(env.LIVEKIT_URL, await hermesToken(roomName), {
autoSubscribe: false,
dynacast: false,
});
await publishHermesIntervention(room, collision, intervention, voiceLine);
} finally {
await room.disconnect().catch(() => {});
}
}
+244
View File
@@ -0,0 +1,244 @@
import type {
Collision,
EngineerContext,
MemberWorkHistory,
MemberWorkHistoryFile,
MemberWorkHistoryRoi,
} from '@podman/shared';
import { getDb } from '../memory/db.js';
import { parseGitStatusPath } from '../graph/live.js';
interface EngineerStateDoc {
_id: string;
podId: string;
name: string;
changedFiles?: string[];
branch?: string | null;
recentCommit?: string | null;
gitUpdatedAt?: Date | string;
updatedAt?: Date | string;
}
interface FileAccumulator {
file: string;
observations: number;
gitChanges: number;
firstSeenAt: number;
lastSeenAt: number;
confidenceSum: number;
confidenceCount: number;
activities: Set<string>;
current: boolean;
}
function toIso(ms: number): string {
return new Date(ms).toISOString();
}
function dateMs(value: string | Date | undefined): number {
if (value instanceof Date) return value.getTime();
if (value) {
const parsed = Date.parse(value);
if (Number.isFinite(parsed)) return parsed;
}
return 0;
}
function clean(value: string | undefined): string {
return value?.trim() ?? '';
}
function sameMember(a: string | undefined, b: string): boolean {
return clean(a).toLowerCase() === b.trim().toLowerCase();
}
function addFile(files: Map<string, FileAccumulator>, file: string, at: number): FileAccumulator {
const existing = files.get(file);
if (existing) {
if (at > 0) {
existing.firstSeenAt = Math.min(existing.firstSeenAt || at, at);
existing.lastSeenAt = Math.max(existing.lastSeenAt, at);
}
return existing;
}
const acc: FileAccumulator = {
file,
observations: 0,
gitChanges: 0,
firstSeenAt: at,
lastSeenAt: at,
confidenceSum: 0,
confidenceCount: 0,
activities: new Set<string>(),
current: false,
};
files.set(file, acc);
return acc;
}
export async function getMemberWorkHistory(
podId: string,
member: string,
options: { hours?: number; limit?: number } = {},
): Promise<MemberWorkHistory> {
const db = await getDb();
const windowHours = Math.min(Math.max(options.hours ?? 24, 1), 168);
const limit = Math.min(Math.max(options.limit ?? 80, 10), 200);
const since = new Date(Date.now() - windowHours * 60 * 60 * 1000).toISOString();
const [observations, gitState, collisions, interventions] = await Promise.all([
db
.collection<EngineerContext>('observations')
.find({ podId, observedAt: { $gte: since } }, { projection: { _id: 0 } })
.sort({ observedAt: -1 })
.limit(500)
.toArray(),
db.collection<EngineerStateDoc>('engineer_states').findOne({
podId,
name: { $regex: `^${member.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')}$`, $options: 'i' },
}),
db
.collection<Collision>('collisions')
.find({ podId, detectedAt: { $gte: since } }, { projection: { _id: 0 } })
.sort({ detectedAt: -1 })
.limit(200)
.toArray(),
db
.collection<{ collisionId: string }>('interventions')
.find({ podId }, { projection: { collisionId: 1, _id: 0 } })
.toArray(),
]);
const files = new Map<string, FileAccumulator>();
const timeline: MemberWorkHistory['timeline'] = [];
const memberObservations = observations.filter((doc) => sameMember(doc.engineerId, member));
for (const doc of memberObservations) {
const file = clean(doc.currentFile);
if (!file) continue;
const at = dateMs(doc.observedAt);
const acc = addFile(files, file, at);
acc.observations += 1;
acc.current ||= timeline.length === 0;
if (typeof doc.confidence === 'number') {
acc.confidenceSum += doc.confidence;
acc.confidenceCount += 1;
}
const activity = clean(doc.activity);
if (activity) acc.activities.add(activity);
timeline.push({
id: `vision:${doc.engineerId}:${doc.observedAt}:${file}`,
at: doc.observedAt,
source: 'vision',
file,
title: activity || `Worked in ${file}`,
detail: clean(doc.currentSymbol) ? `symbol ${doc.currentSymbol}` : undefined,
confidence: doc.confidence,
});
}
const gitAt = dateMs(gitState?.gitUpdatedAt ?? gitState?.updatedAt);
for (const raw of gitState?.changedFiles ?? []) {
const file = parseGitStatusPath(raw);
if (!file) continue;
const acc = addFile(files, file, gitAt || Date.now());
acc.gitChanges += 1;
acc.current = true;
timeline.push({
id: `git:${gitState?._id}:${gitAt}:${file}`,
at: toIso(gitAt || Date.now()),
source: 'git',
file,
title: `Local change in ${file}`,
detail: [gitState?.branch ? `branch ${gitState.branch}` : undefined, gitState?.recentCommit]
.filter(Boolean)
.join(' · '),
});
}
const fileRows: MemberWorkHistoryFile[] = [...files.values()]
.sort((a, b) => b.lastSeenAt - a.lastSeenAt || b.observations - a.observations)
.slice(0, 12)
.map((file) => ({
file: file.file,
observations: file.observations,
gitChanges: file.gitChanges,
firstSeenAt: toIso(file.firstSeenAt || file.lastSeenAt || Date.now()),
lastSeenAt: toIso(file.lastSeenAt || file.firstSeenAt || Date.now()),
confidenceAvg: file.confidenceCount
? Math.round((file.confidenceSum / file.confidenceCount) * 100) / 100
: null,
activities: [...file.activities].slice(0, 3),
current: file.current,
}));
timeline.sort((a, b) => Date.parse(b.at) - Date.parse(a.at));
const interventionIds = new Set(interventions.map((i) => i.collisionId));
const roi = computeRoi(member, collisions, interventionIds, gitState?.changedFiles?.length ?? 0);
return {
podId,
member,
generatedAt: new Date().toISOString(),
windowHours,
totals: {
files: fileRows.length,
observations: memberObservations.length,
gitChanges: gitState?.changedFiles?.length ?? 0,
},
files: fileRows,
timeline: timeline.slice(0, limit),
roi,
};
}
function computeRoi(
member: string,
collisions: Collision[],
interventionCollisionIds: Set<string>,
changedFileCount: number,
): MemberWorkHistoryRoi {
const involved = (c: Collision) =>
c.engineers?.some((e) => sameMember(e, member)) ||
sameMember(c.researcher, member) ||
sameMember(c.editor, member);
const eligible = collisions.filter(
(c) =>
involved(c) &&
interventionCollisionIds.has(c.id) &&
(c.gitOverlap === true || c.severity === 'critical'),
);
const weightOf = (c: Collision): { label: string; minutes: number } => {
if (c.overlapKind === 'research') return { label: 'research overlap', minutes: 10 };
if (c.severity === 'critical') return { label: 'critical same-file', minutes: 45 };
if (c.severity === 'warn') return { label: 'warn same-file', minutes: 20 };
return { label: 'info same-file', minutes: 10 };
};
let savedMinutes = 0;
const groups = new Map<string, { count: number; minutesEach: number }>();
for (const c of eligible) {
const { label, minutes } = weightOf(c);
savedMinutes += minutes / Math.max(1, c.engineers?.length ?? 1);
const g = groups.get(label) ?? { count: 0, minutesEach: minutes };
g.count += 1;
groups.set(label, g);
}
const filesDeconflicted = new Set(eligible.map((c) => c.file)).size;
return {
savedMinutes: Math.round(savedMinutes),
clashesCaught: eligible.length,
filesDeconflicted,
conflictFreeFiles: Math.max(0, changedFileCount - filesDeconflicted),
totalFiles: changedFileCount,
breakdown: [...groups.entries()].map(([label, g]) => ({
label,
count: g.count,
minutesEach: g.minutesEach,
})),
};
}
+197
View File
@@ -0,0 +1,197 @@
import type {
Collision,
EngineerContext,
Intervention,
InterventionOutcome,
PodActivityEvent,
} from '@podman/shared';
import { getDb } from '../memory/db.js';
interface EngineerStateDoc {
_id: string;
podId: string;
name: string;
changedFiles?: string[];
diffStat?: string | null;
recentCommit?: string | null;
branch?: string | null;
gitUpdatedAt?: Date | string;
updatedAt?: Date | string;
}
function toIso(value: Date | string | undefined): string {
if (value instanceof Date) return value.toISOString();
if (value) return new Date(value).toISOString();
return new Date(0).toISOString();
}
function clean(value: string | undefined): string | undefined {
const trimmed = value?.trim();
return trimmed || undefined;
}
function shortFiles(files: string[] | undefined): string {
if (!files?.length) return 'clean working tree';
const sample = files.slice(0, 3).join(', ');
return files.length > 3 ? `${sample}, +${files.length - 3} more` : sample;
}
function observationEvent(doc: EngineerContext): PodActivityEvent {
const file = clean(doc.currentFile);
const symbol = clean(doc.currentSymbol);
return {
id: `observation:${doc.engineerId}:${doc.observedAt}`,
podId: doc.podId,
kind: 'observation',
source: 'vision',
actor: doc.engineerId,
actors: [doc.engineerId],
file,
imageUrl: doc.screenshotDataUrl,
title: file ? `Working in ${file}` : 'Screen context updated',
detail: [
symbol ? `symbol ${symbol}` : undefined,
clean(doc.activity),
doc.hasUnpushedChanges ? 'unpushed changes visible' : undefined,
`confidence ${Math.round(doc.confidence * 100)}%`,
]
.filter(Boolean)
.join(' · '),
severity: doc.hasUnpushedChanges ? 'warn' : 'info',
at: doc.observedAt,
};
}
function gitEvent(doc: EngineerStateDoc): PodActivityEvent {
const changedFiles = doc.changedFiles ?? [];
return {
id: `git:${doc._id}:${toIso(doc.gitUpdatedAt ?? doc.updatedAt)}`,
podId: doc.podId,
kind: 'git',
source: 'git',
actor: doc.name,
actors: [doc.name],
title: changedFiles.length ? `${changedFiles.length} local file changes` : 'Git state is clean',
detail: [
doc.branch ? `branch ${doc.branch}` : undefined,
shortFiles(changedFiles),
doc.recentCommit ? `head ${doc.recentCommit}` : undefined,
]
.filter(Boolean)
.join(' · '),
severity: changedFiles.length ? 'warn' : 'info',
at: toIso(doc.gitUpdatedAt ?? doc.updatedAt),
};
}
function collisionEvent(doc: Collision): PodActivityEvent {
return {
id: `collision:${doc.id}`,
podId: doc.podId,
kind: 'collision',
source: 'memory',
actor: doc.engineers[0],
actors: doc.engineers,
file: doc.file,
title: `${doc.engineers.join(' + ')} conflict on ${doc.file}`,
detail: [
doc.symbol ? `symbol ${doc.symbol}` : undefined,
doc.githubState?.unpushed ? 'unpushed local changes involved' : undefined,
doc.githubState?.openPrs?.length
? `open PRs ${doc.githubState.openPrs.join(', ')}`
: undefined,
]
.filter(Boolean)
.join(' · '),
severity: doc.severity,
at: doc.detectedAt,
};
}
function interventionEvent(doc: Intervention): PodActivityEvent {
return {
id: `intervention:${doc.id}`,
podId: doc.podId,
kind: 'intervention',
source: 'hermes',
title: `Hermes ${doc.status} ${doc.suggestedAction.kind.replaceAll('_', ' ')}`,
detail: doc.message,
severity: doc.status === 'accepted' ? 'success' : doc.status === 'dismissed' ? 'info' : 'warn',
at: doc.createdAt,
};
}
function outcomeEvent(doc: InterventionOutcome): PodActivityEvent {
return {
id: `outcome:${doc.interventionId}:${doc.recordedAt}`,
podId: doc.podId,
kind: 'outcome',
source: 'policy',
title: doc.accepted ? 'Intervention accepted' : 'Intervention dismissed',
detail: doc.wasRealCollision ? 'confirmed real collision' : 'marked as false positive',
severity: doc.accepted ? 'success' : 'info',
at: doc.recordedAt,
};
}
export async function listPodActivity(podId: string, limit = 80): Promise<PodActivityEvent[]> {
const db = await getDb();
const [observations, gitStates, collisions, interventions, outcomes] = await Promise.all([
db
.collection<EngineerContext>('observations')
.find({ podId }, { projection: { _id: 0 } })
.sort({ observedAt: -1 })
.limit(limit)
.toArray(),
db
.collection<EngineerStateDoc>('engineer_states')
.find(
{ podId },
{
projection: {
_id: 1,
podId: 1,
name: 1,
changedFiles: 1,
diffStat: 1,
recentCommit: 1,
branch: 1,
gitUpdatedAt: 1,
updatedAt: 1,
},
},
)
.sort({ gitUpdatedAt: -1 })
.limit(limit)
.toArray(),
db
.collection<Collision>('collisions')
.find({ podId }, { projection: { _id: 0 } })
.sort({ detectedAt: -1 })
.limit(limit)
.toArray(),
db
.collection<Intervention>('interventions')
.find({ podId }, { projection: { _id: 0 } })
.sort({ createdAt: -1 })
.limit(limit)
.toArray(),
db
.collection<InterventionOutcome>('outcomes')
.find({ podId }, { projection: { _id: 0 } })
.sort({ recordedAt: -1 })
.limit(limit)
.toArray(),
]);
return [
...observations.map(observationEvent),
...gitStates.map(gitEvent),
...collisions.map(collisionEvent),
...interventions.map(interventionEvent),
...outcomes.map(outcomeEvent),
]
.filter((event) => event.at !== new Date(0).toISOString())
.sort((a, b) => Date.parse(b.at) - Date.parse(a.at))
.slice(0, limit);
}
+117 -14
View File
@@ -6,6 +6,7 @@ import {
VideoStream,
VideoBufferType,
dispose,
type VideoFrameEvent,
type RemoteTrack,
type RemoteTrackPublication,
type RemoteParticipant,
@@ -19,6 +20,8 @@ import { initMemory } from './memory/db.js';
const POD_ROOM = process.env.POD_ROOM ?? 'demo-pod';
const HERMES_IDENTITY = 'podman-hermes';
const SAMPLE_INTERVAL_MS = 1000; // ~1 fps to the vision model
const SCREEN_THUMBNAIL_WIDTH = 360;
const SHUTDOWN_GRACE_MS = 5000;
async function agentToken(room: string): Promise<string> {
const at = new AccessToken(env.LIVEKIT_API_KEY, env.LIVEKIT_API_SECRET, {
@@ -30,6 +33,20 @@ async function agentToken(room: string): Promise<string> {
return at.toJwt();
}
async function withTimeout<T>(promise: Promise<T>, ms: number): Promise<T | null> {
let timer: NodeJS.Timeout | undefined;
try {
return await Promise.race([
promise,
new Promise<null>((resolve) => {
timer = setTimeout(() => resolve(null), ms);
}),
]);
} finally {
if (timer) clearTimeout(timer);
}
}
async function main() {
// MongoDB is mandatory. Verify the connection before joining the room so bad
// creds / unreachable Atlas fail loudly at boot, not silently mid-demo.
@@ -45,6 +62,57 @@ async function main() {
console.log(`[agent] ${HERMES_IDENTITY} joined room ${POD_ROOM}`);
const lastSent = new Map<string, number>();
const inFlight = new Set<string>();
const activeStreams = new Map<string, ReadableStreamDefaultReader<VideoFrameEvent>>();
const streamKey = (
track: RemoteTrack,
pub: RemoteTrackPublication,
participant: RemoteParticipant,
) => `${participant.identity}:${pub.sid ?? track.sid ?? 'screen'}`;
const stopStream = async (key: string) => {
const reader = activeStreams.get(key);
if (!reader) return;
activeStreams.delete(key);
await reader.cancel().catch(() => {});
try {
reader.releaseLock();
} catch {
/* already released */
}
};
const processFrame = async (engineerId: string, event: VideoFrameEvent) => {
const now = Date.now();
if (now - (lastSent.get(engineerId) ?? 0) < SAMPLE_INTERVAL_MS) return;
if (inFlight.has(engineerId)) return;
lastSent.set(engineerId, now);
inFlight.add(engineerId);
try {
const rgba = event.frame.convert(VideoBufferType.RGBA);
const pixels = Buffer.from(rgba.data);
const raw = { width: rgba.width, height: rgba.height, channels: 4 } as const;
const [jpeg, thumbnail] = await Promise.all([
sharp(pixels, { raw })
.resize({ width: 1280, withoutEnlargement: true })
.jpeg({ quality: 70 })
.toBuffer(),
sharp(pixels, { raw })
.resize({ width: SCREEN_THUMBNAIL_WIDTH, withoutEnlargement: true })
.jpeg({ quality: 42 })
.toBuffer(),
]);
await podman.onScreenFrame(
engineerId,
jpeg,
`data:image/jpeg;base64,${thumbnail.toString('base64')}`,
);
} finally {
inFlight.delete(engineerId);
}
};
room.on(
RoomEvent.TrackSubscribed,
@@ -52,30 +120,65 @@ async function main() {
if (track.kind !== TrackKind.KIND_VIDEO || pub.source !== TrackSource.SOURCE_SCREENSHARE)
return;
const id = participant.identity;
const key = streamKey(track, pub, participant);
const stream = new VideoStream(track);
void stopStream(key);
const reader = stream.getReader();
activeStreams.set(key, reader);
void (async () => {
for await (const event of stream) {
const now = Date.now();
if (now - (lastSent.get(id) ?? 0) < SAMPLE_INTERVAL_MS) continue; // THROTTLE
lastSent.set(id, now);
const rgba = event.frame.convert(VideoBufferType.RGBA);
const jpeg = await sharp(Buffer.from(rgba.data), {
raw: { width: rgba.width, height: rgba.height, channels: 4 },
})
.resize({ width: 1280, withoutEnlargement: true })
.jpeg({ quality: 70 })
.toBuffer();
await podman.onScreenFrame(id, jpeg);
try {
while (activeStreams.get(key) === reader) {
const { done, value } = await reader.read();
if (done) break;
void processFrame(id, value).catch((err) =>
console.error(`[agent] frame sample failed for ${id}: ${(err as Error).message}`),
);
}
} catch (err) {
console.error(`[agent] screen stream failed for ${id}: ${(err as Error).message}`);
} finally {
if (activeStreams.get(key) === reader) activeStreams.delete(key);
await reader.cancel().catch(() => {});
try {
reader.releaseLock();
} catch {
/* already released */
}
}
})();
},
);
room.on(
RoomEvent.TrackUnsubscribed,
(track: RemoteTrack, pub: RemoteTrackPublication, participant: RemoteParticipant) => {
void stopStream(streamKey(track, pub, participant));
},
);
room.on(RoomEvent.ParticipantDisconnected, (participant: RemoteParticipant) => {
for (const key of [...activeStreams.keys()]) {
if (key.startsWith(`${participant.identity}:`)) void stopStream(key);
}
});
let shuttingDown = false;
const shutdown = async () => {
await room.disconnect();
await dispose();
if (shuttingDown) return;
shuttingDown = true;
await withTimeout(
Promise.all([...activeStreams.keys()].map(stopStream)).then(() => room.disconnect()),
SHUTDOWN_GRACE_MS,
);
await withTimeout(dispose(), SHUTDOWN_GRACE_MS);
process.exit(0);
};
room.on(RoomEvent.Disconnected, () => {
if (!shuttingDown) {
console.error('[agent] LiveKit disconnected; exiting so systemd restarts the worker');
process.exit(1);
}
});
process.on('SIGINT', shutdown);
process.on('SIGTERM', shutdown);
}
+172 -27
View File
@@ -1,24 +1,70 @@
import { RoomEvent, type Room } from '@livekit/rtc-node';
import type { EngineerContext, Collision, Intervention, DataMessage } from '@podman/shared';
import { DATA_TOPIC } from '@podman/shared';
import { analyzeFrame } from '../vision/gemini.js';
import { detectCollisions } from '../collision/detector.js';
import { detectResearchOverlaps } from '../collision/research.js';
import { getGithubState } from '../github/client.js';
import {
recordObservation,
recordCollision,
recordIntervention,
recordSuppression,
updateInterventionStatus,
} from '../memory/store.js';
import { getGitStates } from '../memory/db.js';
import { getGitStates, type GitState } from '../memory/db.js';
import { recallSimilar } from '../memory/vectors.js';
import { shouldIntervene, preferredAction } from '../memory/policy.js';
import { speak } from '../voice/live.js';
import { publishHermesMessage } from '../action/hermes.js';
import { publishHermesIntervention } from '../action/hermes.js';
/** Strip a git-status prefix ("M ", "?? ") and reduce a path to its lowercased
* basename — matches comparableFile() in memory/store.ts so keys line up. */
function comparableBasename(raw?: string): string {
return (
(raw ?? '')
.trim()
.replace(/^(\?\?|[MADRCU!]{1,2})\s+/, '')
.split(/[\\/]/)
.pop()
?.toLowerCase() ?? ''
);
}
/** Canonicalize an engineer name for case/whitespace-insensitive matching, so
* "Karti" and "karti" resolve to the same engineer's git state. */
function canonicalName(raw?: string): string {
return (raw ?? '').trim().toLowerCase();
}
/** How long a still-present conflict stays muted after we voice it, before it
* re-alerts. Edge-trigger alone permanently muted files that stay dirty for
* the whole session (e.g. README everyone tests on). This bounds that: voice
* once, then re-alert at most every interval while it persists. */
const CONFLICT_REALERT_MS = Number(process.env.CONFLICT_REALERT_MS ?? '20000');
/** Git ground truth: do ALL involved engineers currently have the collided file
* in their changedFiles? Computed at detection time while git state is fresh. */
function engineersOverlapOnFile(collision: Collision, gitStates: Map<string, GitState>): boolean {
const target = comparableBasename(collision.file);
if (!target || collision.engineers.length < 2) return false;
const byCanon = new Map<string, string[]>();
for (const [name, st] of gitStates) byCanon.set(canonicalName(name), st.changedFiles);
return collision.engineers.every((e) =>
(byCanon.get(canonicalName(e)) ?? []).some((f) => comparableBasename(f) === target),
);
}
export class PodMan {
private contexts = new Map<string, EngineerContext>();
private encoder = new TextEncoder();
/**
* Conflicts we have already voiced, keyed by file + engineer pair (see
* conflictKey), mapped to the time we last voiced them. Edge-triggered:
* speak once when a conflict appears. Re-armed (deleted here) by
* onScreenFrame as soon as a detection cycle no longer sees it, so a
* resolved-then-recurring conflict alerts again. Additionally, a conflict
* that *persists* re-alerts every CONFLICT_REALERT_MS so a perpetually-dirty
* file (README) doesn't go silent forever after the first alert.
*/
private activeConflicts = new Map<string, number>();
constructor(
private room: Room,
@@ -35,15 +81,20 @@ export class PodMan {
if (c)
c.hasUnpushedChanges = msg.report.unpushedCount > 0 || msg.report.dirtyFiles.length > 0;
}
if (msg.type === 'ACK') void updateInterventionStatus(msg.interventionId, msg.status);
if (msg.type === 'ACK') {
void updateInterventionStatus(msg.interventionId, msg.status).catch((err) =>
console.error(`[memory] intervention ack failed: ${(err as Error).message}`),
);
}
} catch {
/* ignore malformed */
}
});
}
async onScreenFrame(engineerId: string, jpeg: Buffer): Promise<void> {
async onScreenFrame(engineerId: string, jpeg: Buffer, screenshotDataUrl?: string): Promise<void> {
const ctx = await analyzeFrame(engineerId, this.podId, jpeg);
if (screenshotDataUrl) ctx.screenshotDataUrl = screenshotDataUrl;
this.contexts.set(engineerId, ctx);
await recordObservation(ctx);
@@ -56,26 +107,123 @@ export class PodMan {
}
const github = await getGithubState(); // cached
const collisions = detectCollisions([...this.contexts.values()], github);
const contexts = [...this.contexts.values()];
const fileCollisions = detectCollisions(contexts, github, gitStates);
const researchCollisions = await detectResearchOverlaps(contexts, gitStates);
const collisions = [...fileCollisions, ...researchCollisions];
// Re-arm: any conflict we previously voiced that is no longer present has
// resolved, so allow it to alert again if it recurs.
const current = new Set(collisions.map((c) => this.conflictKey(c)));
for (const key of this.activeConflicts.keys()) {
if (!current.has(key)) this.activeConflicts.delete(key);
}
// Capture git ground-truth overlap now, while engineer_states are fresh, so
// the outcome-time verifier never depends on a stale sidecar or a late click.
for (const collision of collisions) {
if (collision.overlapKind !== 'research') {
collision.gitOverlap = engineersOverlapOnFile(collision, gitStates);
}
}
for (const collision of collisions) await this.handle(collision);
}
private async handle(collision: Collision): Promise<void> {
const prior = await recallSimilar(collision); // Loop A: exact/vector recall raises confidence
if (prior) collision.severity = 'critical';
if (!shouldIntervene(collision, prior)) return; // Loop B: policy gate
/**
* Stable identity for a conflict, independent of the Date.now() baked into
* collision.id. Mirrors comparableFile() in memory/store.ts so keys line up:
* strip any git-status prefix ("M ", "?? ") and reduce to a lowercased
* basename.
*/
private conflictKey(collision: Collision): string {
const who = [...collision.engineers].map(canonicalName).sort().join('+');
return `${collision.overlapKind ?? 'file'}:${comparableBasename(collision.file)}:${who}`;
}
private async handle(collision: Collision): Promise<void> {
const key = this.conflictKey(collision);
const lastAlerted = this.activeConflicts.get(key);
// Edge-trigger + time-based re-arm: stay quiet right after voicing, but
// re-alert a persistent conflict once CONFLICT_REALERT_MS has elapsed.
if (lastAlerted !== undefined && Date.now() - lastAlerted < CONFLICT_REALERT_MS) return;
const prior = await recallSimilar(collision); // Loop A: exact/vector recall raises confidence
// Only escalate to critical (which triggers the spoken alert) when the
// recalled prior was an *accepted real* collision. Blanket-escalating every
// recall — including dismissed/false-positive priors — masked the learned
// routing in preferredAction and made recalled noise scream "CRITICAL".
// (RSI Step 2 — continual-learning/policy.md:62-63, plan.md:66)
if (prior?.priorOutcome?.accepted && prior?.priorOutcome?.wasRealCollision) {
collision.severity = 'critical';
}
if (!shouldIntervene(collision, prior)) {
// Feature A — make the negative-feedback loop VISIBLE. If we stayed quiet
// *specifically* because this signature was DISMISSED before, record a
// durable suppressed-repeat event (timestamped now, at the repeat) so the
// activity stream shows the learning instead of nothing.
if (prior?.priorOutcome && !prior.priorOutcome.accepted) {
// Mark handled first — like the alert path below — so we record ONE
// suppressed-repeat per recurrence, not once per frame; it re-arms via
// the resolution sweep in onScreenFrame. Awaited like recordCollision so
// the durable learning proof is reliably written.
this.activeConflicts.set(key, Date.now());
await recordSuppression(
collision,
prior.priorOutcome.interventionId,
prior.priorOutcome.recordedAt,
);
}
return; // Loop B: policy gate
}
this.activeConflicts.set(key, Date.now()); // claim + timestamp; re-armed on resolution or after CONFLICT_REALERT_MS
await recordCollision(collision);
const action = preferredAction(collision, prior);
const names = collision.engineers.join(' and ');
const shortFile = collision.file.split('/').pop() ?? collision.file;
const isResearchOverlap = collision.overlapKind === 'research';
if (isResearchOverlap) {
const researcher = collision.researcher ?? collision.engineers[1] ?? 'A teammate';
const editor = collision.editor ?? collision.engineers[0] ?? 'a teammate';
const topic = collision.researchTopic ?? 'the same area';
const source = collision.researchSource ? ` (${collision.researchSource})` : '';
const message = `🤝 ${researcher} is researching ${topic}${source} while ${editor} edits ${shortFile} — sync up before duplicating effort.`;
const voiceLine = `${researcher} is researching ${topic} while ${editor} works on ${shortFile}. Worth a quick sync.`;
const intervention: Intervention = {
id: `int_${Date.now()}`,
collisionId: collision.id,
podId: this.podId,
kind: 'card',
message,
suggestedAction: {
kind: 'ping_teammate',
params: {
file: collision.file,
summary: message,
engineers: collision.engineers,
researchTopic: collision.researchTopic,
researchSource: collision.researchSource,
},
},
status: 'pending',
createdAt: new Date().toISOString(),
};
await recordIntervention(intervention);
await publishHermesIntervention(this.room, collision, intervention, voiceLine);
return;
}
const names = collision.engineers.join(' + ');
// Terse, demo-centered alert — short and direct, not chatty AI prose.
const message =
`${names} are both editing ${collision.file}` +
(collision.githubState?.unpushed ? ' and one has unpushed changes.' : '.') +
(prior?.priorOutcome?.accepted
? ` I've seen this conflict pattern before; last time the team accepted the ${prior.priorIntervention?.suggestedAction.kind.replaceAll('_', ' ') ?? 'suggested'} action.`
: prior
? ` I've seen this conflict pattern before.`
: '');
`Conflict: ${names} both on ${shortFile}` +
(collision.githubState?.unpushed ? ' (unpushed).' : '.') +
(prior ? ' Seen before.' : '');
// Spoken line stays short, but uses natural phrasing for Gemini TTS prosody.
const voiceLine = `${names} are both editing ${shortFile}. Please sync before pushing.`;
const intervention: Intervention = {
id: `int_${Date.now()}`,
@@ -96,12 +244,9 @@ export class PodMan {
};
await recordIntervention(intervention);
const data: DataMessage = { type: 'COLLISION', collision, intervention };
await this.room.localParticipant?.publishData(this.encoder.encode(JSON.stringify(data)), {
reliable: true,
topic: DATA_TOPIC,
});
await publishHermesMessage(this.room, collision, intervention);
if (collision.severity === 'critical') await speak(this.room, message);
// Voice every intervention, not just critical escalations. Priority is set
// by severity inside publishHermesIntervention: critical jumps the queue,
// the rest play sequentially so concurrent alerts don't garble each other.
await publishHermesIntervention(this.room, collision, intervention, voiceLine);
}
}
+30
View File
@@ -0,0 +1,30 @@
import { clerkMiddleware, getAuth } from '@clerk/express';
import type { Request, RequestHandler, Response } from 'express';
import { env } from './env.js';
export interface RequestUserContext {
clerkUserId: string;
}
export const clerkAuthMiddleware: RequestHandler = env.CLERK_SECRET_KEY
? clerkMiddleware()
: (_req, _res, next) => next();
export function requestUser(req: Request): RequestUserContext | null {
if (!env.CLERK_SECRET_KEY) return null;
try {
const auth = getAuth(req);
return auth.isAuthenticated && auth.userId ? { clerkUserId: auth.userId } : null;
} catch {
return null;
}
}
export function requireRequestUser(req: Request, res: Response): RequestUserContext | null {
const user = requestUser(req);
if (!user) {
res.status(401).json({ error: 'sign in required' });
return null;
}
return user;
}
+81 -15
View File
@@ -1,34 +1,100 @@
import type { EngineerContext, Collision, GithubStateSnapshot } from '@podman/shared';
import type { GitState } from '../memory/db.js';
function normalize(path?: string): string | undefined {
if (!path) return undefined;
return path.replace(/^\.?\/?(src\/)?/, 'src/').toLowerCase();
/**
* Collapse any path-ish string to a comparable file key.
*
* Vision reads paths at inconsistent depths ("agent.ts" vs
* "backend/src/agent.ts"), and git status lines carry a status prefix
* ("M README.md", "?? test.txt"). Reduce both to a lowercased basename so the
* same file matches regardless of how it was observed. Basename matching can
* over-group two same-named files in different dirs, but for live coordination
* that bias toward firing is the right trade.
*/
function fileKey(raw?: string): string | undefined {
if (!raw) return undefined;
const stripped = raw.trim().replace(/^(\?\?|[MADRCU!]{1,2})\s+/, ''); // drop git status prefix
const base = stripped.split(/[\\/]/).pop()?.trim();
if (!base) return undefined;
return base.toLowerCase();
}
/**
* Canonicalize an engineer identity for case/whitespace-insensitive matching.
* Vision reports a display name ("Karti") while the git watcher reports a handle
* ("karti"); without this the same human collides with themselves.
*/
function canonicalName(raw: string): string {
return raw.trim().toLowerCase();
}
interface Touch {
engineerId: string;
unpushed: boolean;
display: string; // original path/name to show in the card
}
/**
* Detect same-file collisions from two fused signals:
* 1. Vision — what each engineer currently has on screen.
* 2. Git ground truth — each engineer's dirty/unpushed `changedFiles`.
*
* Git overlap is deterministic and does not require both engineers to have the
* file on screen at the same instant, so it is the reliable demo path.
*/
export function detectCollisions(
contexts: EngineerContext[],
github: GithubStateSnapshot,
gitStates?: Map<string, GitState>,
): Collision[] {
const byFile = new Map<string, EngineerContext[]>();
const byFile = new Map<string, Touch[]>();
const add = (key: string | undefined, touch: Touch): void => {
if (!key) return;
(byFile.get(key) ?? byFile.set(key, []).get(key)!).push(touch);
};
// Signal 1: live vision context.
for (const c of contexts) {
const f = normalize(c.currentFile);
if (!f) continue;
(byFile.get(f) ?? byFile.set(f, []).get(f)!).push(c);
add(fileKey(c.currentFile), {
engineerId: c.engineerId,
unpushed: c.hasUnpushedChanges === true,
display: c.currentFile ?? '',
});
}
// Signal 2: git ground truth (a dirty changed file is unpushed by definition).
if (gitStates) {
for (const [engineerId, git] of gitStates) {
for (const changed of git.changedFiles) {
add(fileKey(changed), { engineerId, unpushed: true, display: changed });
}
}
}
const out: Collision[] = [];
for (const [file, group] of byFile) {
const engineers = [...new Set(group.map((g) => g.engineerId))];
if (engineers.length < 2) continue;
for (const [, touches] of byFile) {
// Distinct PEOPLE, case/whitespace-insensitive — one display name per person.
// Vision "Karti" and git "karti" are the same human, not a collision.
const byPerson = new Map<string, string>(); // canonical id -> display name
for (const t of touches) {
const id = canonicalName(t.engineerId);
if (id && !byPerson.has(id)) byPerson.set(id, t.engineerId);
}
if (byPerson.size < 2) continue; // need two distinct people on one file
const engineers = [...byPerson.values()];
const anyUnpushed = group.some((g) => g.hasUnpushedChanges) || github.unpushed === true;
const anyUnpushed = touches.some((t) => t.unpushed) || github.unpushed === true;
if (!anyUnpushed) continue; // the crux GitHub alone cannot answer
// Show the most specific path we saw for this file.
const display =
touches.map((t) => t.display).sort((a, b) => b.length - a.length)[0] ?? touches[0]!.display;
out.push({
id: `col_${file}_${Date.now()}`,
podId: group[0]!.podId,
file,
symbol: group.find((g) => g.currentSymbol)?.currentSymbol,
id: `col_${fileKey(display)}_${Date.now()}`,
podId: contexts[0]?.podId ?? 'demo-pod',
file: display,
symbol: contexts.find((c) => c.currentSymbol)?.currentSymbol,
engineers,
severity: 'warn',
githubState: { ...github, unpushed: anyUnpushed },
+147
View File
@@ -0,0 +1,147 @@
import type { Collision, EngineerContext } from '@podman/shared';
import { env } from '../env.js';
import type { GitState } from '../memory/db.js';
import { semanticSimilarity } from '../memory/vectors.js';
export interface ResearchOpts {
similarity?: (a: string, b: string) => Promise<number | null>;
threshold?: number;
}
interface EditorFile {
engineerId: string;
file: string;
symbol?: string;
activity?: string;
}
interface Candidate {
collision: Collision;
score: number;
}
function stripGitPrefix(raw: string): string {
return raw.trim().replace(/^(\?\?|[MADRCU!]{1,2})\s+/, '');
}
/** Case/whitespace-insensitive identity so vision "Karti" and git "karti" are
* recognized as the same person and never flagged researching-vs-editing self. */
function canonicalName(raw: string): string {
return raw.trim().toLowerCase();
}
function fileStem(raw: string): string {
const base = stripGitPrefix(raw).split(/[\\/]/).pop()?.trim().toLowerCase() ?? '';
return base.replace(/\.[^.]+$/, '');
}
function words(raw: string): string[] {
return raw
.toLowerCase()
.replace(/[^a-z0-9]+/g, ' ')
.split(/\s+/)
.filter((word) => word.length >= 3);
}
function uniqueTokens(raw: string): Set<string> {
const tokens = new Set(words(raw));
for (const word of [...tokens]) {
if (word.endsWith('kit')) tokens.add(word.replace(/kit$/, ''));
if (word.endsWith('s')) tokens.add(word.slice(0, -1));
}
return tokens;
}
function fallbackMatches(researchText: string, fileText: string): boolean {
const research = uniqueTokens(researchText);
const file = uniqueTokens(fileText);
for (const token of file) {
if (research.has(token)) return true;
}
return false;
}
function collectEditorFiles(
contexts: EngineerContext[],
gitStates: Map<string, GitState> | undefined,
): EditorFile[] {
const files = new Map<string, EditorFile>();
const add = (editor: EditorFile): void => {
const stem = fileStem(editor.file);
if (!stem) return;
files.set(`${editor.engineerId}:${stripGitPrefix(editor.file)}`, editor);
};
for (const [engineerId, git] of gitStates ?? []) {
for (const changed of git.changedFiles) {
add({ engineerId, file: changed });
}
}
for (const context of contexts) {
if (context.mode !== 'research' && context.currentFile) {
add({
engineerId: context.engineerId,
file: context.currentFile,
symbol: context.currentSymbol,
activity: context.activity,
});
}
}
return [...files.values()];
}
export async function detectResearchOverlaps(
contexts: EngineerContext[],
gitStates: Map<string, GitState> | undefined,
opts: ResearchOpts = {},
): Promise<Collision[]> {
const similarity = opts.similarity ?? semanticSimilarity;
const threshold = opts.threshold ?? env.RESEARCH_OVERLAP_THRESHOLD;
const researchers = contexts.filter((c) => c.mode === 'research' && c.researchTopic);
const editorFiles = collectEditorFiles(contexts, gitStates);
const bestByResearcher = new Map<string, Candidate>();
for (const researcher of researchers) {
const topic = researcher.researchTopic?.trim();
if (!topic) continue;
const source = researcher.researchSource?.trim();
const researchText = [topic, source].filter(Boolean).join(' ');
for (const editor of editorFiles) {
if (canonicalName(editor.engineerId) === canonicalName(researcher.engineerId)) continue;
const stem = fileStem(editor.file);
if (!stem) continue;
const fileText = [stem, editor.symbol, editor.activity].filter(Boolean).join(' ');
const score = await similarity(researchText, fileText);
const matched = score === null ? fallbackMatches(researchText, fileText) : score >= threshold;
if (!matched) continue;
const rank = score ?? 1;
const existing = bestByResearcher.get(researcher.engineerId);
if (existing && existing.score >= rank) continue;
bestByResearcher.set(researcher.engineerId, {
score: rank,
collision: {
id: `col_research_${stem}_${Date.now()}`,
podId: researcher.podId,
file: editor.file,
symbol: editor.symbol,
engineers: [editor.engineerId, researcher.engineerId],
severity: 'warn',
overlapKind: 'research',
researchTopic: topic,
...(source ? { researchSource: source } : {}),
researcher: researcher.engineerId,
editor: editor.engineerId,
detectedAt: new Date().toISOString(),
},
});
}
}
return [...bestByResearcher.values()].map((candidate) => candidate.collision);
}
+13 -1
View File
@@ -1,4 +1,6 @@
import 'dotenv/config';
import { config } from 'dotenv';
config({ path: ['.env.local', '../.env.local', '.env', '../.env'] });
function req(name: string): string {
const v = process.env[name];
@@ -21,10 +23,17 @@ export const env = {
LIVEKIT_URL: req('LIVEKIT_URL'),
LIVEKIT_API_KEY: req('LIVEKIT_API_KEY'),
LIVEKIT_API_SECRET: req('LIVEKIT_API_SECRET'),
LIVEKIT_AGENT_NAME: opt('LIVEKIT_AGENT_NAME'),
LIVEKIT_CONVERSATION_AGENT_NAME: opt(
'LIVEKIT_CONVERSATION_AGENT_NAME',
'podman-live-conversation',
),
// Gemini
GEMINI_API_KEY: reqAny('GEMINI_API_KEY', ['GOOGLE_API_KEY', 'GOOGLE_GENERATIVE_AI_API_KEY']),
GEMINI_VISION_MODEL: opt('GEMINI_VISION_MODEL', 'gemini-2.0-flash'),
GEMINI_LIVE_MODEL: opt('GEMINI_LIVE_MODEL', 'gemini-3.1-flash-tts-preview'),
GEMINI_CONVERSATION_MODEL: opt('GEMINI_CONVERSATION_MODEL', 'gemini-3.1-flash-live-preview'),
GEMINI_TTS_VOICE: opt('GEMINI_TTS_VOICE', 'Charon'),
GEMINI_EMBEDDING_MODEL: opt('GEMINI_EMBEDDING_MODEL', 'gemini-embedding-001'),
// GitHub
GITHUB_TOKEN: req('GITHUB_TOKEN'),
@@ -35,7 +44,10 @@ export const env = {
VOYAGE_EMBEDDING_MODEL: opt('VOYAGE_EMBEDDING_MODEL', 'voyage-4-lite'),
// Server
PORT: Number(opt('PORT', '8787')),
CLERK_SECRET_KEY: opt('CLERK_SECRET_KEY'),
NUDGE_COOLDOWN_MS: Number(opt('NUDGE_COOLDOWN_MS', '180000')),
RESEARCH_OVERLAP_THRESHOLD: Number(opt('RESEARCH_OVERLAP_THRESHOLD', '0.6')),
INTERNAL_AGENT_TOKEN: opt('INTERNAL_AGENT_TOKEN'),
} as const;
export function repoParts(): { owner: string; repo: string } {
+68 -37
View File
@@ -8,45 +8,9 @@ import type { PodGraph } from '@podman/shared';
* auth* — is the continual-learning story the demo lights up.
*/
export function createDemoPodGraph(podId: string): PodGraph {
const base = Date.now();
const at = (secAgo: number): string => new Date(base - secAgo * 1000).toISOString();
return {
podId,
generatedAt: new Date().toISOString(),
loop: [
{ key: 'observe', title: 'OBSERVE', value: '5', detail: '~5/s vision contexts', active: false },
{ key: 'store', title: 'STORE', value: '124', detail: 'memory vectors · Atlas', active: false },
{ key: 'predict', title: 'PREDICT', value: '1', detail: 'open risk path', active: true },
{ key: 'outcome', title: 'OUTCOME', value: '1/0', detail: 'accepted · dismissed', active: false },
{ key: 'adapt', title: 'ADAPT', value: '3', detail: 'learned owners', active: false },
],
activity: [
{
id: 'demo-learn',
at: at(20),
kind: 'learned_from',
text: 'Memory updated: Karti owns auth.ts (confidence ↑)',
},
{ id: 'demo-out', at: at(24), kind: 'outcome', text: 'Intervention accepted by the pod' },
{
id: 'demo-warn',
at: at(40),
kind: 'warns',
text: 'PodMan: "Karti & Yahya are both in auth.ts — open a sync PR?" → card sent',
},
{
id: 'demo-col',
at: at(58),
kind: 'collision',
text: 'Critical overlap on auth.ts · Karti + Yahya',
},
{
id: 'demo-edit',
at: at(72),
kind: 'editing',
text: 'Yahya opened auth.ts — unpushed changes',
},
],
// Kept consistent with the graph below (3 owner engineers, 1 collision file,
// 1 of 1 interventions accepted) so the numbers never contradict the picture.
metrics: [
@@ -66,6 +30,73 @@ export function createDemoPodGraph(podId: string): PodGraph {
detail: 'Interventions accepted vs total this session.',
},
],
loop: {
activeStep: 'adapt',
steps: [
{
key: 'observe',
label: 'Observe',
value: '2',
detail: 'Screen context and local git state show two active editors.',
status: 'complete',
},
{
key: 'store',
label: 'Store',
value: '7',
detail: 'Observations, collisions, interventions, and outcomes are in MongoDB.',
status: 'complete',
},
{
key: 'predict',
label: 'Predict',
value: '2',
detail: 'Same-file risk paths are detected before push.',
status: 'complete',
},
{
key: 'outcome',
label: 'Outcome',
value: '6',
detail: 'Accepted and dismissed outcomes supervise future routing.',
status: 'complete',
},
{
key: 'adapt',
label: 'Adapt',
value: '1',
detail: 'Accepted real collision created a learned_from edge.',
status: 'complete',
},
],
},
activity: [
{
id: 'demo-learned-auth',
at: new Date().toISOString(),
kind: 'learned',
title: 'Learned Karti owns auth.ts',
detail: 'Accepted sync PR outcome created a durable learned_from path.',
nodeId: 'engineer:karti',
edgeId: 'e7',
},
{
id: 'demo-intervention-sync-pr',
at: new Date().toISOString(),
kind: 'intervention',
title: 'Intervention: sync PR',
detail: 'PodMan offered a small coordination card before voice.',
nodeId: 'intervention:sync-pr',
},
{
id: 'demo-collision-auth',
at: new Date().toISOString(),
kind: 'collision',
title: 'Collision risk on auth.ts',
detail: 'Karti and Yahya converged on unpushed work.',
nodeId: 'collision:auth',
},
],
nodes: [
{
id: 'engineer:shakthi',
@@ -224,7 +255,7 @@ export function createDemoPodGraph(podId: string): PodGraph {
source: 'collision:auth',
target: 'intervention:sync-pr',
kind: 'warns',
label: 'nudges',
label: 'routes',
strength: 0.9,
},
{
+195 -229
View File
@@ -3,16 +3,11 @@ import type {
PodGraphNode,
PodGraphEdge,
PodGraphMetric,
PodLearningLoop,
PodGraphActivity,
PodGraphNodeKind,
PodGraphEdgeKind,
PodGraphNodeStatus,
LearningStage,
LearningStageKey,
ActivityEvent,
EngineerContext,
Collision,
Intervention,
InterventionOutcome,
} from '@podman/shared';
import { collections, getGitStates, getDb } from '../memory/db.js';
@@ -155,183 +150,75 @@ function layout(nodes: PodGraphNode[]): void {
const SEVERITY_WEIGHT: Record<string, number> = { info: 0.4, warn: 0.7, critical: 1 };
/** Parse any timestamp-ish value to epoch ms (0 when missing/unparseable). */
function ms(t: string | Date | null | undefined): number {
if (!t) return 0;
const v = new Date(t).getTime();
return Number.isFinite(v) ? v : 0;
}
const OBSERVE_WINDOW_MS = 60_000;
/**
* Live counts for the learning-loop rail (observe→store→predict→outcome→adapt).
* The "active" stage is the one whose latest underlying event is most recent —
* with deeper stages winning ties so the rail lights up at the furthest point
* the pod reached this session. Additive: derived from already-fetched docs.
*/
function buildLoop(opts: {
now: number;
observations: EngineerContext[];
collisions: Collision[];
outcomes: InterventionOutcome[];
riskPaths: number;
vectorCount: number;
learnedOwners: number;
}): LearningStage[] {
const { now, observations, collisions, outcomes, riskPaths, vectorCount, learnedOwners } = opts;
const recentObs = observations.filter((o) => now - ms(o.observedAt) < OBSERVE_WINDOW_MS).length;
const rate = (recentObs / 60).toFixed(1);
const accepted = outcomes.filter((o) => o.accepted).length;
const dismissed = outcomes.filter((o) => !o.accepted).length;
// Latest event time per stage; `store` sits just behind `predict` so a shared
// collision timestamp resolves to PREDICT rather than STORE.
const latestObs = Math.max(0, ...observations.map((o) => ms(o.observedAt)));
const latestCol = Math.max(0, ...collisions.map((c) => ms(c.detectedAt)));
const latestOut = Math.max(0, ...outcomes.map((o) => ms(o.recordedAt)));
const latestAdapt = Math.max(
0,
...outcomes.filter((o) => o.accepted && o.wasRealCollision).map((o) => ms(o.recordedAt)),
);
const refs: Array<[LearningStageKey, number]> = [
['observe', latestObs],
['store', latestCol ? latestCol - 1 : 0],
['predict', latestCol],
['outcome', latestOut],
['adapt', latestAdapt],
];
let activeKey: LearningStageKey = 'observe';
let best = 0;
for (const [k, t] of refs) {
if (t > 0 && t >= best) {
best = t;
activeKey = k;
}
}
const stages: Array<Omit<LearningStage, 'active'>> = [
{ key: 'observe', title: 'OBSERVE', value: String(recentObs), detail: `~${rate}/s vision contexts` },
{ key: 'store', title: 'STORE', value: String(vectorCount), detail: 'memory vectors · Atlas' },
function buildLoop(input: {
observations: number;
gitStates: number;
collisions: number;
interventions: number;
outcomes: number;
acceptedReal: number;
learnedEdges: number;
}): PodLearningLoop {
const stored = input.observations + input.gitStates + input.interventions + input.outcomes;
return {
activeStep:
input.acceptedReal > 0
? 'adapt'
: input.outcomes > 0
? 'outcome'
: input.collisions > 0
? 'predict'
: input.observations + input.gitStates > 0
? 'store'
: 'observe',
steps: [
{
key: 'observe',
label: 'Observe',
value: String(input.observations + input.gitStates),
detail: 'Recent vision observations plus local git-state reports.',
status: input.observations + input.gitStates > 0 ? 'complete' : 'quiet',
},
{
key: 'store',
label: 'Store',
value: String(stored),
detail: 'MongoDB records available to recall for this pod.',
status: stored > 0 ? 'complete' : 'quiet',
},
{
key: 'predict',
title: 'PREDICT',
value: String(riskPaths),
detail: `open risk path${riskPaths === 1 ? '' : 's'}`,
label: 'Predict',
value: String(input.collisions),
detail: 'Distinct collision signatures detected from live work.',
status: input.collisions > 0 ? 'complete' : 'quiet',
},
{
key: 'outcome',
label: 'Outcome',
value: String(input.outcomes),
detail: 'Accepted and dismissed intervention outcomes.',
status: input.outcomes > 0 ? 'complete' : 'quiet',
},
{ key: 'outcome', title: 'OUTCOME', value: `${accepted}/${dismissed}`, detail: 'accepted · dismissed' },
{
key: 'adapt',
title: 'ADAPT',
value: String(learnedOwners),
detail: `learned owner${learnedOwners === 1 ? '' : 's'}`,
label: 'Adapt',
value: String(input.learnedEdges),
detail: 'Learned graph edges created from accepted real outcomes.',
status: input.acceptedReal > 0 ? 'complete' : 'planned',
},
];
return stages.map((s) => ({ ...s, active: s.key === activeKey }));
],
};
}
/**
* Merge + time-sort recent events into the activity stream feed. Reuses the same
* de-noise (isFilePath / ENGINEER_NOISE / signature collapse) as the graph so
* the feed never shows junk paths or test-artifact engineers. Capped to 8.
*/
function buildActivity(opts: {
observations: EngineerContext[];
collisions: Collision[];
interventions: Intervention[];
outcomes: InterventionOutcome[];
ownership: Record<string, string>;
}): ActivityEvent[] {
const { observations, collisions, interventions, outcomes, ownership } = opts;
const cleanEng = (n: string): boolean => Boolean(n) && !ENGINEER_NOISE.test(n);
const out: ActivityEvent[] = [];
// editing — newest observation per (engineer, file); observations arrive desc.
const seenEdit = new Set<string>();
for (const o of observations) {
if (!o.engineerId || !cleanEng(o.engineerId)) continue;
const file = o.currentFile ? normalizeFile(o.currentFile) : '';
if (!isFilePath(file)) continue;
const key = `${o.engineerId.toLowerCase()}|${file}`;
if (seenEdit.has(key)) continue;
seenEdit.add(key);
out.push({
id: `edit:${o.engineerId}:${file}`,
at: o.observedAt,
kind: 'editing',
text: `${o.engineerId} opened ${shortLabel(file)}${
o.hasUnpushedChanges ? ' — unpushed changes' : ''
}`,
});
}
// collision — collapse by signature, newest first.
const seenCol = new Set<string>();
for (const c of collisions) {
const file = normalizeFile(c.file);
if (!isFilePath(file)) continue;
const sig = (c as { memorySignature?: string }).memorySignature ?? `${file}#${c.symbol ?? ''}`;
if (seenCol.has(sig)) continue;
seenCol.add(sig);
const engs = c.engineers.filter(cleanEng);
if (!engs.length) continue;
out.push({
id: `col:${c.id}`,
at: c.detectedAt,
kind: 'collision',
text: `${c.severity === 'critical' ? 'Critical overlap' : 'Overlap'} on ${shortLabel(
file,
)} · ${engs.join(' + ')}`,
});
}
// warns — interventions PodMan raised.
for (const iv of interventions) {
if (!iv.message) continue;
const msg = iv.message.length > 64 ? `${iv.message.slice(0, 61)}` : iv.message;
out.push({
id: `warn:${iv.id}`,
at: iv.createdAt,
kind: 'warns',
text: `PodMan: "${msg}" → card sent`,
});
}
// outcome + learned_from — the supervised learning beat.
const colById = new Map(collisions.map((c) => [c.id, c]));
const ivById = new Map(interventions.map((i) => [i.id, i]));
for (const o of outcomes) {
if (!o.accepted) continue;
out.push({
id: `out:${o.interventionId}`,
at: o.recordedAt,
kind: 'outcome',
text: 'Intervention accepted by the pod',
});
if (!o.wasRealCollision) continue;
const iv = ivById.get(o.interventionId);
const col = iv ? colById.get(iv.collisionId) : colById.get(o.collisionId);
if (!col) continue;
const file = normalizeFile(col.file);
if (!isFilePath(file)) continue;
const owner =
(o as { learnedOwner?: string }).learnedOwner ??
ownership[file] ??
col.engineers.find(cleanEng) ??
col.engineers[0];
if (!owner) continue;
out.push({
id: `learn:${o.interventionId}`,
at: o.recordedAt,
kind: 'learned_from',
text: `Memory updated: ${owner} owns ${shortLabel(file)} (confidence ↑)`,
});
}
out.sort((a, b) => ms(b.at) - ms(a.at));
return out.slice(0, 8);
function pushActivity(
activity: PodGraphActivity[],
item: PodGraphActivity,
seen: Set<string>,
): void {
if (seen.has(item.id)) return;
seen.add(item.id);
activity.push(item);
}
export async function materializePodGraph(podId: string): Promise<PodGraph | null> {
@@ -360,6 +247,8 @@ export async function materializePodGraph(podId: string): Promise<PodGraph | nul
}
const b: Builder = { nodes: new Map(), edges: new Map() };
const activity: PodGraphActivity[] = [];
const activityIds = new Set<string>();
const now = Date.now();
// 1. Baseline engineer nodes from the roster.
@@ -379,6 +268,18 @@ export async function materializePodGraph(podId: string): Promise<PodGraph | nul
if (isFilePath(file)) {
const f = upsertNode(b, 'file', file, { label: shortLabel(file), summary: file });
upsertEdge(b, eng, f, 'editing', o.activity ?? 'edits', Math.max(0.4, o.confidence ?? 0.5));
pushActivity(
activity,
{
id: `editing:${o.engineerId}:${file}:${String(o.observedAt ?? '')}`,
at: String(o.observedAt ?? new Date().toISOString()),
kind: 'editing',
title: `${o.engineerId} editing ${shortLabel(file)}`,
detail: o.activity ?? 'Vision observed active work.',
nodeId: f,
},
activityIds,
);
}
}
@@ -437,6 +338,18 @@ export async function materializePodGraph(podId: string): Promise<PodGraph | nul
const eng = upsertNode(b, 'engineer', name, { label: name });
upsertEdge(b, eng, cNode, 'collides', 'in', SEVERITY_WEIGHT[col.severity] ?? 0.7);
}
pushActivity(
activity,
{
id: `collision:${col.id}`,
at: col.detectedAt,
kind: 'collision',
title: `Collision risk on ${shortLabel(file)}`,
detail: `${col.engineers.join(' + ')} converged on ${file}.`,
nodeId: cNode,
},
activityIds,
);
sigToNode.set(sig, cNode);
colNodeFor.set(col.id, cNode);
if (!isPriority) distinctCollisions++;
@@ -479,7 +392,19 @@ export async function materializePodGraph(podId: string): Promise<PodGraph | nul
: 'watch',
summary: iv.message,
});
upsertEdge(b, colNode, ivNode, 'warns', 'nudges', 0.85);
upsertEdge(b, colNode, ivNode, 'warns', 'routes', 0.85);
pushActivity(
activity,
{
id: `intervention:${iv.id}`,
at: iv.createdAt,
kind: 'intervention',
title: `Intervention: ${b.nodes.get(ivNode)?.label ?? iv.kind}`,
detail: iv.message,
nodeId: ivNode,
},
activityIds,
);
ivNodeForCol.set(colNode, ivNode);
}
@@ -502,8 +427,66 @@ export async function materializePodGraph(podId: string): Promise<PodGraph | nul
if (ivNode) {
const ivObj = b.nodes.get(ivNode);
if (ivObj) ivObj.status = 'learned';
const before = b.edges.size;
upsertEdge(b, ivNode, engNode, 'learned_from', `learned: owns ${file}`, 0.6);
const edgeId = `${'learned_from'}:${ivNode}->${engNode}`;
pushActivity(
activity,
{
id: `learned:${out.interventionId}:${owner}:${file}`,
at: out.recordedAt,
kind: 'learned',
title: `Learned ${owner} owns ${shortLabel(file)}`,
detail: 'Accepted real outcome created a durable learned_from path.',
nodeId: engNode,
edgeId: before === b.edges.size ? undefined : edgeId,
},
activityIds,
);
}
pushActivity(
activity,
{
id: `outcome:${out.interventionId}:${out.recordedAt}`,
at: out.recordedAt,
kind: 'outcome',
title: out.accepted ? 'Outcome accepted' : 'Outcome dismissed',
detail: out.wasRealCollision ? 'Marked as a real collision.' : 'Marked as noise.',
},
activityIds,
);
}
// 7. Suppressed repeats: PodMan stayed quiet on a recurring collision because
// the signature was dismissed before — the negative-feedback loop made visible
// (Feature A). Stamped at repeat time, so it sorts as recent activity.
const suppressionDocs = await c.suppressions
.find({ podId })
.sort({ suppressedAt: -1 })
.limit(50)
.toArray();
// Collapse to one beat per file (keep the most recent — docs are sorted desc)
// so pre-fix duplicate rows never render as spam. The stable per-file id also
// dedupes through pushActivity's `seen` set.
const seenSuppressedFiles = new Set<string>();
for (const s of suppressionDocs) {
const sFile = normalizeFile(s.file);
if (!isFilePath(sFile)) continue;
const fileKey = sFile.toLowerCase();
if (seenSuppressedFiles.has(fileKey)) continue;
seenSuppressedFiles.add(fileKey);
const sEngs = (s.engineers ?? []).join(' + ') || 'teammates';
pushActivity(
activity,
{
id: `suppressed:${fileKey}`,
at: s.suppressedAt,
kind: 'suppressed',
title: `Suppressed — ${shortLabel(sFile)} repeat silenced`,
detail: `${sEngs} on ${sFile} recurred, but it was dismissed before — PodMan stayed quiet.`,
},
activityIds,
);
}
// Prune test-artifact engineers, then anything left orphaned by that.
@@ -544,27 +527,35 @@ export async function materializePodGraph(podId: string): Promise<PodGraph | nul
const nodes = [...b.nodes.values()];
// No real activity beyond the bare roster -> let the caller fall back to demo.
const hasActivity = nodes.some((n) => n.kind !== 'engineer');
// Suppression beats are real negative-feedback proof even when they add no
// nodes/edges, so they satisfy the gate too — a clean pod with preserved
// suppressions must not fall back to the demo graph and hide the proof.
const hasSuppressed = activity.some((a) => a.kind === 'suppressed');
const hasActivity = hasSuppressed || nodes.some((n) => n.kind !== 'engineer');
if (!hasActivity) return null;
layout(nodes);
// Metrics are derived from the FINAL de-noised graph (not raw docs) so the
// numbers match what's actually on screen. Counting raw collision signatures /
// accepted-outcome rows inflates them with test churn (e.g. 50 "risk paths" for
// 2 files), which reads as fake — these count distinct visible entities instead.
const finalEdges = [...b.edges.values()];
const acceptedReal = outcomeDocs.filter((o) => o.accepted && o.wasRealCollision).length;
const totalOutcomes = outcomeDocs.length;
// Raw distinct collision signatures — kept for the learning-loop throughput view.
const riskPaths = new Set(
collisionDocs.map(
(col) =>
(col as { memorySignature?: string }).memorySignature ??
`${normalizeFile(col.file)}#${col.symbol ?? ''}`,
),
).size;
// Open risk paths = distinct files carrying a surviving collision (the triangles).
// Headline metric cards are derived from the FINAL de-noised graph so they match
// what's drawn. Counting raw collision signatures / accepted-outcome rows inflates
// them with test churn (e.g. 50 "risk paths" for 4 files), which reads as fake.
const finalEdges = [...b.edges.values()];
const riskFiles = new Set<string>();
for (const e of finalEdges) {
if (e.kind === 'touches' && b.nodes.get(e.source)?.kind === 'file') riskFiles.add(e.source);
}
const collisionNodeCount = nodes.filter((n) => n.kind === 'collision').length;
const riskPaths = riskFiles.size || collisionNodeCount;
// Learned owners = distinct engineers PodMan retained as owners from accepted
// interventions (the owns / learned_from edges actually drawn).
const openRiskPaths = riskFiles.size || nodes.filter((n) => n.kind === 'collision').length;
const ownerSet = new Set<string>();
for (const e of finalEdges) {
if (e.kind === 'learned_from') ownerSet.add(e.target);
@@ -572,10 +563,6 @@ export async function materializePodGraph(podId: string): Promise<PodGraph | nul
}
const learnedOwners = [...ownerSet].filter((id) => b.nodes.get(id)?.kind === 'engineer').length;
const acceptedReal = outcomeDocs.filter((o) => o.accepted && o.wasRealCollision).length;
const totalOutcomes = outcomeDocs.length;
const acceptRate = totalOutcomes ? Math.round((acceptedReal / totalOutcomes) * 100) : null;
const metrics: PodGraphMetric[] = [
{
label: 'Learned owners',
@@ -584,46 +571,17 @@ export async function materializePodGraph(podId: string): Promise<PodGraph | nul
},
{
label: 'Open risk paths',
value: String(riskPaths),
detail: `${riskPaths === 1 ? 'File' : 'Files'} with two or more converging editors.`,
value: String(openRiskPaths),
detail: `${openRiskPaths === 1 ? 'File' : 'Files'} with two or more converging editors.`,
},
{
label: 'Accept rate',
value: acceptRate == null ? '—' : `${acceptRate}%`,
value: totalOutcomes ? `${Math.round((acceptedReal / totalOutcomes) * 100)}%` : '—',
detail: 'Interventions accepted vs total this session.',
},
];
// Stored vectors for the STORE stage: prefer a real memory_vectors count,
// fall back to collisions carrying an embedding, then to collision count.
let vectorCount = 0;
try {
vectorCount = await db.collection('memory_vectors').countDocuments({ podId });
} catch {
/* memory_vectors is optional */
}
if (!vectorCount)
vectorCount = collisionDocs.filter(
(c) => (c as { embedding?: number[] }).embedding?.length,
).length;
if (!vectorCount) vectorCount = collisionDocs.length;
const loop = buildLoop({
now,
observations,
collisions: collisionDocs,
outcomes: outcomeDocs,
riskPaths,
vectorCount,
learnedOwners,
});
const activity = buildActivity({
observations,
collisions: collisionDocs,
interventions: interventionDocs,
outcomes: outcomeDocs,
ownership,
});
const learnedEdges = [...b.edges.values()].filter((e) => e.kind === 'learned_from').length;
activity.sort((a, z) => String(z.at).localeCompare(String(a.at)));
return {
podId,
@@ -631,7 +589,15 @@ export async function materializePodGraph(podId: string): Promise<PodGraph | nul
nodes,
edges: [...b.edges.values()],
metrics,
loop,
activity,
loop: buildLoop({
observations: observations.length,
gitStates: gitStates.size,
collisions: riskPaths,
interventions: interventionDocs.length,
outcomes: totalOutcomes,
acceptedReal,
learnedEdges,
}),
activity: activity.slice(0, 12),
};
}
+349
View File
@@ -0,0 +1,349 @@
import { randomUUID } from 'node:crypto';
import { execFile } from 'node:child_process';
import { promisify } from 'node:util';
import { Room as LiveKitRoom } from '@livekit/rtc-node';
import { AccessToken } from 'livekit-server-sdk';
import {
DATA_TOPIC,
type DataMessage,
type HermesJob,
type HermesJobEvent,
type HermesJobEventType,
type HermesJobInput,
type HermesJobStatus,
type HermesRiskLevel,
} from '@podman/shared';
import { env, repoParts } from '../env.js';
import { getDb } from '../memory/db.js';
const execFileAsync = promisify(execFile);
const encoder = new TextEncoder();
const MAX_OUTPUT = 3_000;
const COMMAND_TIMEOUT_MS = 45_000;
const runners = new Map<string, AbortController>();
function now(): string {
return new Date().toISOString();
}
function truncate(value: string): string {
return value.length > MAX_OUTPUT ? `${value.slice(0, MAX_OUTPUT)}\n...[truncated]` : value;
}
function redact(value: string): string {
return value
.replace(/AIza[0-9A-Za-z_-]{20,}/g, '[redacted-google-key]')
.replace(/API[_-]?SECRET=[^\s]+/gi, 'API_SECRET=[redacted]')
.replace(/TOKEN=[^\s]+/gi, 'TOKEN=[redacted]')
.replace(/mongodb(\+srv)?:\/\/[^@\s]+@/gi, 'mongodb$1://[redacted]@');
}
function normalizeRisk(value: unknown): HermesRiskLevel {
return value === 'safe_write' ||
value === 'commit_allowed' ||
value === 'deploy_allowed' ||
value === 'read_only'
? value
: 'read_only';
}
async function hermesJobs() {
return (await getDb()).collection<HermesJob>('hermes_jobs');
}
async function hermesJobEvents() {
return (await getDb()).collection<HermesJobEvent>('hermes_job_events');
}
export async function ensureHermesJobIndexes(): Promise<void> {
const db = await getDb();
await Promise.allSettled([
db.collection('hermes_jobs').createIndex({ id: 1 }, { unique: true }),
db.collection('hermes_jobs').createIndex({ sessionId: 1, status: 1, updatedAt: -1 }),
db.collection('hermes_jobs').createIndex({ podId: 1, updatedAt: -1 }),
db.collection('hermes_job_events').createIndex({ jobId: 1, createdAt: 1 }),
db.collection('hermes_job_events').createIndex({ sessionId: 1, createdAt: -1 }),
]);
}
export async function createHermesJob(input: Partial<HermesJobInput>): Promise<HermesJob> {
const prompt = typeof input.prompt === 'string' ? input.prompt.trim() : '';
if (!prompt) throw new Error('prompt is required');
const createdAt = now();
const job: HermesJob = {
id: `hermes_job_${randomUUID()}`,
podId: input.podId || 'demo-pod',
identity: input.identity || 'developer',
sessionId: input.sessionId || 'unknown',
conversationRoom: input.conversationRoom,
prompt,
contextScope: input.contextScope || 'current_repo',
targetRepository: input.targetRepository || env.GITHUB_REPO,
riskLevel: normalizeRisk(input.riskLevel),
requiresConfirmation: input.requiresConfirmation === true,
successCriteria: Array.isArray(input.successCriteria)
? input.successCriteria.map(String).filter(Boolean).slice(0, 8)
: ['Hermes reports what it inspected and what changed.'],
parentJobId: input.parentJobId,
status: 'queued',
createdAt,
updatedAt: createdAt,
};
await (await hermesJobs()).insertOne(job);
await appendHermesJobEvent(job.id, 'accepted', 'Hermes accepted the task.', {
riskLevel: job.riskLevel,
contextScope: job.contextScope,
});
void runHermesJob(job.id);
return job;
}
export async function getHermesJob(jobId: string): Promise<HermesJob | null> {
return (await hermesJobs()).findOne({ id: jobId }, { projection: { _id: 0 } });
}
export async function getActiveHermesJobForSession(sessionId: string): Promise<HermesJob | null> {
return (await hermesJobs()).findOne(
{ sessionId, status: { $in: ['queued', 'running', 'waiting_for_confirmation', 'aborting'] } },
{ projection: { _id: 0 }, sort: { updatedAt: -1 } },
);
}
export async function getLatestHermesJobForSession(sessionId: string): Promise<HermesJob | null> {
return (await hermesJobs()).findOne(
{ sessionId },
{ projection: { _id: 0 }, sort: { updatedAt: -1 } },
);
}
export async function listHermesJobEvents(jobId: string, limit = 40): Promise<HermesJobEvent[]> {
return (await hermesJobEvents())
.find({ jobId }, { projection: { _id: 0 } })
.sort({ createdAt: 1 })
.limit(Math.min(limit, 200))
.toArray();
}
export async function appendHermesJobEvent(
jobId: string,
type: HermesJobEventType,
message: string,
data?: Record<string, unknown>,
): Promise<HermesJobEvent> {
const job = await getHermesJob(jobId);
if (!job) throw new Error('job not found');
const event: HermesJobEvent = {
id: `hermes_evt_${randomUUID()}`,
jobId,
podId: job.podId,
sessionId: job.sessionId,
type,
message: redact(truncate(message)),
data,
createdAt: now(),
};
await (await hermesJobEvents()).insertOne(event);
await (
await hermesJobs()
).updateOne(
{ id: jobId },
{ $set: { updatedAt: event.createdAt, lastHeartbeatAt: event.createdAt } },
);
if (job.conversationRoom) {
void publishHermesJobEvent(job.conversationRoom, event).catch((err) =>
console.warn(`[hermes-job] data publish failed: ${(err as Error).message}`),
);
}
return event;
}
export async function abortHermesJob(jobId: string): Promise<HermesJob | null> {
const job = await getHermesJob(jobId);
if (!job) return null;
const abortAt = now();
await (
await hermesJobs()
).updateOne(
{ id: jobId },
{ $set: { status: 'aborting', abortRequestedAt: abortAt, updatedAt: abortAt } },
);
runners.get(jobId)?.abort();
await appendHermesJobEvent(jobId, 'heartbeat', 'Hermes is aborting the current job.');
return getHermesJob(jobId);
}
async function setStatus(jobId: string, status: HermesJobStatus, patch: Partial<HermesJob> = {}) {
await (
await hermesJobs()
).updateOne({ id: jobId }, { $set: { status, updatedAt: now(), ...patch } });
}
async function publishHermesJobEvent(roomName: string, event: HermesJobEvent): Promise<void> {
const room = new LiveKitRoom();
try {
const at = new AccessToken(env.LIVEKIT_API_KEY, env.LIVEKIT_API_SECRET, {
identity: `podman-hermes-job-${Date.now()}`,
name: 'PodMan Hermes jobs',
ttl: '5m',
});
at.addGrant({
roomJoin: true,
room: roomName,
canPublish: true,
canSubscribe: false,
canPublishData: true,
});
await room.connect(env.LIVEKIT_URL, await at.toJwt(), {
autoSubscribe: false,
dynacast: false,
});
const data: DataMessage = { type: 'HERMES_JOB_EVENT', event };
await room.localParticipant?.publishData(encoder.encode(JSON.stringify(data)), {
reliable: true,
topic: DATA_TOPIC,
});
} finally {
await room.disconnect().catch(() => {});
}
}
async function runCommand(
jobId: string,
label: string,
command: string,
args: string[],
signal: AbortSignal,
): Promise<string> {
await appendHermesJobEvent(jobId, 'step_started', `${label} started.`);
const started = Date.now();
const { stdout, stderr } = await execFileAsync(command, args, {
cwd: process.cwd(),
timeout: COMMAND_TIMEOUT_MS,
signal,
maxBuffer: 1024 * 1024,
});
const output = redact(truncate([stdout, stderr].filter(Boolean).join('\n').trim()));
await appendHermesJobEvent(jobId, 'step_output', output || `${label} produced no output.`, {
label,
durationMs: Date.now() - started,
});
await appendHermesJobEvent(jobId, 'step_completed', `${label} completed.`);
return output;
}
function wantsBuild(prompt: string, criteria: string[]): boolean {
const haystack = `${prompt} ${criteria.join(' ')}`.toLowerCase();
return /build|typecheck|test|lint|broken|failing|verify/.test(haystack);
}
function wantsMongo(prompt: string, scope: string): boolean {
return scope === 'mongodb' || /mongo|database|telemetry|logs?|memory/.test(prompt.toLowerCase());
}
function wantsGithub(prompt: string, scope: string): boolean {
return (
scope === 'github' || /github|branch|pr|pull request|commit|diff/.test(prompt.toLowerCase())
);
}
async function inspectMongo(jobId: string) {
await appendHermesJobEvent(jobId, 'step_started', 'MongoDB inspection started.');
const db = await getDb();
const [observations, collisions, interventions, outcomes, jobs] = await Promise.all([
db.collection('observations').estimatedDocumentCount(),
db.collection('collisions').estimatedDocumentCount(),
db.collection('interventions').estimatedDocumentCount(),
db.collection('outcomes').estimatedDocumentCount(),
db.collection('hermes_jobs').estimatedDocumentCount(),
]);
await appendHermesJobEvent(
jobId,
'step_output',
`MongoDB is reachable. Counts: observations=${observations}, collisions=${collisions}, interventions=${interventions}, outcomes=${outcomes}, hermes_jobs=${jobs}.`,
);
await appendHermesJobEvent(jobId, 'step_completed', 'MongoDB inspection completed.');
}
async function inspectGithub(jobId: string) {
await appendHermesJobEvent(jobId, 'step_started', 'GitHub repository inspection started.');
const { owner, repo } = repoParts();
const res = await fetch(`https://api.github.com/repos/${owner}/${repo}`, {
headers: {
accept: 'application/vnd.github+json',
authorization: `Bearer ${env.GITHUB_TOKEN}`,
'x-github-api-version': '2022-11-28',
},
});
if (!res.ok) throw new Error(`GitHub repo check returned ${res.status}`);
const body = (await res.json()) as {
full_name?: string;
default_branch?: string;
open_issues_count?: number;
};
await appendHermesJobEvent(
jobId,
'step_output',
`GitHub ${body.full_name ?? `${owner}/${repo}`} is reachable. Default branch=${body.default_branch ?? 'unknown'}, open issue count=${body.open_issues_count ?? 0}.`,
);
await appendHermesJobEvent(jobId, 'step_completed', 'GitHub repository inspection completed.');
}
async function runHermesJob(jobId: string): Promise<void> {
const job = await getHermesJob(jobId);
if (!job) return;
const controller = new AbortController();
runners.set(jobId, controller);
try {
await setStatus(jobId, 'running', { startedAt: now() });
await appendHermesJobEvent(jobId, 'heartbeat', 'Hermes is gathering repository context.');
const outputs: string[] = [];
outputs.push(
await runCommand(
jobId,
'Git status',
'git',
['status', '--short', '--branch'],
controller.signal,
),
);
outputs.push(
await runCommand(jobId, 'Git diff summary', 'git', ['diff', '--stat'], controller.signal),
);
if (wantsGithub(job.prompt, job.contextScope)) await inspectGithub(jobId);
if (wantsMongo(job.prompt, job.contextScope)) await inspectMongo(jobId);
if (wantsBuild(job.prompt, job.successCriteria)) {
outputs.push(
await runCommand(jobId, 'TypeScript typecheck', 'pnpm', ['typecheck'], controller.signal),
);
}
if (job.riskLevel === 'deploy_allowed' && job.requiresConfirmation) {
await setStatus(jobId, 'waiting_for_confirmation');
await appendHermesJobEvent(
jobId,
'needs_confirmation',
'Hermes needs confirmation before deploy-level actions.',
);
return;
}
const finalSummary = `Hermes completed the task. It inspected repository state${wantsMongo(job.prompt, job.contextScope) ? ', MongoDB' : ''}${wantsGithub(job.prompt, job.contextScope) ? ', and GitHub' : ''}. ${outputs.some((o) => /error|failed/i.test(o)) ? 'Review the recorded output for warnings.' : 'No blocking error was reported by the completed checks.'}`;
await setStatus(jobId, 'completed', { completedAt: now(), finalSummary });
await appendHermesJobEvent(jobId, 'completed', finalSummary);
} catch (err) {
const aborted = controller.signal.aborted;
const message = aborted
? 'Hermes aborted the job before making further changes.'
: (err as Error).message;
await setStatus(jobId, aborted ? 'aborted' : 'failed', {
completedAt: now(),
finalSummary: message,
error: aborted ? undefined : message,
});
await appendHermesJobEvent(jobId, aborted ? 'aborted' : 'failed', message);
} finally {
runners.delete(jobId);
}
}
+95
View File
@@ -0,0 +1,95 @@
import { getDb } from '../memory/db.js';
import { getMemberWorkHistory } from '../activity/member-history.js';
import {
buildUserLearningProfile,
getUserLearningProfileByIdentity,
} from '../memory/user-learning.js';
const DEFAULT_LIMIT = 8;
function sinceIso(hours: number): string {
return new Date(Date.now() - hours * 60 * 60 * 1000).toISOString();
}
export async function getLiveConversationContext(podId: string, identity: string) {
const db = await getDb();
const since = sinceIso(12);
const [pod, history, gitState, collisions, interventions, outcomes, userLearningProfile] =
await Promise.all([
db.collection('pods').findOne({ id: podId }, { projection: { _id: 0 } }),
getMemberWorkHistory(podId, identity, { hours: 24, limit: 30 }).catch(() => null),
db.collection('engineer_states').findOne(
{ podId, name: identity },
{
projection: {
_id: 0,
name: 1,
branch: 1,
changedFiles: 1,
recentCommit: 1,
gitUpdatedAt: 1,
},
},
),
db
.collection('collisions')
.find({ podId, detectedAt: { $gte: since } }, { projection: { _id: 0, embedding: 0 } })
.sort({ detectedAt: -1 })
.limit(DEFAULT_LIMIT)
.toArray(),
db
.collection('interventions')
.find({ podId, createdAt: { $gte: since } }, { projection: { _id: 0 } })
.sort({ createdAt: -1 })
.limit(DEFAULT_LIMIT)
.toArray(),
db
.collection('outcomes')
.find({ podId, recordedAt: { $gte: since } }, { projection: { _id: 0 } })
.sort({ recordedAt: -1 })
.limit(DEFAULT_LIMIT)
.toArray(),
getUserLearningProfileByIdentity(identity).catch(() => null),
]);
return {
pod,
identity,
generatedAt: new Date().toISOString(),
currentGitState: gitState,
userLearningProfile,
memberHistory: history,
recentCollisions: collisions,
recentInterventions: interventions,
recentOutcomes: outcomes,
};
}
export async function recordLiveConversationNote(input: {
podId: string;
sessionId: string;
identity?: string;
note: string;
kind?: string;
}) {
const note = input.note.trim();
if (!note) throw new Error('note is required');
const doc = {
podId: input.podId,
sessionId: input.sessionId,
identity: input.identity,
kind: input.kind || 'summary',
note: note.slice(0, 4000),
createdAt: new Date().toISOString(),
};
await (await getDb()).collection('conversation_notes').insertOne(doc);
if (input.identity) {
const profile = await getUserLearningProfileByIdentity(input.identity);
if (profile) {
void buildUserLearningProfile(profile.clerkUserId).catch((err) =>
console.warn(`[memory] note user learning refresh failed: ${(err as Error).message}`),
);
}
}
return { ...doc, _id: undefined };
}
+193
View File
@@ -0,0 +1,193 @@
import { randomUUID } from 'node:crypto';
import { AccessToken, RoomAgentDispatch, RoomConfiguration } from 'livekit-server-sdk';
import { Room as LiveKitRoom } from '@livekit/rtc-node';
import type { Collision, DataMessage, Intervention, LiveConversationEvent } from '@podman/shared';
import { DATA_TOPIC } from '@podman/shared';
import { env } from '../env.js';
import { closeRoom } from '../livekit/rooms.js';
import { speakInRoom } from '../voice/live.js';
const encoder = new TextEncoder();
const DEFAULT_AGENT = 'podman-live-conversation';
export interface LiveConversationSession {
sessionId: string;
podId: string;
identity: string;
displayName: string;
room: string;
url: string;
startedAt: string;
lastEventAt?: string;
endedAt?: string;
}
const sessions = new Map<string, LiveConversationSession>();
function sessionKey(podId: string, identity: string): string {
return `${podId}:${identity.toLowerCase()}`;
}
function cleanPart(value: string): string {
return value
.trim()
.toLowerCase()
.replace(/[^a-z0-9_-]+/g, '-')
.replace(/^-+|-+$/g, '')
.slice(0, 48);
}
function agentName(): string {
return env.LIVEKIT_CONVERSATION_AGENT_NAME || DEFAULT_AGENT;
}
function tokenFor(room: string, identity: string, name: string, metadata: object): Promise<string> {
const at = new AccessToken(env.LIVEKIT_API_KEY, env.LIVEKIT_API_SECRET, {
identity,
name,
ttl: '4h',
metadata: JSON.stringify(metadata),
});
at.addGrant({ roomJoin: true, room, canPublish: true, canSubscribe: true, canPublishData: true });
return at.toJwt();
}
export async function startLiveConversation(input: {
podId: string;
identity: string;
displayName?: string;
}): Promise<LiveConversationSession & { token: string }> {
const identity = input.identity.trim();
if (!identity) throw new Error('identity is required');
const existing = activeLiveConversation(input.podId, identity);
if (existing) {
return {
...existing,
token: await tokenFor(existing.room, identity, existing.displayName, {
podId: input.podId,
identity,
sessionId: existing.sessionId,
mode: 'podman-live-conversation',
}),
};
}
const sessionId = randomUUID();
const room = `podman-live:${cleanPart(input.podId)}:${cleanPart(identity)}:${sessionId.slice(0, 8)}`;
const displayName = input.displayName?.trim() || identity;
const metadata = { podId: input.podId, identity, sessionId, mode: 'podman-live-conversation' };
const at = new AccessToken(env.LIVEKIT_API_KEY, env.LIVEKIT_API_SECRET, {
identity,
name: displayName,
ttl: '4h',
metadata: JSON.stringify(metadata),
});
at.addGrant({ roomJoin: true, room, canPublish: true, canSubscribe: true, canPublishData: true });
at.roomConfig = new RoomConfiguration({
name: room,
emptyTimeout: 60,
departureTimeout: 15,
agents: [
new RoomAgentDispatch({
agentName: agentName(),
metadata: JSON.stringify(metadata),
}),
],
});
const session: LiveConversationSession = {
sessionId,
podId: input.podId,
identity,
displayName,
room,
url: env.LIVEKIT_URL,
startedAt: new Date().toISOString(),
};
sessions.set(sessionKey(input.podId, identity), session);
return { ...session, token: await at.toJwt() };
}
export function activeLiveConversation(
podId: string,
identity: string,
): LiveConversationSession | null {
const session = sessions.get(sessionKey(podId, identity));
return session && !session.endedAt ? session : null;
}
export function listActiveLiveConversations(podId: string): LiveConversationSession[] {
return [...sessions.values()].filter((session) => session.podId === podId && !session.endedAt);
}
export async function stopLiveConversation(
podId: string,
sessionId: string,
): Promise<LiveConversationSession | null> {
const session = [...sessions.values()].find(
(candidate) => candidate.podId === podId && candidate.sessionId === sessionId,
);
if (!session) return null;
session.endedAt = new Date().toISOString();
await closeRoom(session.room);
return session;
}
async function publishPrivateConversationEvent(
roomName: string,
event: LiveConversationEvent,
): Promise<void> {
const room = new LiveKitRoom();
try {
const token = await tokenFor(roomName, `podman-live-router-${Date.now()}`, 'PodMan live router', {
mode: 'podman-live-router',
});
await room.connect(env.LIVEKIT_URL, token, { autoSubscribe: false, dynacast: false });
const data: DataMessage = { type: 'LIVE_CONVERSATION_EVENT', event };
await room.localParticipant?.publishData(encoder.encode(JSON.stringify(data)), {
reliable: true,
topic: DATA_TOPIC,
});
} finally {
await room.disconnect().catch(() => {});
}
}
export async function notifyCriticalLiveConversations(
collision: Collision,
intervention: Intervention,
voiceLine?: string,
): Promise<void> {
if (collision.severity !== 'critical') return;
const recipients = new Set(collision.engineers.map((name) => name.toLowerCase()));
const active = listActiveLiveConversations(collision.podId).filter((session) =>
recipients.has(session.identity.toLowerCase()),
);
if (active.length === 0) return;
await Promise.allSettled(
active.map(async (session) => {
const createdAt = new Date().toISOString();
session.lastEventAt = createdAt;
const summary =
voiceLine ||
`Critical collision in ${collision.file}. ${collision.engineers.join(
' and ',
)} should sync before pushing.`;
await publishPrivateConversationEvent(session.room, {
id: `live_evt_${Date.now()}_${session.sessionId.slice(0, 8)}`,
podId: collision.podId,
sessionId: session.sessionId,
kind: 'critical_collision',
severity: 'critical',
summary,
interrupt: true,
createdAt,
collisionId: collision.id,
interventionId: intervention.id,
});
await speakInRoom(session.room, summary, { priority: 'critical' });
}),
);
}
+71
View File
@@ -31,12 +31,40 @@ export async function closeMemory(): Promise<void> {
await client.close();
}
/** A durable record that PodMan stayed quiet on a recurring collision because
* its signature was previously dismissed — the negative-feedback loop made
* auditable. Timestamped at the repeat; materialized as `suppressed` activity
* in graph/live.ts. The suppressed collision is never written to `collisions`
* (the agent returns before recordCollision), so this is its own record. */
export interface SuppressionDoc {
id: string;
podId: string;
collisionId: string;
file: string;
engineers: string[];
priorInterventionId?: string;
priorDismissedAt?: string;
suppressedAt: string;
}
export interface PodCollections {
pods: Collection<Pod>;
observations: Collection<EngineerContext>;
collisions: Collection<Collision>;
interventions: Collection<Intervention>;
outcomes: Collection<InterventionOutcome>;
suppressions: Collection<SuppressionDoc>;
}
export interface UserPodContextDoc {
id: string;
clerkUserId: string;
podId: string;
memberName?: string;
action: string;
source: 'clerk';
observedAt: string;
metadata?: Record<string, string | number | boolean | null>;
}
export async function collections(): Promise<PodCollections> {
@@ -47,6 +75,7 @@ export async function collections(): Promise<PodCollections> {
collisions: db.collection<Collision>('collisions'),
interventions: db.collection<Intervention>('interventions'),
outcomes: db.collection<InterventionOutcome>('outcomes'),
suppressions: db.collection<SuppressionDoc>('suppressions'),
};
}
@@ -57,6 +86,8 @@ export interface GitState {
gitUpdatedAt: Date | null;
}
const GIT_STATE_TTL_MS = Number(process.env.GIT_STATE_TTL_MS ?? '120000');
/** Fetch latest git state per engineer for a pod from the engineer_states collection.
* Returns a map keyed by engineer name (matches --name arg used in podman-agent.mjs). */
export async function getGitStates(podId: string): Promise<Map<string, GitState>> {
@@ -71,7 +102,17 @@ export async function getGitStates(podId: string): Promise<Map<string, GitState>
}>('engineer_states');
const docs = await col.find({ podId }).toArray();
const map = new Map<string, GitState>();
const now = Date.now();
for (const doc of docs) {
const updatedAt = doc.gitUpdatedAt ? new Date(doc.gitUpdatedAt) : null;
if (
updatedAt &&
!Number.isNaN(updatedAt.getTime()) &&
GIT_STATE_TTL_MS > 0 &&
now - updatedAt.getTime() > GIT_STATE_TTL_MS
) {
continue;
}
map.set(doc.name, {
changedFiles: doc.changedFiles ?? [],
branch: doc.branch ?? null,
@@ -102,6 +143,36 @@ export async function initMemory(): Promise<void> {
['collisions.file', () => c.collisions.createIndex({ podId: 1, file: 1, detectedAt: -1 })],
['interventions.collisionId', () => c.interventions.createIndex({ collisionId: 1 })],
['outcomes.interventionId', () => c.outcomes.createIndex({ interventionId: 1 })],
['suppressions.podId', () => c.suppressions.createIndex({ podId: 1, suppressedAt: -1 })],
['hermes_jobs.id', () => db.collection('hermes_jobs').createIndex({ id: 1 }, { unique: true })],
[
'hermes_jobs.session',
() => db.collection('hermes_jobs').createIndex({ sessionId: 1, status: 1, updatedAt: -1 }),
],
[
'hermes_job_events.job',
() => db.collection('hermes_job_events').createIndex({ jobId: 1, createdAt: 1 }),
],
[
'user_pod_context.user',
() => db.collection('user_pod_context').createIndex({ clerkUserId: 1, observedAt: -1 }),
],
[
'user_pod_context.pod',
() => db.collection('user_pod_context').createIndex({ podId: 1, observedAt: -1 }),
],
[
'user_learning_profiles.user',
() => db.collection('user_learning_profiles').createIndex({ clerkUserId: 1 }, { unique: true }),
],
[
'user_learning_profiles.updated',
() => db.collection('user_learning_profiles').createIndex({ updatedAt: -1 }),
],
[
'conversation_notes.identity',
() => db.collection('conversation_notes').createIndex({ identity: 1, createdAt: -1 }),
],
];
for (const [name, make] of indexes) {
try {
+16 -21
View File
@@ -1,27 +1,22 @@
import type { Collision, SuggestedActionKind } from '@podman/shared';
import type { RecalledCollision } from './vectors.js';
const lastNudgeByPod = new Map<string, number>();
function cooldownMs(): number {
return Number(process.env.NUDGE_COOLDOWN_MS ?? '180000');
}
/** Policy gate: combines severity, exact recall outcomes, and per-pod cooldown. */
export function shouldIntervene(collision: Collision, prior: RecalledCollision | null): boolean {
if (collision.severity === 'info') return false;
const priorOutcome = prior?.priorOutcome;
if (priorOutcome && !priorOutcome.accepted && !priorOutcome.wasRealCollision) return false;
const cooldown = cooldownMs();
const last = lastNudgeByPod.get(collision.podId) ?? 0;
if (cooldown > 0 && Date.now() - last < cooldown && collision.severity !== 'critical') {
return false;
}
lastNudgeByPod.set(collision.podId, Date.now());
return true;
/**
* Policy gate — suppression fully disabled for the demo.
*
* Every real collision surfaces an intervention. Accept/Dismiss are still
* recorded as outcomes but carry NO gating meaning: dismissing one collision
* never silences future ones, and there is no per-pod cooldown throttling
* consecutive conflicts. The only filter is severity: `info` collisions are
* informational, not actionable, so they don't nudge. The same-collision
* single-shot dedupe lives in `PodmanAgent.handle` (activeConflicts), so
* removing the cooldown does not cause repeat spam of an unresolved collision.
*
* (Prior dismissal-based suppression over-generalized: one README discard
* permanently muted all README clashes via loose file/vector recall.)
*/
export function shouldIntervene(collision: Collision, _prior: RecalledCollision | null): boolean {
return collision.severity !== 'info';
}
/** Preferred action selection based on collision severity and prior accepted actions. */
+152 -6
View File
@@ -5,8 +5,20 @@ import type {
InterventionOutcome,
InterventionStatus,
} from '@podman/shared';
import { collections } from './db.js';
import { collections, getDb, getGitStates, type UserPodContextDoc } from './db.js';
import { enrichCollisionMemory } from './vectors.js';
import { buildUserLearningProfile } from './user-learning.js';
function comparableFile(raw?: string): string {
return (
(raw ?? '')
.trim()
.replace(/^(\?\?|[MADRCU!]{1,2})\s+/, '')
.split(/[\\/]/)
.pop()
?.toLowerCase() ?? ''
);
}
/**
* Continual-learning memory: persist observations, collisions, interventions,
@@ -29,6 +41,30 @@ export async function recordObservation(ctx: EngineerContext): Promise<void> {
);
}
export async function recordUserPodContext(input: {
clerkUserId: string;
podId: string;
memberName?: string;
action: string;
metadata?: UserPodContextDoc['metadata'];
}): Promise<void> {
await persist('user pod context', async () =>
(await getDb()).collection<UserPodContextDoc>('user_pod_context').insertOne({
id: `upc_${Date.now()}_${Math.random().toString(36).slice(2, 8)}`,
clerkUserId: input.clerkUserId,
podId: input.podId,
memberName: input.memberName,
action: input.action,
source: 'clerk',
observedAt: new Date().toISOString(),
metadata: input.metadata,
}),
);
void buildUserLearningProfile(input.clerkUserId).catch((err) =>
console.warn(`[memory] user learning refresh failed: ${(err as Error).message}`),
);
}
export async function recordCollision(collision: Collision): Promise<void> {
await persist('collision', async () =>
(await collections()).collisions.insertOne(await enrichCollisionMemory(collision)),
@@ -41,6 +77,56 @@ export async function recordIntervention(intervention: Intervention): Promise<vo
);
}
/**
* Feature A — record that PodMan stayed quiet on a recurring collision because
* its signature was previously dismissed. Timestamped at the repeat (now), so
* the negative-feedback beat surfaces as recent `suppressed` activity rather
* than being re-synthesized from the old dismissal row.
*/
export async function recordSuppression(
collision: Collision,
priorInterventionId?: string,
priorDismissedAt?: string,
): Promise<void> {
await persist('suppression', async () =>
(await collections()).suppressions.insertOne({
id: `supp_${Date.now()}`,
podId: collision.podId,
collisionId: collision.id,
file: collision.file,
engineers: collision.engineers,
priorInterventionId,
priorDismissedAt,
suppressedAt: new Date().toISOString(),
}),
);
}
export async function hasRecentInterventionForCollision(
collision: Collision,
windowMs = Number(process.env.NUDGE_COOLDOWN_MS ?? '180000'),
): Promise<boolean> {
if (windowMs <= 0) return false;
const c = await collections();
const since = new Date(Date.now() - windowMs).toISOString();
const recent = await c.collisions
.find({ podId: collision.podId, detectedAt: { $gte: since } })
.sort({ detectedAt: -1 })
.limit(100)
.toArray();
const targetFile = comparableFile(collision.file);
for (const match of recent) {
if (match.id === collision.id || comparableFile(match.file) !== targetFile) continue;
const existing = await c.interventions.findOne({
collisionId: match.id,
createdAt: { $gte: since },
});
if (existing) return true;
}
return false;
}
export async function updateInterventionStatus(
interventionId: string,
status: InterventionStatus,
@@ -50,13 +136,53 @@ export async function updateInterventionStatus(
);
}
/**
* Step 3 — derive whether a flagged collision was REAL from git ground truth,
* instead of trusting the client (which historically hardcoded `true`). A
* collision counts as real only if BOTH named engineers currently have the
* collided file in their git `changedFiles`. Conservative: returns false when
* the collision is orphaned/missing or git state is stale/unavailable.
* Verifier supervision per docs/continual-learning/spec.md:98-108, policy.md:35-42.
*/
export async function deriveWasRealCollision(outcome: InterventionOutcome): Promise<boolean> {
try {
const c = await collections();
const collision = await c.collisions.findOne({ id: outcome.collisionId });
if (!collision) return false;
// Prefer the overlap evidence captured at detection time (fresh git state):
// immune to late clicks, stale sidecars, and the engineer_states TTL.
if (typeof collision.gitOverlap === 'boolean') return collision.gitOverlap;
// Fallback for collisions detected before gitOverlap was captured: re-derive
// from latest git state, matching engineers on case/whitespace-canonical names.
if (!Array.isArray(collision.engineers) || collision.engineers.length < 2) return false;
const target = comparableFile(collision.file);
if (!target) return false;
const byCanon = new Map<string, string[]>();
for (const [name, st] of await getGitStates(outcome.podId)) {
byCanon.set(name.trim().toLowerCase(), st.changedFiles);
}
return collision.engineers.every((e) =>
(byCanon.get(e.trim().toLowerCase()) ?? []).some((f) => comparableFile(f) === target),
);
} catch (err) {
console.error(`[memory] wasRealCollision verifier failed: ${(err as Error).message}`);
return false;
}
}
export async function recordOutcome(outcome: InterventionOutcome): Promise<void> {
// Backend is authoritative for wasRealCollision: derive it from git overlap
// rather than trusting the client-supplied value. (RSI Step 3)
const verified: InterventionOutcome = {
...outcome,
wasRealCollision: await deriveWasRealCollision(outcome),
};
await persist('outcome', async () => {
const c = await collections();
await c.outcomes.insertOne({ ...outcome });
await c.outcomes.insertOne({ ...verified });
await c.interventions.updateOne(
{ id: outcome.interventionId },
{ $set: { status: outcome.accepted ? 'accepted' : 'dismissed' } },
{ id: verified.interventionId },
{ $set: { status: verified.accepted ? 'accepted' : 'dismissed' } },
);
});
}
@@ -64,11 +190,31 @@ export async function recordOutcome(outcome: InterventionOutcome): Promise<void>
/** Document counts per collection — used by the /api/memory/stats endpoint. */
export async function memoryStats(): Promise<Record<string, number>> {
const c = await collections();
const [observations, collisions, interventions, outcomes] = await Promise.all([
const db = await getDb();
const [
observations,
collisions,
interventions,
outcomes,
suppressions,
userPodContext,
userLearningProfiles,
] = await Promise.all([
c.observations.estimatedDocumentCount(),
c.collisions.estimatedDocumentCount(),
c.interventions.estimatedDocumentCount(),
c.outcomes.estimatedDocumentCount(),
c.suppressions.estimatedDocumentCount(),
db.collection<UserPodContextDoc>('user_pod_context').estimatedDocumentCount(),
db.collection('user_learning_profiles').estimatedDocumentCount(),
]);
return { observations, collisions, interventions, outcomes };
return {
observations,
collisions,
interventions,
outcomes,
suppressions,
userPodContext,
userLearningProfiles,
};
}
+282
View File
@@ -0,0 +1,282 @@
import type { Collision, EngineerContext, HermesJob, Pod, UserLearningProfile } from '@podman/shared';
import { getDb, type UserPodContextDoc } from './db.js';
interface EngineerStateDoc {
podId: string;
name: string;
changedFiles?: string[];
branch?: string | null;
recentCommit?: string | null;
gitUpdatedAt?: Date | string;
}
interface ConversationNoteDoc {
podId: string;
identity?: string;
kind?: string;
note?: string;
createdAt?: string;
}
function clean(value: unknown): string | undefined {
return typeof value === 'string' && value.trim() ? value.trim() : undefined;
}
function toIso(value: unknown): string {
if (value instanceof Date) return value.toISOString();
if (typeof value === 'string' && !Number.isNaN(Date.parse(value))) return new Date(value).toISOString();
return new Date().toISOString();
}
function uniq(values: Array<string | undefined>): string[] {
const seen = new Set<string>();
const out: string[] = [];
for (const value of values) {
const item = value?.trim();
if (!item) continue;
const key = item.toLowerCase();
if (seen.has(key)) continue;
seen.add(key);
out.push(item);
}
return out;
}
function top(values: string[], limit: number): string[] {
const counts = new Map<string, number>();
for (const value of values) {
const item = value.trim();
if (!item) continue;
counts.set(item, (counts.get(item) ?? 0) + 1);
}
return [...counts.entries()]
.sort((a, b) => b[1] - a[1] || a[0].localeCompare(b[0]))
.slice(0, limit)
.map(([value]) => value);
}
function lowerRegex(values: string[]): RegExp[] {
return values.map((value) => new RegExp(`^${value.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')}$`, 'i'));
}
function noteGoals(notes: ConversationNoteDoc[]): string[] {
const goalNotes = notes
.filter((note) => /goal|plan|todo|next|preference|prefers|likes|wants|needs/i.test(`${note.kind} ${note.note}`))
.map((note) => clean(note.note)?.slice(0, 180))
.filter(Boolean) as string[];
return uniq(goalNotes).slice(0, 8);
}
function inferCollaborationStyle(input: {
contexts: UserPodContextDoc[];
observations: EngineerContext[];
collisions: Collision[];
outcomes: Array<{ accepted?: boolean; wasRealCollision?: boolean }>;
notes: ConversationNoteDoc[];
jobs: HermesJob[];
}): string[] {
const style: string[] = [];
const actions = input.contexts.map((ctx) => ctx.action);
const pods = new Set(input.contexts.map((ctx) => ctx.podId));
if (pods.size > 1) style.push(`Works across ${pods.size} pods and carries context between rooms.`);
if (actions.filter((action) => action === 'joined_pod').length >= 3)
style.push('Frequently joins live rooms and collaborates synchronously.');
if (input.collisions.length > 0)
style.push(`Has been involved in ${input.collisions.length} coordination risk signals.`);
const accepted = input.outcomes.filter((outcome) => outcome.accepted).length;
if (accepted > 0) style.push(`Accepts Hermes help when useful (${accepted} accepted outcomes).`);
if (input.jobs.length > 0) style.push(`Delegates complex work to Hermes (${input.jobs.length} jobs).`);
if (input.notes.some((note) => /concise|short|brief/i.test(note.note ?? '')))
style.push('Prefers concise coordination.');
return style.slice(0, 8);
}
function inferWorkingStyle(input: {
observations: EngineerContext[];
gitStates: EngineerStateDoc[];
}): string[] {
const style: string[] = [];
const modes = top(input.observations.map((obs) => obs.mode ?? '').filter(Boolean), 2);
const activities = top(input.observations.map((obs) => obs.activity ?? '').filter(Boolean), 5);
const files = top(
[
...input.observations.map((obs) => obs.currentFile ?? ''),
...input.gitStates.flatMap((state) => state.changedFiles ?? []),
].filter(Boolean),
6,
);
if (modes.length) style.push(`Usually seen in ${modes.join(' and ')} mode.`);
if (activities.length) style.push(`Common work patterns: ${activities.join(', ')}.`);
if (files.length) style.push(`Frequently touches ${files.join(', ')}.`);
return style.slice(0, 8);
}
function inferKnowledge(input: {
observations: EngineerContext[];
gitStates: EngineerStateDoc[];
notes: ConversationNoteDoc[];
}): string[] {
const topics = top(
[
...input.observations.map((obs) => obs.researchTopic ?? ''),
...input.observations.map((obs) => obs.researchSource ?? ''),
...input.observations.map((obs) => obs.currentSymbol ?? ''),
...input.gitStates.flatMap((state) => state.changedFiles ?? []),
...input.notes
.filter((note) => /learn|knows|expert|worked on|decision/i.test(note.note ?? ''))
.map((note) => note.note?.slice(0, 120) ?? ''),
].filter(Boolean),
10,
);
return topics;
}
export async function buildUserLearningProfile(clerkUserId: string): Promise<UserLearningProfile | null> {
const db = await getDb();
const contexts = await db
.collection<UserPodContextDoc>('user_pod_context')
.find({ clerkUserId }, { projection: { _id: 0 } })
.sort({ observedAt: -1 })
.limit(500)
.toArray();
if (!contexts.length) return null;
const latest = contexts[0];
const identities = uniq([
...contexts.map((ctx) => ctx.memberName),
...contexts.map((ctx) => clean(ctx.metadata?.identity)),
...contexts.map((ctx) => clean(ctx.metadata?.email)),
]);
const email = clean(latest?.metadata?.email) ?? identities.find((item) => item.includes('@'));
const displayName = latest?.memberName ?? identities.find((item) => !item.includes('@')) ?? email;
const imageUrl = clean(latest?.metadata?.imageUrl);
const identityRegex = lowerRegex(identities);
const podIds = uniq(contexts.map((ctx) => ctx.podId));
const [pods, observations, gitStates, collisions, outcomes, notes, jobs] = await Promise.all([
db
.collection<Pod>('pods')
.find({ id: { $in: podIds } }, { projection: { _id: 0 } })
.toArray(),
db
.collection<EngineerContext>('observations')
.find({ engineerId: { $in: identities } }, { projection: { _id: 0, screenshotDataUrl: 0 } })
.sort({ observedAt: -1 })
.limit(500)
.toArray(),
db
.collection<EngineerStateDoc>('engineer_states')
.find({ name: { $in: identityRegex } }, { projection: { _id: 0 } })
.limit(100)
.toArray(),
db
.collection<Collision>('collisions')
.find({ engineers: { $in: identities } }, { projection: { _id: 0, embedding: 0 } })
.sort({ detectedAt: -1 })
.limit(100)
.toArray(),
db
.collection<{ podId: string; accepted?: boolean; wasRealCollision?: boolean }>('outcomes')
.find({ podId: { $in: podIds } }, { projection: { _id: 0 } })
.sort({ recordedAt: -1 })
.limit(100)
.toArray(),
db
.collection<ConversationNoteDoc>('conversation_notes')
.find({ identity: { $in: identities } }, { projection: { _id: 0 } })
.sort({ createdAt: -1 })
.limit(100)
.toArray(),
db
.collection<HermesJob>('hermes_jobs')
.find({ identity: { $in: identities } }, { projection: { _id: 0 } })
.sort({ createdAt: -1 })
.limit(100)
.toArray(),
]);
const podById = new Map(pods.map((pod) => [pod.id, pod]));
const podsSummary = podIds.map((podId) => {
const rows = contexts.filter((ctx) => ctx.podId === podId);
const sorted = [...rows].sort((a, b) => Date.parse(a.observedAt) - Date.parse(b.observedAt));
return {
podId,
podName: podById.get(podId)?.name,
visits: rows.filter((row) => row.action === 'joined_pod').length,
actions: top(rows.map((row) => row.action), 6),
firstSeenAt: toIso(sorted[0]?.observedAt),
lastSeenAt: toIso(sorted.at(-1)?.observedAt),
};
});
const profile: UserLearningProfile = {
clerkUserId,
displayName,
email,
imageUrl,
identities,
pods: podsSummary,
recentWork: observations.slice(0, 20).map((obs) => ({
podId: obs.podId,
file: obs.currentFile,
activity: obs.activity ?? obs.researchTopic,
at: obs.observedAt,
})),
collaborationStyle: inferCollaborationStyle({ contexts, observations, collisions, outcomes, notes, jobs }),
workingStyle: inferWorkingStyle({ observations, gitStates }),
goals: noteGoals(notes),
knowledge: inferKnowledge({ observations, gitStates, notes }),
counts: {
podActions: contexts.length,
observations: observations.length,
gitStates: gitStates.length,
collisionsInvolved: collisions.length,
outcomes: outcomes.length,
conversationNotes: notes.length,
hermesJobs: jobs.length,
},
updatedAt: new Date().toISOString(),
};
await db
.collection<UserLearningProfile>('user_learning_profiles')
.updateOne({ clerkUserId }, { $set: profile }, { upsert: true });
return profile;
}
export async function refreshUserLearningProfiles(): Promise<UserLearningProfile[]> {
const db = await getDb();
const ids = await db.collection<UserPodContextDoc>('user_pod_context').distinct('clerkUserId');
const profiles = await Promise.all(ids.map((id) => buildUserLearningProfile(String(id))));
return profiles.filter(Boolean) as UserLearningProfile[];
}
export async function getUserLearningProfileByIdentity(identity: string): Promise<UserLearningProfile | null> {
const db = await getDb();
const direct = await db
.collection<UserLearningProfile>('user_learning_profiles')
.findOne({ identities: { $regex: `^${identity.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')}$`, $options: 'i' } }, { projection: { _id: 0 } });
if (direct) return direct;
const context = await db
.collection<UserPodContextDoc>('user_pod_context')
.findOne(
{
$or: [
{ memberName: { $regex: `^${identity.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')}$`, $options: 'i' } },
{ 'metadata.identity': { $regex: `^${identity.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')}$`, $options: 'i' } },
],
},
{ sort: { observedAt: -1 }, projection: { _id: 0 } },
);
return context ? buildUserLearningProfile(context.clerkUserId) : null;
}
export async function listUserLearningProfiles(): Promise<UserLearningProfile[]> {
const db = await getDb();
return db
.collection<UserLearningProfile>('user_learning_profiles')
.find({}, { projection: { _id: 0 } })
.sort({ updatedAt: -1 })
.toArray();
}
+7 -1
View File
@@ -49,7 +49,7 @@ function memoryText(collision: Collision): string {
.join('\n');
}
function cosine(a: number[], b: number[]): number {
export function cosine(a: number[], b: number[]): number {
const n = Math.min(a.length, b.length);
let dot = 0;
let aNorm = 0;
@@ -69,6 +69,12 @@ async function embed(text: string, inputType: 'document' | 'query'): Promise<num
return (await embedWithVoyage(text, inputType)) ?? embedWithGemini(text, inputType);
}
export async function semanticSimilarity(a: string, b: string): Promise<number | null> {
const [va, vb] = await Promise.all([embed(a, 'query'), embed(b, 'document')]);
if (!va || !vb) return null;
return cosine(va, vb);
}
async function embedWithVoyage(
text: string,
inputType: 'document' | 'query',
+59 -6
View File
@@ -1,5 +1,5 @@
import type { Pod, PodInput } from '@podman/shared';
import { collections } from '../memory/db.js';
import { collections, getDb, type UserPodContextDoc } from '../memory/db.js';
const NO_ID = { projection: { _id: 0 } } as const;
@@ -59,14 +59,67 @@ function now(): string {
return new Date().toISOString();
}
function regexExact(value: string): RegExp {
return new RegExp(`^${value.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')}$`, 'i');
}
function profileString(value: unknown): string | undefined {
return typeof value === 'string' && value.trim() ? value.trim() : undefined;
}
function matchesMember(doc: UserPodContextDoc, member: string): boolean {
const key = member.trim().toLowerCase();
return [doc.memberName, doc.metadata?.identity, doc.metadata?.email]
.filter(Boolean)
.some((value) => String(value).trim().toLowerCase() === key);
}
async function hydrateMemberProfiles(pod: Pod): Promise<Pod> {
if (!pod.members.length) return pod;
const db = await getDb();
const matchers = pod.members.map(regexExact);
const docs = await db
.collection<UserPodContextDoc>('user_pod_context')
.find(
{
$or: [
{ memberName: { $in: matchers } },
{ 'metadata.identity': { $in: matchers } },
{ 'metadata.email': { $in: matchers } },
],
},
{ projection: { _id: 0 } },
)
.sort({ observedAt: -1 })
.limit(500)
.toArray();
const memberProfiles: NonNullable<Pod['memberProfiles']> = {};
for (const member of pod.members) {
const doc = docs.find((row) => matchesMember(row, member));
const imageUrl = profileString(doc?.metadata?.imageUrl);
if (!doc || !imageUrl) continue;
memberProfiles[member] = {
displayName: doc.memberName ?? member,
email: profileString(doc.metadata?.email),
imageUrl,
};
}
return Object.keys(memberProfiles).length ? { ...pod, memberProfiles } : pod;
}
async function hydratePods(pods: Pod[]): Promise<Pod[]> {
return Promise.all(pods.map((pod) => hydrateMemberProfiles(pod)));
}
export async function listPods(): Promise<Pod[]> {
const c = await collections();
return c.pods.find({}, NO_ID).sort({ createdAt: 1 }).toArray();
return hydratePods(await c.pods.find({}, NO_ID).sort({ createdAt: 1 }).toArray());
}
export async function getPod(id: string): Promise<Pod | null> {
const c = await collections();
return c.pods.findOne({ id }, NO_ID);
const pod = await c.pods.findOne({ id }, NO_ID);
return pod ? hydrateMemberProfiles(pod) : null;
}
export async function createPod(input: PodInput): Promise<Pod> {
@@ -92,7 +145,7 @@ export async function createPod(input: PodInput): Promise<Pod> {
};
try {
await c.pods.insertOne({ ...pod });
return pod;
return hydrateMemberProfiles(pod);
} catch (err) {
if (isDuplicateKey(err)) continue;
throw err;
@@ -119,7 +172,7 @@ export async function updatePod(id: string, patch: PodInput): Promise<Pod | null
{ $set: set },
{ returnDocument: 'after', projection: { _id: 0 } },
);
return updated ?? null;
return updated ? hydrateMemberProfiles(updated) : null;
}
export async function deletePod(id: string): Promise<boolean> {
@@ -135,7 +188,7 @@ async function setMembers(id: string, members: string[]): Promise<Pod | null> {
{ $set: { members, updatedAt: now() } },
{ returnDocument: 'after', projection: { _id: 0 } },
);
return updated ?? null;
return updated ? hydrateMemberProfiles(updated) : null;
}
export async function addMember(id: string, rawName: unknown): Promise<Pod | null> {
+514 -8
View File
@@ -1,11 +1,20 @@
import express from 'express';
import cors from 'cors';
import { createServer } from 'node:http';
import type { Socket } from 'node:net';
import { WebSocketServer } from 'ws';
import { AccessToken, RoomConfiguration } from 'livekit-server-sdk';
import { AccessToken, RoomAgentDispatch, RoomConfiguration } from 'livekit-server-sdk';
import { env } from './env.js';
import { createSyncPr } from './github/client.js';
import { recordOutcome, memoryStats } from './memory/store.js';
import {
recordCollision,
recordIntervention,
recordOutcome,
hasRecentInterventionForCollision,
memoryStats,
recordUserPodContext,
} from './memory/store.js';
import { clerkAuthMiddleware, requestUser } from './auth.js';
import { closeMemory, initMemory } from './memory/db.js';
import {
listPods,
@@ -19,27 +28,120 @@ import {
} from './pods/store.js';
import { getPresence, closeRoom } from './livekit/rooms.js';
import { loadPodGraph, reachFrom } from './graph/store.js';
import type { InterventionOutcome } from '@podman/shared';
import { listPodActivity } from './activity/store.js';
import { getMemberWorkHistory } from './activity/member-history.js';
import { speakInRoom } from './voice/live.js';
import { getPodMusic } from './voice/music.js';
import { notifyHermesInterventionInRoom } from './action/hermes.js';
import {
activeLiveConversation,
startLiveConversation,
stopLiveConversation,
} from './live-conversation/sessions.js';
import {
getLiveConversationContext,
recordLiveConversationNote,
} from './live-conversation/context.js';
import {
listUserLearningProfiles,
refreshUserLearningProfiles,
} from './memory/user-learning.js';
import {
abortHermesJob,
appendHermesJobEvent,
createHermesJob,
getActiveHermesJobForSession,
getHermesJob,
getLatestHermesJobForSession,
listHermesJobEvents,
} from './hermes/jobs.js';
import type {
Collision,
HermesJobEventType,
Intervention,
InterventionOutcome,
SuggestedActionKind,
} from '@podman/shared';
const app = express();
app.use(cors());
app.use(express.json());
app.use(clerkAuthMiddleware);
app.get('/health', (_req, res) => res.json({ ok: true }));
function stringArray(value: unknown): string[] {
return Array.isArray(value) ? value.map((item) => String(item).trim()).filter(Boolean) : [];
}
function suggestedAction(value: unknown): SuggestedActionKind {
return value === 'open_sync_pr' || value === 'ping_teammate' || value === 'none'
? value
: 'ping_teammate';
}
function stringMeta(value: unknown): string | undefined {
return typeof value === 'string' && value.trim() ? value.trim() : undefined;
}
function hermesJobEventType(value: unknown): HermesJobEventType | null {
return value === 'accepted' ||
value === 'heartbeat' ||
value === 'step_started' ||
value === 'step_output' ||
value === 'needs_confirmation' ||
value === 'step_completed' ||
value === 'aborted' ||
value === 'failed' ||
value === 'completed'
? value
: null;
}
// Mint a LiveKit token for an engineer joining a pod.
app.post('/api/token', async (req, res) => {
const { room, identity, name, githubLogin } = req.body ?? {};
if (!room || !identity) return res.status(400).json({ error: 'room+identity required' });
const user = requestUser(req);
const at = new AccessToken(env.LIVEKIT_API_KEY, env.LIVEKIT_API_SECRET, {
identity,
name,
ttl: '4h',
metadata: JSON.stringify({ githubLogin: githubLogin ?? name }),
metadata: JSON.stringify({
githubLogin: githubLogin ?? name,
email: stringMeta(req.body?.profile?.email),
imageUrl: stringMeta(req.body?.profile?.imageUrl),
}),
});
at.addGrant({ roomJoin: true, room, canPublish: true, canSubscribe: true, canPublishData: true });
const agents = env.LIVEKIT_AGENT_NAME
? [
new RoomAgentDispatch({
agentName: env.LIVEKIT_AGENT_NAME,
metadata: JSON.stringify({ podId: room }),
}),
]
: undefined;
// Auto-clean the room: close 60s after it empties, drop a participant 20s
// after they disconnect. Applied when LiveKit auto-creates the room.
at.roomConfig = new RoomConfiguration({ name: room, emptyTimeout: 60, departureTimeout: 20 });
at.roomConfig = new RoomConfiguration({
name: room,
emptyTimeout: 60,
departureTimeout: 20,
agents,
});
if (user) {
await recordUserPodContext({
clerkUserId: user.clerkUserId,
podId: String(room),
memberName: typeof name === 'string' ? name : String(identity),
action: 'joined_pod',
metadata: {
identity: String(identity),
email: stringMeta(req.body?.profile?.email) ?? null,
imageUrl: stringMeta(req.body?.profile?.imageUrl) ?? null,
},
});
}
res.json({ token: await at.toJwt(), url: env.LIVEKIT_URL });
});
@@ -77,6 +179,22 @@ app.get('/api/memory/stats', async (_req, res) => {
}
});
app.get('/api/memory/users', async (_req, res) => {
try {
res.json(await listUserLearningProfiles());
} catch (e) {
res.status(500).json({ error: (e as Error).message });
}
});
app.post('/api/memory/users/refresh', async (_req, res) => {
try {
res.json({ profiles: await refreshUserLearningProfiles() });
} catch (e) {
res.status(500).json({ error: (e as Error).message });
}
});
// Live presence: who is currently connected in each pod's LiveKit room.
app.get('/api/presence', async (_req, res) => {
try {
@@ -97,7 +215,24 @@ app.get('/api/pods', async (_req, res) => {
app.post('/api/pods', async (req, res) => {
try {
res.status(201).json(await createPod(req.body ?? {}));
const pod = await createPod(req.body ?? {});
const user = requestUser(req);
if (user) {
await recordUserPodContext({
clerkUserId: user.clerkUserId,
podId: pod.id,
memberName: stringMeta(req.body?.profile?.displayName),
action: 'created_pod',
metadata: {
podName: pod.name,
repo: pod.repo,
identity: stringMeta(req.body?.profile?.displayName) ?? null,
email: stringMeta(req.body?.profile?.email) ?? null,
imageUrl: stringMeta(req.body?.profile?.imageUrl) ?? null,
},
});
}
res.status(201).json(await getPod(pod.id));
} catch (e) {
res.status(400).json({ error: (e as Error).message });
}
@@ -113,6 +248,15 @@ app.patch('/api/pods/:id', async (req, res) => {
try {
const pod = await updatePod(req.params.id, req.body ?? {});
if (!pod) return res.status(404).json({ error: 'pod not found' });
const user = requestUser(req);
if (user) {
await recordUserPodContext({
clerkUserId: user.clerkUserId,
podId: pod.id,
action: 'updated_pod',
metadata: { podName: pod.name, repo: pod.repo },
});
}
res.json(pod);
} catch (e) {
res.status(400).json({ error: (e as Error).message });
@@ -130,18 +274,325 @@ app.post('/api/pods/:id/members', async (req, res) => {
try {
const pod = await addMember(req.params.id, req.body?.name ?? '');
if (!pod) return res.status(404).json({ error: 'pod not found' });
res.json(pod);
const user = requestUser(req);
if (user) {
await recordUserPodContext({
clerkUserId: user.clerkUserId,
podId: pod.id,
memberName: typeof req.body?.name === 'string' ? req.body.name.trim() : undefined,
action: 'added_member',
metadata: {
identity: typeof req.body?.name === 'string' ? req.body.name.trim() : null,
email: stringMeta(req.body?.profile?.email) ?? null,
imageUrl: stringMeta(req.body?.profile?.imageUrl) ?? null,
},
});
}
res.json(await getPod(pod.id));
} catch (e) {
res.status(400).json({ error: (e as Error).message });
}
});
app.post('/api/pods/:id/voice-test', async (req, res) => {
const podId = req.params.id;
const message =
typeof req.body?.message === 'string' && req.body.message.trim()
? req.body.message.trim()
: 'PodMan voice test. Gemini TTS is playing through LiveKit.';
try {
await speakInRoom(podId, message);
res.json({ ok: true });
} catch (e) {
res.status(500).json({ error: (e as Error).message });
}
});
app.post('/api/pods/:id/live-conversation/start', async (req, res) => {
try {
const identity = typeof req.body?.identity === 'string' ? req.body.identity.trim() : '';
const displayName =
typeof req.body?.displayName === 'string' ? req.body.displayName.trim() : identity;
if (!identity) return res.status(400).json({ error: 'identity is required' });
const pod = await getPod(req.params.id);
if (!pod) return res.status(404).json({ error: 'pod not found' });
res.json(await startLiveConversation({ podId: req.params.id, identity, displayName }));
} catch (e) {
res.status(500).json({ error: (e as Error).message });
}
});
app.post('/api/pods/:id/live-conversation/:sessionId/stop', async (req, res) => {
try {
const session = await stopLiveConversation(req.params.id, req.params.sessionId);
if (!session) return res.status(404).json({ error: 'session not found' });
res.json({ ok: true, session });
} catch (e) {
res.status(500).json({ error: (e as Error).message });
}
});
app.get('/api/pods/:id/live-conversation/status', (req, res) => {
const identity = typeof req.query.identity === 'string' ? req.query.identity.trim() : '';
if (!identity) return res.status(400).json({ error: 'identity is required' });
res.json({ active: activeLiveConversation(req.params.id, identity) });
});
app.get('/api/pods/:id/live-conversation/:sessionId/hermes-job', async (req, res) => {
try {
const job = await getLatestHermesJobForSession(req.params.sessionId);
if (!job || job.podId !== req.params.id) return res.json({ job: null, events: [] });
res.json({ job, events: await listHermesJobEvents(job.id, 12) });
} catch (e) {
res.status(500).json({ error: (e as Error).message });
}
});
app.post('/api/pods/:id/live-conversation/:sessionId/hermes-job/abort', async (req, res) => {
try {
const job = await getActiveHermesJobForSession(req.params.sessionId);
if (!job || job.podId !== req.params.id)
return res.status(404).json({ error: 'active job not found' });
res.json({ job: await abortHermesJob(job.id) });
} catch (e) {
res.status(500).json({ error: (e as Error).message });
}
});
function requireInternalAgent(req: express.Request, res: express.Response): boolean {
const expected = env.INTERNAL_AGENT_TOKEN;
if (!expected) {
res.status(503).json({ error: 'INTERNAL_AGENT_TOKEN is not configured' });
return false;
}
const header = req.header('authorization') ?? '';
const actual = header.startsWith('Bearer ') ? header.slice('Bearer '.length) : '';
if (actual !== expected) {
res.status(401).json({ error: 'unauthorized' });
return false;
}
return true;
}
app.get('/api/internal/pods/:id/live-context', async (req, res) => {
if (!requireInternalAgent(req, res)) return;
try {
const identity = typeof req.query.identity === 'string' ? req.query.identity.trim() : '';
if (!identity) return res.status(400).json({ error: 'identity is required' });
res.json(await getLiveConversationContext(req.params.id, identity));
} catch (e) {
res.status(500).json({ error: (e as Error).message });
}
});
app.post('/api/internal/pods/:id/live-conversation/:sessionId/note', async (req, res) => {
if (!requireInternalAgent(req, res)) return;
try {
const note = typeof req.body?.note === 'string' ? req.body.note : '';
const identity = typeof req.body?.identity === 'string' ? req.body.identity : undefined;
const kind = typeof req.body?.kind === 'string' ? req.body.kind : undefined;
const saved = await recordLiveConversationNote({
podId: req.params.id,
sessionId: req.params.sessionId,
identity,
kind,
note,
});
res.status(201).json(saved);
} catch (e) {
res.status(400).json({ error: (e as Error).message });
}
});
app.post('/api/internal/hermes/jobs', async (req, res) => {
if (!requireInternalAgent(req, res)) return;
try {
res.status(202).json(await createHermesJob(req.body ?? {}));
} catch (e) {
res.status(400).json({ error: (e as Error).message });
}
});
app.get('/api/internal/hermes/jobs/:jobId', async (req, res) => {
if (!requireInternalAgent(req, res)) return;
const job = await getHermesJob(req.params.jobId);
if (!job) return res.status(404).json({ error: 'job not found' });
res.json(job);
});
app.post('/api/internal/hermes/jobs/:jobId/abort', async (req, res) => {
if (!requireInternalAgent(req, res)) return;
const job = await abortHermesJob(req.params.jobId);
if (!job) return res.status(404).json({ error: 'job not found' });
res.json(job);
});
app.get('/api/internal/hermes/jobs/:jobId/events', async (req, res) => {
if (!requireInternalAgent(req, res)) return;
const job = await getHermesJob(req.params.jobId);
if (!job) return res.status(404).json({ error: 'job not found' });
res.json(await listHermesJobEvents(req.params.jobId, 100));
});
app.post('/api/internal/hermes/jobs/:jobId/events', async (req, res) => {
if (!requireInternalAgent(req, res)) return;
try {
const { type, message, data } = req.body ?? {};
const eventType = hermesJobEventType(type);
if (!eventType || typeof message !== 'string') {
return res.status(400).json({ error: 'type and message are required' });
}
res.status(201).json(await appendHermesJobEvent(req.params.jobId, eventType, message, data));
} catch (e) {
res.status(400).json({ error: (e as Error).message });
}
});
app.get('/api/internal/hermes/jobs/:jobId/events/stream', async (req, res) => {
if (!requireInternalAgent(req, res)) return;
const job = await getHermesJob(req.params.jobId);
if (!job) return res.status(404).json({ error: 'job not found' });
res.setHeader('Content-Type', 'text/event-stream');
res.setHeader('Cache-Control', 'no-cache, no-transform');
res.setHeader('Connection', 'keep-alive');
res.flushHeaders?.();
let closed = false;
let lastIds = new Set<string>();
const send = async () => {
if (closed) return;
try {
const events = await listHermesJobEvents(req.params.jobId, 100);
const fresh = events.filter((event) => !lastIds.has(event.id));
lastIds = new Set(events.map((event) => event.id));
for (const event of fresh) {
res.write(`event: job-event\n`);
res.write(`data: ${JSON.stringify(event)}\n\n`);
}
const current = await getHermesJob(req.params.jobId);
if (current && ['completed', 'failed', 'aborted'].includes(current.status)) {
res.write(`event: done\n`);
res.write(`data: ${JSON.stringify(current)}\n\n`);
closed = true;
res.end();
} else {
res.write(`: keepalive ${Date.now()}\n\n`);
}
} catch (e) {
res.write(`event: error\n`);
res.write(`data: ${JSON.stringify({ error: (e as Error).message })}\n\n`);
}
};
await send();
const interval = setInterval(() => void send(), 1500);
req.on('close', () => {
closed = true;
clearInterval(interval);
});
});
// Per-pod background music (Lyria), generated once and cached. Streams MP3 the
// frontend loops as a pod-wide LiveKit track (replaces the synthesized beat).
app.get('/api/pods/:id/music', async (req, res) => {
try {
const pod = await getPod(req.params.id);
if (!pod) return res.status(404).json({ error: 'pod not found' });
const mp3 = await getPodMusic(pod.id, pod.name);
res.set('Content-Type', 'audio/mpeg');
res.set('Cache-Control', 'public, max-age=86400');
res.send(mp3);
} catch (e) {
res.status(500).json({ error: (e as Error).message });
}
});
app.post('/api/pods/:id/hermes/notify', async (req, res) => {
const podId = req.params.id;
const pod = await getPod(podId);
if (!pod) return res.status(404).json({ error: 'pod not found' });
const body = req.body ?? {};
const message = typeof body.message === 'string' ? body.message.trim() : '';
if (!message) return res.status(400).json({ error: 'message is required' });
const now = new Date().toISOString();
const engineers = stringArray(body.engineers);
const recipients = engineers.length ? engineers : pod.members.slice(0, 2);
const file =
typeof body.file === 'string' && body.file.trim() ? body.file.trim() : 'Hermes signal';
const urgent = body.urgency === 'urgent' || body.severity === 'critical';
const collision: Collision = {
id:
typeof body.collisionId === 'string' && body.collisionId
? body.collisionId
: `col_${Date.now()}`,
podId,
file,
symbol: typeof body.symbol === 'string' && body.symbol ? body.symbol : undefined,
engineers: recipients,
severity: urgent ? 'critical' : 'warn',
githubState: { unpushed: body.unpushed !== false },
detectedAt: now,
};
const intervention: Intervention = {
id:
typeof body.interventionId === 'string' && body.interventionId
? body.interventionId
: `int_${Date.now()}`,
collisionId: collision.id,
podId,
kind: urgent ? 'voice' : 'card',
message,
suggestedAction: {
kind: suggestedAction(body.suggestedAction),
params: {
file,
engineers: recipients,
source: 'local-hermes',
},
},
status: 'pending',
createdAt: now,
};
const voiceLine =
urgent && body.speak !== false
? typeof body.voiceLine === 'string' && body.voiceLine.trim()
? body.voiceLine.trim()
: message
: undefined;
try {
if (body.force !== true && (await hasRecentInterventionForCollision(collision))) {
return res.status(202).json({ ok: true, collision, intervention, livekit: 'suppressed' });
}
await recordCollision(collision);
await recordIntervention(intervention);
if (body.dryRun === true) {
return res.status(202).json({ ok: true, collision, intervention, livekit: 'dry-run' });
}
await notifyHermesInterventionInRoom(podId, collision, intervention, voiceLine);
res.status(202).json({ ok: true, collision, intervention, livekit: 'notified' });
} catch (e) {
res.status(500).json({ error: (e as Error).message });
}
});
app.delete('/api/pods/:id/members/:name', async (req, res) => {
const pod = await removeMember(req.params.id, req.params.name);
if (!pod) return res.status(404).json({ error: 'pod not found' });
res.json(pod);
});
app.get('/api/pods/:id/members/:name/history', async (req, res) => {
try {
const hours = Number(req.query.hours ?? 24) || 24;
const limit = Number(req.query.limit ?? 80) || 80;
res.json(await getMemberWorkHistory(req.params.id, req.params.name, { hours, limit }));
} catch (e) {
res.status(500).json({ error: (e as Error).message });
}
});
// --- Continual-learning graph (team_model view) ---
app.get('/api/pods/:id/graph', async (req, res) => {
try {
@@ -159,7 +610,58 @@ app.get('/api/pods/:id/graph/reach/:node', async (req, res) => {
}
});
app.get('/api/pods/:id/activity', async (req, res) => {
try {
const limit = Math.min(Number(req.query.limit ?? 80) || 80, 200);
res.json(await listPodActivity(req.params.id, limit));
} catch (e) {
res.status(500).json({ error: (e as Error).message });
}
});
app.get('/api/pods/:id/activity/stream', async (req, res) => {
res.setHeader('Content-Type', 'text/event-stream');
res.setHeader('Cache-Control', 'no-cache, no-transform');
res.setHeader('Connection', 'keep-alive');
res.flushHeaders?.();
let closed = false;
let lastPayload = '';
const send = async () => {
if (closed) return;
try {
const events = await listPodActivity(req.params.id, 80);
const payload = JSON.stringify(events);
if (payload !== lastPayload) {
lastPayload = payload;
res.write(`event: snapshot\n`);
res.write(`data: ${payload}\n\n`);
} else {
res.write(`: keepalive ${Date.now()}\n\n`);
}
} catch (e) {
res.write(`event: error\n`);
res.write(`data: ${JSON.stringify({ error: (e as Error).message })}\n\n`);
}
};
await send();
const interval = setInterval(() => void send(), 1500);
req.on('close', () => {
closed = true;
clearInterval(interval);
});
});
const http = createServer(app);
const sockets = new Set<Socket>();
http.on('connection', (socket) => {
sockets.add(socket);
socket.on('close', () => sockets.delete(socket));
});
// ws relay: the agent pushes collision/intervention JSON here; PWAs subscribed by pod receive it.
const wss = new WebSocketServer({ server: http, path: '/api/events' });
@@ -191,7 +693,11 @@ async function shutdown(signal: NodeJS.Signals): Promise<void> {
console.log(`[server] ${signal} received; shutting down`);
for (const client of clients) client.close();
wss.close();
await new Promise<void>((resolve) => http.close(() => resolve()));
for (const socket of sockets) socket.destroy();
await Promise.race([
new Promise<void>((resolve) => http.close(() => resolve())),
new Promise<void>((resolve) => setTimeout(resolve, 5000)),
]);
await closeMemory().catch((e) => console.warn(`[memory] close failed: ${(e as Error).message}`));
process.exit(0);
}
+23 -1
View File
@@ -7,6 +7,11 @@ const ai = new GoogleGenAI({ apiKey: env.GEMINI_API_KEY });
const SCHEMA = {
type: Type.OBJECT,
properties: {
mode: {
type: Type.STRING,
description:
'editing when an IDE/editor/terminal is primary; research for browser docs/SDK pages',
},
currentFile: {
type: Type.STRING,
description: 'open file path if visible, e.g. src/auth/session.ts',
@@ -20,13 +25,24 @@ const SCHEMA = {
type: Type.BOOLEAN,
description: 'dirty git gutter / modified markers visible',
},
researchTopic: {
type: Type.STRING,
description: 'topic being researched when mode is research, e.g. LiveKit agents setup',
},
researchSource: {
type: Type.STRING,
description: 'source domain when mode is research, e.g. docs.livekit.io',
},
confidence: { type: Type.NUMBER, description: '0..1 confidence in this read' },
},
propertyOrdering: [
'mode',
'currentFile',
'currentSymbol',
'activity',
'hasUnpushedChanges',
'researchTopic',
'researchSource',
'confidence',
],
} as const;
@@ -43,7 +59,10 @@ export async function analyzeFrame(
role: 'user',
parts: [
{
text: "You are PodMan watching an engineer's screen. Identify what file/symbol they are working on and whether there are uncommitted edits. JSON only.",
text:
"You are PodMan watching an engineer's screen. Return JSON only. " +
"If the primary window is an IDE/editor/terminal, set mode='editing' and identify the file, symbol, activity, and whether uncommitted edits are visible. " +
"If the primary window is a browser/docs/SDK/reference page, set mode='research', leave currentFile empty unless a file path is clearly visible, and extract researchTopic plus researchSource as the source domain.",
},
{ inlineData: { mimeType: 'image/jpeg', data: jpeg.toString('base64') } },
],
@@ -63,6 +82,9 @@ export async function analyzeFrame(
currentFile: parsed.currentFile,
currentSymbol: parsed.currentSymbol,
activity: parsed.activity,
mode: parsed.mode,
researchTopic: parsed.researchTopic,
researchSource: parsed.researchSource,
hasUnpushedChanges: parsed.hasUnpushedChanges,
confidence: parsed.confidence ?? 0.5,
observedAt: new Date().toISOString(),
+171 -30
View File
@@ -3,19 +3,46 @@ import {
AudioFrame,
AudioSource,
LocalAudioTrack,
Room,
TrackPublishOptions,
TrackSource,
type Room,
type LocalParticipant,
} from '@livekit/rtc-node';
import { GoogleGenAI, Modality, type LiveServerMessage, type Session } from '@google/genai';
import { AccessToken } from 'livekit-server-sdk';
import { DATA_TOPIC, type DataMessage } from '@podman/shared';
import { env } from '../env.js';
const SAMPLE_RATE = 24_000;
const CHANNELS = 1;
const FRAME_SAMPLES = SAMPLE_RATE / 10;
const SUBSCRIBER_READY_MS = 1_500;
const AUDIO_PREROLL_MS = 800;
const AUDIO_TAIL_MS = 1_500;
const AUDIO_HOLD_MS = 5_000;
const VOICE_QUEUE_MS = 60_000;
const VOICE_TRACK_PREFIX = 'podman-hermes-voice';
const encoder = new TextEncoder();
const ai = new GoogleGenAI({ apiKey: env.GEMINI_API_KEY });
let voiceQueue: Promise<void> = Promise.resolve();
export interface SpeakOptions {
priority?: 'normal' | 'critical';
}
function delay(ms: number): Promise<void> {
return new Promise((resolve) => setTimeout(resolve, ms));
}
function ttsPrompt(message: string): string {
return [
'Speak this PodMan coordination alert as a calm, natural engineering teammate.',
'Use warm human pacing, clear pronunciation, and a brief pause after the first sentence.',
'Do not add extra words, labels, markdown, or sound effects.',
'',
message,
].join('\n');
}
async function publishVoiceCue(room: Room, message: string): Promise<void> {
const cue: DataMessage = { type: 'VOICE_CUE', text: message };
@@ -25,12 +52,24 @@ async function publishVoiceCue(room: Room, message: string): Promise<void> {
});
}
async function unpublishVoiceTracks(localParticipant: LocalParticipant): Promise<void> {
const publications = Array.from(localParticipant.trackPublications.values()).filter(
(publication) => publication.name?.startsWith(VOICE_TRACK_PREFIX) && publication.sid,
);
for (const publication of publications) {
await localParticipant.unpublishTrack(publication.sid!, true).catch((err) => {
console.warn(`[voice] stale track cleanup failed: ${(err as Error).message}`);
});
}
}
function audioFrameFromBase64(data: string, mimeType?: string): AudioFrame | null {
if (mimeType && !mimeType.includes('audio')) return null;
const buf = Buffer.from(data, 'base64');
if (buf.byteLength < 2) return null;
const bytes = buf.byteLength % 2 === 0 ? buf : buf.subarray(0, buf.byteLength - 1);
const samples = new Int16Array(bytes.buffer, bytes.byteOffset, bytes.byteLength / 2);
const samples = new Int16Array(bytes.byteLength / 2);
for (let i = 0; i < samples.length; i += 1) samples[i] = bytes.readInt16LE(i * 2);
return new AudioFrame(samples, SAMPLE_RATE, CHANNELS, samples.length / CHANNELS);
}
@@ -53,7 +92,7 @@ function framesFromPcmBase64(data: string, mimeType?: string): AudioFrame[] {
const samples = frame.data;
const frames: AudioFrame[] = [];
for (let offset = 0; offset < samples.length; offset += FRAME_SAMPLES) {
const chunk = samples.subarray(offset, Math.min(offset + FRAME_SAMPLES, samples.length));
const chunk = samples.slice(offset, Math.min(offset + FRAME_SAMPLES, samples.length));
frames.push(new AudioFrame(chunk, SAMPLE_RATE, CHANNELS, chunk.length / CHANNELS));
}
return frames;
@@ -62,10 +101,11 @@ function framesFromPcmBase64(data: string, mimeType?: string): AudioFrame[] {
async function generateTtsFrames(message: string): Promise<AudioFrame[]> {
const res = await ai.models.generateContent({
model: env.GEMINI_LIVE_MODEL,
contents: [{ parts: [{ text: message }] }],
contents: [{ parts: [{ text: ttsPrompt(message) }] }],
config: {
responseModalities: [Modality.AUDIO],
speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: 'Kore' } } },
speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: env.GEMINI_TTS_VOICE } } },
temperature: 0.8,
},
});
const parts = res.candidates?.[0]?.content?.parts ?? [];
@@ -74,29 +114,60 @@ async function generateTtsFrames(message: string): Promise<AudioFrame[]> {
);
}
async function speakWithTts(source: AudioSource, message: string): Promise<void> {
for (const frame of await generateTtsFrames(message)) {
await source.captureFrame(frame);
}
function fallbackVoiceLine(message: string): string {
const clean = message.replace(/^heads up[.!]?\s*/i, '').trim();
if (clean && clean !== message) return clean;
return 'PodMan noticed a critical conflict. Please sync with the team before pushing.';
}
async function speakWithLive(source: AudioSource, message: string): Promise<void> {
async function speakWithTts(source: AudioSource, message: string): Promise<number> {
let frames: AudioFrame[];
try {
frames = await generateTtsFrames(message);
} catch (err) {
const fallback = fallbackVoiceLine(message);
console.warn(`[voice] Gemini TTS retrying with fallback line: ${(err as Error).message}`);
frames = await generateTtsFrames(fallback);
}
if (frames.length === 0) throw new Error('Gemini TTS returned no audio frames');
const durationMs = frames.reduce(
(sum, frame) => sum + (frame.samplesPerChannel / frame.sampleRate) * 1000,
0,
);
console.log(
`[voice] publishing Gemini TTS audio frames=${frames.length} durationMs=${Math.round(durationMs)}`,
);
for (const frame of frames) {
await source.captureFrame(frame);
}
return durationMs;
}
async function speakWithLive(source: AudioSource, message: string): Promise<number> {
let durationMs = 0;
let done: () => void = () => {};
const donePromise = new Promise<void>((resolve) => {
done = resolve;
});
const session: Session = await ai.live.connect({
model: env.GEMINI_LIVE_MODEL,
config: { responseModalities: [Modality.AUDIO] },
config: {
responseModalities: [Modality.AUDIO],
speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: env.GEMINI_TTS_VOICE } } },
temperature: 0.8,
},
callbacks: {
onmessage: (event) => {
void (async () => {
for (const frame of audioFrames(event)) await source.captureFrame(frame);
for (const frame of audioFrames(event)) {
durationMs += (frame.samplesPerChannel / frame.sampleRate) * 1000;
await source.captureFrame(frame);
}
if (event.serverContent?.turnComplete || event.serverContent?.generationComplete) done();
})();
},
onerror: (event) => {
console.warn(`[voice] Gemini Live error: ${event.message}`);
console.warn(`[voice] Gemini voice error: ${event.message}`);
done();
},
onclose: done,
@@ -104,36 +175,106 @@ async function speakWithLive(source: AudioSource, message: string): Promise<void
});
session.sendClientContent({
turns: [{ role: 'user', parts: [{ text: message }] }],
turns: [{ role: 'user', parts: [{ text: ttsPrompt(message) }] }],
turnComplete: true,
});
await Promise.race([donePromise, new Promise((resolve) => setTimeout(resolve, 15_000))]);
session.close();
return durationMs;
}
/**
* Speak a message into the LiveKit room using Gemini Live audio. A data-channel
* VOICE_CUE is sent first so clients still get the cue if audio generation or
* publishing fails.
*/
export async function speak(room: Room, message: string): Promise<void> {
await publishVoiceCue(room, message);
if (!room.localParticipant) return;
async function waitForVoicePlayout(source: AudioSource): Promise<void> {
if (source.queuedDuration <= 0) return;
const queuedMs = Math.round(source.queuedDuration);
console.log(`[voice] waiting for queued audio playout queuedMs=${queuedMs}`);
await Promise.race([
source.waitForPlayout(),
new Promise((resolve) => setTimeout(resolve, VOICE_QUEUE_MS + 2_000)),
]);
console.log('[voice] queued audio playout complete');
}
const source = new AudioSource(SAMPLE_RATE, CHANNELS);
const track = LocalAudioTrack.createAudioTrack('podman-hermes-voice', source);
async function captureSilence(source: AudioSource, durationMs: number): Promise<void> {
const totalSamples = Math.max(1, Math.round((SAMPLE_RATE * durationMs) / 1000));
for (let offset = 0; offset < totalSamples; offset += FRAME_SAMPLES) {
const samples = Math.min(FRAME_SAMPLES, totalSamples - offset);
await source.captureFrame(new AudioFrame(new Int16Array(samples), SAMPLE_RATE, CHANNELS, samples));
}
}
async function speakAudio(room: Room, message: string): Promise<void> {
const localParticipant = room.localParticipant;
if (!localParticipant) return;
await unpublishVoiceTracks(localParticipant);
const source = new AudioSource(SAMPLE_RATE, CHANNELS, VOICE_QUEUE_MS);
const track = LocalAudioTrack.createAudioTrack(`${VOICE_TRACK_PREFIX}-${Date.now()}`, source);
const options = new TrackPublishOptions();
options.source = TrackSource.SOURCE_MICROPHONE;
let publicationSid: string | undefined;
try {
const publication = await room.localParticipant.publishTrack(track, options);
if (env.GEMINI_LIVE_MODEL.includes('tts')) await speakWithTts(source, message);
else await speakWithLive(source, message);
if (publication.sid) await room.localParticipant.unpublishTrack(publication.sid, true);
await source.close();
const publication = await localParticipant.publishTrack(track, options);
publicationSid = publication.sid;
await delay(SUBSCRIBER_READY_MS);
await captureSilence(source, AUDIO_PREROLL_MS);
const audioDurationMs = env.GEMINI_LIVE_MODEL.includes('tts')
? await speakWithTts(source, message)
: await speakWithLive(source, message);
await captureSilence(source, AUDIO_TAIL_MS);
await waitForVoicePlayout(source);
const manualHoldMs = Math.ceil(audioDurationMs + AUDIO_TAIL_MS + AUDIO_HOLD_MS);
console.log(`[voice] holding track for subscriber playout holdMs=${manualHoldMs}`);
await delay(manualHoldMs);
} catch (err) {
console.warn(`[voice] Gemini Live publish failed: ${(err as Error).message}`);
console.warn(`[voice] Gemini voice publish failed: ${(err as Error).message}`);
} finally {
if (publicationSid) {
await localParticipant.unpublishTrack(publicationSid, true).catch((err) => {
console.warn(`[voice] track unpublish failed: ${(err as Error).message}`);
});
}
await source.close().catch(() => {});
}
}
/**
* Speak a message into the LiveKit room using Gemini audio. A data-channel
* VOICE_CUE is sent first so clients still get the cue if audio generation or
* publishing fails.
*/
export async function speak(room: Room, message: string, options: SpeakOptions = {}): Promise<void> {
await publishVoiceCue(room, message);
if (options.priority === 'critical') {
await speakAudio(room, message);
return;
}
voiceQueue = voiceQueue.catch(() => {}).then(() => speakAudio(room, message));
await voiceQueue;
}
export async function speakInRoom(
roomName: string,
message: string,
options: SpeakOptions = {},
): Promise<void> {
const room = new Room();
try {
const at = new AccessToken(env.LIVEKIT_API_KEY, env.LIVEKIT_API_SECRET, {
identity: `podman-voice-${Date.now()}`,
name: 'PodMan voice',
ttl: '5m',
});
at.addGrant({
roomJoin: true,
room: roomName,
canPublish: true,
canSubscribe: true,
canPublishData: true,
});
await room.connect(env.LIVEKIT_URL, await at.toJwt());
await speak(room, message, options);
} finally {
await room.disconnect().catch(() => {});
}
}
+94
View File
@@ -0,0 +1,94 @@
import { Buffer } from 'node:buffer';
import { getDb } from '../memory/db.js';
import { env } from '../env.js';
// Lyria 3 is reached via the Gemini "interactions" endpoint (not :predict, which
// is the Vertex path). The clip model returns a ~30s base64 MP3.
const MUSIC_MODEL = process.env.GEMINI_MUSIC_MODEL ?? 'lyria-3-clip-preview';
const INTERACTIONS_URL = 'https://generativelanguage.googleapis.com/v1beta/interactions';
interface PodMusicDoc {
podId: string;
name: string; // pod name the vocal was generated for
model: string;
mp3Base64: string;
createdAt: string;
}
interface InteractionContent {
type?: string;
data?: string;
text?: string;
}
interface InteractionResponse {
steps?: Array<{ content?: InteractionContent[] }>;
output_audio?: { data?: string };
}
/**
* Background "hold music" prompt: opens with the pod name sung once, then a calm
* instrumental bed that loops. Keep it unobtrusive this is fill, not a song.
*/
function musicPrompt(podName: string): string {
return [
'Calm soothing instrumental background hold music for a tech app, like gentle on-hold lobby music.',
`It opens in the first three seconds with a soft gentle voice clearly saying the words "${podName}" one time,`,
'and after that opening it is purely instrumental with warm electric piano, gentle synth pads and a soft relaxed beat.',
'Unobtrusive, pleasant and steady with no climax, designed to loop seamlessly as quiet background fill.',
'No other lyrics or vocals after the opening.',
].join(' ');
}
function extractAudioBase64(data: InteractionResponse): string | null {
for (const step of data.steps ?? []) {
for (const c of step.content ?? []) {
if (c.type === 'audio' && c.data) return c.data;
}
}
return data.output_audio?.data ?? null;
}
async function generate(podName: string): Promise<Buffer> {
const res = await fetch(INTERACTIONS_URL, {
method: 'POST',
headers: { 'Content-Type': 'application/json', 'x-goog-api-key': env.GEMINI_API_KEY },
body: JSON.stringify({ model: MUSIC_MODEL, input: musicPrompt(podName) }),
});
if (!res.ok) {
throw new Error(`Lyria ${res.status}: ${(await res.text()).slice(0, 300)}`);
}
const data = (await res.json()) as InteractionResponse;
const b64 = extractAudioBase64(data);
if (!b64) throw new Error('Lyria returned no audio');
return Buffer.from(b64, 'base64');
}
/**
* The pod's background-music MP3, generated by Lyria on first request and cached
* in the `pod_music` collection. Regenerated if the pod name changes so the sung
* name stays correct. Lyria generation is slow (~20s); the cache makes every
* call after the first instant.
*/
export async function getPodMusic(podId: string, podName: string): Promise<Buffer> {
const db = await getDb();
const col = db.collection<PodMusicDoc>('pod_music');
const cached = await col.findOne({ podId });
if (cached && cached.name === podName && cached.model === MUSIC_MODEL && cached.mp3Base64) {
return Buffer.from(cached.mp3Base64, 'base64');
}
const mp3 = await generate(podName);
await col.updateOne(
{ podId },
{
$set: {
podId,
name: podName,
model: MUSIC_MODEL,
mp3Base64: mp3.toString('base64'),
createdAt: new Date().toISOString(),
},
},
{ upsert: true },
);
return mp3;
}
-293
View File
@@ -1,293 +0,0 @@
# Team Memory Graph — Redesign Brief (fresh-session handoff)
> **You are a fresh Claude Code session with no prior context. Read this whole file first.**
> Your job: rebuild the **live "Team memory" graph view** so the **light/real-data** version is as
> polished and functional as the original **dark Bauhaus mock**, and make the graph **dynamic**
> (force-directed + animated), not the current dead static-column layout.
> **Do not rewrite the backend materializer — it is good.** The problem is 100% the frontend rendering.
---
## 0. Mission (one paragraph)
PodMan's "Team memory" is a per-pod graph of who owns/edits which files, where work collides, and what
PodMan learned from accepted interventions — the continual-learning loop made legible in 10 seconds.
A **dark Bauhaus mock** of this view looks great (clean 3-panel layout, a learning-loop rail, an activity
stream, a readable graph). The **shipped light version on real data looks terrible** (a hairball of red
edges, overlapping labels, a static lifeless layout, and it's missing the learning-loop rail + activity
stream entirely). Make the light version match the mock's structure/polish/functionality, in the app's
**light shadcn theme**, and make the **graph dynamic** (organic force-directed layout, draggable,
animated transitions). Keep using **real data** from the existing materializer.
---
## 1. The two reference points
### A. The dark Bauhaus mock = what "good" looks like (target structure)
A single dark card titled **"PODMAN — CONTINUAL-LEARNING OBSERVATORY"** with a `LIVE · POD demo-pod`
status. Layout:
- **Left rail — WORKFLOW METRICS**: a vertical stack of bordered cards, each a big numeral + an
UPPERCASE tracked label + a one-line detail, with a colored left-accent bar:
`03 PODS WATCHED`, `05 ENGINEERS LIVE`, `02 COLLISIONS OPEN (▲ auth.ts critical)`,
`01 INTERVENTION SENT`, `86% ACCEPT RATE (▲ +14% this session)`, `124 MEMORY VECTORS`.
- **Center — the GRAPH**: sparse, geometric, readable. Node shapes encode kind
(engineer = filled square, file = outlined square, feature = circle, collision = triangle,
intervention = diamond). One **risk path is lit** (Karti+Yahya → auth.ts → collision → sync PR →
`learned_from`), everything else dimmed. Edges color-coded (`collides` red, `warns` amber/orange,
`learned_from` dashed violet, `owns` blue, `editing` paper, `touches` grey).
- **Right rail — LEARNING LOOP**: a vertical 5-step stepper with the active step highlighted/pulsing:
`01 OBSERVE (vision → 5 contexts/s)``02 STORE (124 vectors · Atlas)`
`03 PREDICT (2 collisions flagged)` [active] → `04 OUTCOME (1 accepted · 0 dismissed)`
`05 ADAPT (Karti→auth ownership +)`. Arrows between steps.
- **Bottom-left — ACTIVITY STREAM**: a time-stamped feed with colored kind-tags:
`15:48 EDITING Yahya opened auth.ts — unpushed changes detected`,
`15:48 COLLISION Critical overlap on auth.ts · Karti + Yahya`,
`15:49 WARNS PodMan spoke: "open a sync PR?" → card sent`,
`15:49 OUTCOME Sync PR accepted by the pod`,
`15:49 LEARNED_FROM Memory updated: Karti owns auth (confidence ↑)`.
- **Bottom-right — SELECTED NODE**: click a node → kind / name / relationships count / severity / a
one-line "why" (`Two engineers editing the same file before push — the signal git can't see.`).
- **Legend**: engineer / file / feature / collision / intervention · collides / warns / learned_from.
It reads in 10 seconds because it is **sparse, color-coded, and tells the loop story** with the rails +
stream, not just a node blob. (The full mock HTML/CSS is reproduced in **Appendix A** — port its
structure to light shadcn.)
### B. The shipped light version = what's wrong (the thing to fix)
Same data, but: a **hairball** — every engineer fans red `collides` edges to ~6 collision triangles
(all labeled `sync PR`); **file labels overlap** in a dim middle column; the graph uses a **static
deterministic column layout** (`x` by kind, `y` evenly spread) so it looks dead/lifeless; and it is
**missing the LEARNING LOOP rail and the ACTIVITY STREAM** entirely — it's just a metrics rail + the bare
graph + a selected-node panel. (Two already-fixed-on-branch items: garbage collision labels
`infra/README.md### Running the git watcher` and full-path bleed — see PR #7 / commit `a21a289`,
`shortLabel` in `live.ts`. Build on top of that, don't redo it.)
---
## 2. The gap to close (light vs mock)
| Mock has | Light version | Action |
| ---------------------------------------- | ------------------------------- | ---------------------------------- |
| Workflow metrics rail | ✅ has it (3 metrics) | keep; restyle to match |
| **Learning loop rail (observe→…→adapt)** | ❌ missing | **build it** (needs live counts) |
| **Activity stream feed** | ❌ missing | **build it** (needs an event feed) |
| Selected-node panel | ✅ has it | keep |
| **Dynamic / animated graph** | ❌ static columns | **replace the layout** |
| Sparse, lit "risk path" | partial (Risk-path mode exists) | improve emphasis + spacing |
| Legend | ✅ | keep |
---
## 3. Current architecture (build on this — do NOT rewrite the materializer)
**Backend (good, keep):**
- `backend/src/graph/live.ts``materializePodGraph(podId)`: builds the graph from the real Mongo
collections (`pods`, `engineer_states`, `observations`, `collisions`, `interventions`, `outcomes`).
It already de-noises hard: collapses collisions by `memorySignature`, caps to 8, collapses
interventions to one per collision, filters junk files (`isFilePath`), prunes test-artifact engineers
(`ENGINEER_NOISE`), caps files to 9, short labels (`shortLabel`). Output is ~20 clean nodes for
`demo-pod`. **This is solid — extend it, don't replace it.**
- `backend/src/graph/store.ts``loadPodGraph(podId)`: live materializer → seeded `team_model.graph`
→ demo fallback (`createDemoPodGraph`). Plus `reachFrom` (`$graphLookup`).
- Route: `GET /api/pods/:id/graph` returns `PodGraph` (also `/graph/reach/:node`).
- `backend/src/memory/db.ts``collections()`, `getGitStates(podId)`, `getDb()`.
- WS bus: `backend/src/server.ts` hosts `ws /api/events` (the agent + `/api/outcome` broadcast here).
**Frontend (this is where the work is):**
- `frontend/src/components/GraphView.tsx`**the thing you redesign** (~90% of the work). Currently:
fetches `/api/pods/:id/graph`, renders a bespoke SVG with the static column layout, Risk/Learning/Whole
toggles, a metrics rail, a selected-node panel, a legend. Composed from shadcn primitives (`Button`,
`Badge`) + Tailwind utilities. Theme-aware via shadcn tokens.
- `frontend/src/lib/graph.ts``fetchPodGraph(podId)`.
- Opened from each `PodCard`'s `⋯` menu → "Team memory" (`onOpenGraph(pod.id)` in
`frontend/src/App.tsx`). It is a conditional render (no route).
**Data contract** (`shared/src/graph.ts`):
```ts
PodGraph = { podId, generatedAt, nodes: PodGraphNode[], edges: PodGraphEdge[], metrics: PodGraphMetric[] }
PodGraphNode = { id, kind, label, summary, weight 0..1, status: 'stable'|'active'|'risk'|'learned', x, y }
// kind: 'engineer'|'feature'|'file'|'collision'|'intervention'
PodGraphEdge = { id, source, target, kind, label, strength 0..1 }
// kind: 'owns'|'editing'|'touches'|'collides'|'warns'|'learned_from'
PodGraphMetric = { label, value, detail }
```
**Theme / components (HARD RULE):** the app is **light shadcn**, built from the **ruixen registry**
add primitives with `npx shadcn@latest add "https://ruixen.com/r/[component]"` and compose from
`@/components/ui/*` (`Button`, `Badge`, `Card`, `Tabs`, `ToggleGroup`, etc.) using the design tokens
(`var(--card)` / `--foreground` / `--muted-foreground` / `--border`, `--chart-1..5`). Only the SVG/canvas
graph is bespoke. Match `frontend/src/App.tsx`'s `StatPill`/`BriefLine` utility patterns.
---
## 4. Target design (build this)
A light shadcn page with the **mock's structure**:
```
┌───────────────────────────────────────────────────────────────────────┐
│ Header: "Team memory · What PodMan learned · <pod>" [← Pods] │
├───────────────────────────────────────────────────────────────────────┤
│ Toggles: Risk path | Learning edges | Whole graph (keep) │
├──────────────┬──────────────────────────────────┬─────────────────────┤
│ WORKFLOW │ │ LEARNING LOOP │
│ METRICS │ DYNAMIC GRAPH CANVAS │ observe→store→ │
│ (cards) │ (force-directed + animated) │ predict→outcome→ │
│ │ │ adapt (active pulses)│
├──────────────┴──────────────────────────────────┴─────────────────────┤
│ ACTIVITY STREAM (time-tagged feed) │ SELECTED NODE (detail) │
└───────────────────────────────────────────────────────────────────────┘
```
- **Light shadcn** throughout (theme-aware; follows dark mode if the app ever toggles). Keep the
geometric **node-shape + color encoding** (it's the legible part) but on light surfaces with the
app's hues (engineer blue `#2563eb`, file slate outline `#475569`, feature amber `#d97706`,
collision red `#dc2626`, intervention violet `#7c3aed`; edges: collides red, warns amber,
learned_from dashed violet, owns blue, editing slate, touches faint slate).
- **Default to "Risk path"**: light the collision→intervention→`learned_from` chain; dim the rest.
---
## 5. Make the graph DYNAMIC (the headline new requirement)
The static column layout (`live.ts` `layout()` sets `x`/`y` by kind) looks dead. Replace the frontend
rendering with a **dynamic** graph. Pick one (recommended order):
1. **`d3-force` force-directed (recommended).** Add `d3-force` (small). Run a force simulation on the
`PodGraph` nodes/edges: link force (by `edge.strength`), charge/repulsion, center, collision radius
(by `node.weight`). Render nodes/edges as SVG, update positions per tick. Make nodes **draggable**
(pin on drag). Animate new nodes/edges fading in on data refresh, and the `learned_from` dashed
stroke animating. Ignore the server's `x`/`y` (or use them as initial positions). Keep node shapes.
2. `react-force-graph` / `force-graph` (canvas) — heavier, faster for big graphs; overkill at ~20 nodes
but fine.
3. A custom animated **layered** layout (engineers → files → collisions → interventions columns, but
with curved edges, eased position transitions on refresh, and gentle idle motion). Lighter-weight
than d3-force; still feels alive if you animate transitions.
**Realtime/dynamic data:** poll `GET /api/pods/:id/graph` every ~5s and **animate the diff** between
snapshots (don't hard-replace). Optionally subscribe to `ws /api/events` for instant nudges. New
collisions/interventions should visibly animate in; the `learned_from` edge + gold node should pop on a
new accepted outcome.
**De-hairball:** even force-directed, ~6 collisions × 3 engineers = many `collides` edges. Mitigate:
bundle/curve edges, lower non-risk edge opacity, default to Risk-path emphasis, size nodes by `weight`,
and keep label collision-avoidance (offset labels, hide on overlap, show on hover/select).
---
## 6. Data for the new panels (extend the materializer or add endpoints)
The mock's **Learning Loop** and **Activity Stream** need data the current `PodGraph` doesn't carry. Two
options: (a) extend `materializePodGraph` to also return `loop` + `activity`, or (b) add small endpoints.
Recommended: extend the return type (additive to `shared/src/graph.ts`).
- **Learning loop counts** (`observe→store→predict→outcome→adapt`):
- observe = recent `observations` count (e.g. last 60s) / rate
- store = `memory_vectors` or `collisions.embedding` count (Voyage vectors)
- predict = open `collisions` (distinct signatures) count
- outcome = `outcomes` accepted vs dismissed counts
- adapt = `team_model.ownership` entries / learned owners count
- mark the "active" stage = the most recent activity.
- **Activity stream**: merge + time-sort recent events from `collisions.detectedAt`,
`interventions.createdAt`, `outcomes.recordedAt`, `engineer_states.gitUpdatedAt` → a typed feed
`{ at, kind: 'editing'|'collision'|'warns'|'outcome'|'learned_from', text }`. Cap to ~8 most recent.
(`backend/src/memory/db.ts` `collections()` gives you `observations/collisions/interventions/outcomes`;
`getGitStates` gives engineer_states; `team_model` is `db.collection('team_model')`.)
---
## 7. Files to touch
- **`frontend/src/components/GraphView.tsx`** — the redesign (force-directed graph + 3-panel layout +
learning-loop rail + activity stream). May split into `GraphCanvas.tsx`, `LearningLoop.tsx`,
`ActivityStream.tsx`, `MetricsRail.tsx`.
- **`frontend/src/lib/graph.ts`** — add fetches for loop/activity if you add endpoints.
- **`backend/src/graph/live.ts`** (extend, don't rewrite) — emit `loop` + `activity` in the result;
keep all the de-noise.
- **`shared/src/graph.ts`** — add `loop`/`activity` types to `PodGraph` (additive).
- **deps**`d3-force` (+ `@types/d3-force`) via pnpm in `frontend`.
- Possibly add a ruixen primitive (e.g. `timeline`, `stepper`) via the shadcn CLI if one fits.
---
## 8. Constraints & gotchas (READ — these will bite you)
- **The materializer is good — do not rewrite it.** It already de-noises (caps, collapse-by-signature,
engineer/file filters, short labels). The bad UI is the **frontend layout/render**, not the data.
- **PWA service worker caches aggressively** — after any deploy, hard-refresh (Cmd-Shift-R) or test in a
private window, or you'll think nothing changed.
- **Deploy = merge to `main`** (DO `deploy_on_push: true`). `main` is shared by ~4 engineers and moves
fast. Work on a branch, open a PR, merge. Don't push to `main` directly.
- **`learned_from` "money" edge won't render on `demo-pod`** right now — its one accepted outcome is
orphaned (points at a collision deleted by test churn). It needs **one intact accept flow** (real
collision → intervention → someone clicks Accept) to draw. To demo, seed a clean chain or clear test
docs (writes to shared Atlas — confirm scope first).
- **Atlas creds rotate frequently**`podman/.env`'s `MONGODB_URI` may be stale; the **deployed env**
has the working one. If local Mongo auth fails, that's why.
- **Build verification**: some sandboxes can't run `pnpm`/`vite`/shadcn deps (`lucide-react`,
`@radix-ui`). Verify the frontend with `pnpm build` in a real env or CI before merging. Backend
typecheck excludes uninstalled `ws`/`sharp`/`@livekit/rtc-node` noise.
- **Compose from ruixen/shadcn primitives** (`npx shadcn add ruixen.com/r/[component]`,
`@/components/ui/*`); only the SVG/canvas graph is bespoke. Match `StatPill`/`BriefLine` in `App.tsx`.
- **Light theme + tokens** — never hardcode dark colors for chrome; use `var(--card)/--foreground/…`.
Keep fixed semantic hues only for the node/edge kind encoding.
---
## 9. Acceptance criteria
- Light Team-memory page matches the mock's structure: **metrics rail + dynamic graph + learning-loop
rail + activity stream + selected-node panel + legend**.
- **Graph is dynamic**: force-directed (or animated layered), **draggable**, **animates** new
nodes/edges and refresh transitions; **no overlapping labels, no hairball**.
- Reads in 10s; the **risk/money path is obvious** by default.
- **Light shadcn** theme, theme-aware; composed from ruixen primitives.
- Uses **real data** from `materializePodGraph`; graceful demo fallback when empty.
- `pnpm build` + typechecks pass; deploys; verified after a hard-refresh.
---
## 10. Suggested first moves for the new session
1. Read this file + `docs/graph.md` + `docs/live-ui-spec.md` (R1/R2 sections) + `CLAUDE.md`.
2. `git fetch`; branch off `main` (or `feat/live-graph-glue`, which has the latest graph work).
3. Hit the live data once: `curl https://165-22-129-249.sslip.io/api/pods/demo-pod/graph` — that's the
real `PodGraph` you'll render.
4. Build a `d3-force` `GraphCanvas` first (replace the static layout), get it draggable + animated.
5. Add `LearningLoop` + `ActivityStream` (extend the materializer to feed them).
6. Polish to the mock; `pnpm build`; PR → main → redeploy → hard-refresh.
---
## Appendix A — the dark mock (reference structure to port to light)
The mock is a single dark card. Structure + the exact content to reproduce (in light shadcn):
- Header: brand glyph (blue square + amber circle + red triangle + outlined square) + `PODMAN /
CONTINUAL-LEARNING OBSERVATORY` + `● LIVE · POD demo-pod`.
- Grid `180px 1fr 196px`: **metrics rail** | **graph** | **learning-loop rail**.
- Metrics cards: big `Archivo`-weight numeral, uppercase tracked label, muted detail, colored
left-accent (blue/red/yellow/violet/green).
- Graph: SVG, geometric node shapes by kind, color-coded edges, one lit risk path, dim others; click a
node → highlight its incident edges + neighbors, fill the selected-node panel.
- Learning-loop rail: 5 bordered steps with number + UPPERCASE title + muted sub; the active step has a
pulsing left bar; `↓` arrows between.
- Legend row (node kinds + edge kinds).
- Activity stream (time + colored tag + text) and selected-node panel below.
Palette used (port to shadcn tokens for chrome; keep these as node/edge hues):
`bg #0c0c0e`, `panel #141417`, `line #2a2a31`, `paper/text #ECE7DA`, `muted #8d897e`,
`blue #3B5BFF`, `red #E2403A`, `amber #F6C445`, `violet #8b6cff`, `green #46c07a`.
For light: chrome → `var(--card)/--foreground/--border/--muted-foreground`; node/edge hues →
blue `#2563eb`, slate `#475569`, amber `#d97706`, red `#dc2626`, violet `#7c3aed` (light-readable).
> The dark mock was a `show_widget` demo (not a saved file). If you want the literal HTML/CSS, ask the
> user to paste it, or reconstruct from this appendix — the **structure + content above is the spec**.
> Goal: same structure, same legibility, **light theme + dynamic graph**.
-483
View File
@@ -1,483 +0,0 @@
# Team Memory Graph - Codex Redesign Brief
> **Read this whole file before touching code.**
> This is a fresh-session handoff for rebuilding PodMan's live **Team memory** graph UI.
> The goal is not to tweak labels or add another filter. The goal is to make the
> real-data light UI as polished, legible, and functional as the original dark
> Bauhaus mock, while keeping the graph backed by live MongoDB data.
## 0. Mission
PodMan's Team memory view should make the recursive self-improvement loop visible:
who is working, which files overlap, where collisions happen, which intervention was
sent, and what PodMan learned from accepted outcomes.
The current light real-data implementation proves the backend can materialize a graph,
but the UI does not yet tell the story. It still reads as a static node-link diagram:
edges dominate, labels collide, the layout feels fixed, and the important loop
`observe -> store -> predict -> outcome -> adapt` is not visible.
Rebuild the Team memory experience so it has the narrative clarity of the dark
Bauhaus mock, in the app's light shadcn/ruixen visual system, with a dynamic graph
that animates and responds to live data changes.
## 1. Current State
Branch context:
- Work is on `feat/live-graph-glue`.
- The live graph backend exists and should be reused.
- A PR for the live graph glue already exists, and later commits have continued
refining readability.
- The current file requested by the user is this document:
`codex_team_memory_redisgn.md`.
Important existing files:
- `backend/src/graph/live.ts`
Builds `PodGraph` from real Mongo collections:
`pods`, `engineer_states`, `observations`, `collisions`, `interventions`,
and `outcomes`.
- `backend/src/graph/store.ts`
Loads live graph first, then seeded `team_model.graph`, then demo fallback.
- `shared/src/graph.ts`
Defines the graph contract.
- `frontend/src/components/GraphView.tsx`
Current frontend graph rendering. This is the main file to redesign.
- `frontend/src/lib/graph.ts`
Fetches the graph.
- `frontend/src/App.tsx` and `frontend/src/components/PodCard.tsx`
Open Team memory per pod.
Do **not** start by rewriting the backend materializer. It already does the most
important real-data work: filtering noisy files, collapsing repeated collisions,
capping graph size, shortening labels, and pruning test engineers. The redesign is
primarily a frontend information-architecture and interaction problem.
## 2. Reference Screens
### Current Light Real-Data UI
The light UI is technically real and connected to live data, but it fails visually.
Observed problems:
- The graph is too static and column-like.
- Red collision edges dominate the canvas.
- Labels overlap and fight for attention.
- Interventions repeat as a row of identical diamonds.
- The right panel says "It learned" but does not explain the actual workflow state.
- The screen lacks an activity stream.
- The screen lacks the explicit learning-loop rail from the dark mock.
- The viewer cannot quickly answer:
- What happened?
- Who collided?
- What did PodMan do?
- Did the team accept it?
- What changed in memory?
The current light version proves data plumbing. It does not yet work as a demo
surface.
### Dark Bauhaus Mock
The dark mock is the quality target. Do not copy the dark palette wholesale, but
copy the structure, density, and storytelling.
The mock has:
- A strong title bar:
`PODMAN / CONTINUAL-LEARNING OBSERVATORY`
- A live status indicator:
`LIVE - POD demo-pod`
- A left metrics rail:
workflow metrics as compact, high-contrast cards.
- A center graph:
sparse, geometric, readable, with one primary path emphasized.
- A right learning-loop rail:
`Observe -> Store -> Predict -> Outcome -> Adapt`
- A bottom activity stream:
timestamped events with type badges.
- A selected-node detail panel:
kind, relationships, severity, explanation.
- A legend:
node shapes and edge colors.
The mock works because it is not just a graph. It is an observatory. It tells the
loop story.
## 3. Product Goal
Team memory should be the "it learned" surface.
In a 10-second demo, a viewer should understand:
1. Two engineers are converging on the same file.
2. PodMan detected the risk before a push.
3. PodMan suggested an intervention.
4. The team accepted or dismissed the intervention.
5. PodMan retained that outcome as memory.
6. Future collisions become more informed.
The graph should support that story, not overwhelm it.
## 4. Target Layout
Build a light shadcn page with the same conceptual structure as the dark mock.
```text
+--------------------------------------------------------------------------+
| Header: Team memory - What PodMan learned - <pod> [<- Pods] |
+--------------------------------------------------------------------------+
| Mode controls: Risk path | Learning edges | Whole graph |
+---------------+--------------------------------------+-------------------+
| Workflow | | Learning loop |
| metrics | Dynamic graph canvas | Observe |
| cards | | Store |
| | | Predict |
| | | Outcome |
| | | Adapt |
+---------------+--------------------------------------+-------------------+
| Activity stream | Selected node details |
+--------------------------------------------------------------------------+
```
Required panels:
- **Header**
- Pod name / id.
- Live/generated timestamp.
- Back to pods action.
- **Mode controls**
- Risk path.
- Learning edges.
- Whole graph.
- **Workflow metrics rail**
- Learned owners.
- Open risk paths.
- Accept rate.
- Optional: observations, interventions, memory vectors if available.
- **Dynamic graph canvas**
- Force-directed or animated layered graph.
- Geometric node shapes.
- Curved or bundled edges.
- Labels should not overlap by default.
- Hover/select reveals full details.
- **Learning loop rail**
- Observe.
- Store.
- Predict.
- Outcome.
- Adapt.
- Active/current step should pulse or be highlighted.
- **Activity stream**
- Recent editing, collision, warning, outcome, learned events.
- Compact rows with timestamp + colored type badge.
- **Selected node**
- Default state explains the loop.
- Selected state shows node kind, name, relationships, severity/status, and
why this node matters.
## 5. Visual Direction
Use the app's light shadcn/ruixen design system for chrome.
Hard rules:
- Use `@/components/ui/*` primitives where possible.
- If a primitive is missing, add it through:
`npx shadcn@latest add "https://ruixen.com/r/[component]"`
- Do not make the entire UI a bespoke CSS island.
- The graph canvas itself may be bespoke SVG/canvas.
- The rest should be composed from cards, badges, buttons, tabs/toggles, and
utility classes consistent with `App.tsx`.
Keep semantic graph colors:
- Engineer: blue.
- File: slate outline.
- Feature: amber circle.
- Collision: red triangle.
- Intervention: violet diamond.
- `collides`: red edge.
- `warns`: amber/orange edge.
- `learned_from`: dashed violet edge.
- `owns`: blue edge.
- `editing` / `touches`: muted slate.
Use light surfaces:
- Background: app background token.
- Panels: `card`.
- Borders: `border`.
- Text: `foreground`.
- Supporting copy: `muted-foreground`.
The result should feel like the dark mock translated into the app's light command
center, not a random analytics dashboard.
## 6. Dynamic Graph Requirement
The current graph is too static. Replace or augment the static column layout.
Preferred implementation:
- Use `d3-force` in the frontend.
- Initialize nodes from server `x/y` when useful, but let the simulation settle.
- Use:
- link force by edge strength.
- charge force for separation.
- center force.
- collision force based on node radius.
- optional x/y bias by kind to preserve rough story flow.
- Make nodes draggable.
- Preserve node shape encoding.
- Animate:
- new nodes fading/scaling in.
- new edges drawing/fading in.
- `learned_from` dashed edge flowing or pulsing.
- active collision/intervention pulse.
If `d3-force` is too much for the current branch, use an animated layered layout:
- Engineers left.
- Files mid-left.
- Collisions center/right.
- Interventions right.
- Curved edges.
- Smooth transitions between graph snapshots.
- Gentle idle motion only if it helps.
Do not leave the final version as static fixed columns.
## 7. De-Hairball Rules
Default screen should show the risk path, not every possible relationship.
Rules:
- Default mode: `Risk path`.
- Whole graph can exist, but it is not the demo default.
- Dim non-selected/non-risk edges aggressively.
- Use curved edges or edge bundling.
- Hide low-priority labels until hover/select.
- Prefer file basename/short path on canvas.
- Put full path in selected-node panel.
- Group repeated collisions by signature.
- Cap visible collisions/interventions for demo readability.
- Preserve all data in the payload; choose a readable default projection.
The graph is not an exhaustive database browser. It is a story-first visualization.
## 8. Data Model To Use
Current `PodGraph` contract:
```ts
interface PodGraph {
podId: string;
generatedAt: string;
nodes: PodGraphNode[];
edges: PodGraphEdge[];
metrics: PodGraphMetric[];
}
```
Node kinds:
- `engineer`
- `file`
- `feature`
- `collision`
- `intervention`
Edge kinds:
- `owns`
- `editing`
- `touches`
- `collides`
- `warns`
- `learned_from`
Statuses:
- `stable`
- `active`
- `risk`
- `learned`
Existing collections behind the materializer:
- `pods`
- `engineer_states`
- `observations`
- `collisions`
- `interventions`
- `outcomes`
Important caveat:
The current `demo-pod` accepted outcome chain may be orphaned from test churn.
If `learned_from` does not show, confirm whether there is an intact:
```text
collision -> intervention -> accepted outcome
```
Do not assume the UI is broken until this data chain is verified.
## 9. Extend Data For Missing Panels
The current graph contract does not fully support the dark mock's learning-loop
rail or activity stream.
Recommended additive extension:
```ts
interface PodGraphLoopStep {
id: 'observe' | 'store' | 'predict' | 'outcome' | 'adapt';
label: string;
value: string;
detail: string;
status: 'idle' | 'active' | 'complete';
}
interface PodGraphActivity {
id: string;
at: string;
kind: 'editing' | 'collision' | 'warns' | 'outcome' | 'learned_from';
label: string;
detail: string;
nodeId?: string;
edgeId?: string;
}
interface PodGraph {
...
loop?: PodGraphLoopStep[];
activity?: PodGraphActivity[];
}
```
Possible data mappings:
- Observe:
recent `observations`.
- Store:
stored observations / vectorized collisions / memory documents.
- Predict:
distinct live collisions.
- Outcome:
accepted vs dismissed outcomes.
- Adapt:
learned owners / `learned_from` edges / `team_model.ownership`.
Activity stream source:
- `engineer_states.gitUpdatedAt` -> editing/git state.
- `collisions.detectedAt` -> collision.
- `interventions.createdAt` -> warns/intervention.
- `outcomes.recordedAt` -> outcome.
- accepted real outcome -> learned_from/adapt event.
Cap activity rows to 8-10.
## 10. Suggested Implementation Plan
1. Create a new branch from the current graph branch or latest `main`.
2. Read:
- this file.
- `claude_team_memory_redesign.md`.
- `docs/graph.md`.
- `docs/live-ui-spec.md` if present.
- `frontend/src/components/GraphView.tsx`.
- `backend/src/graph/live.ts`.
3. Add graph UI subcomponents:
- `MetricsRail`.
- `GraphCanvas`.
- `LearningLoopRail`.
- `ActivityStream`.
- `SelectedNodePanel`.
4. Implement the dynamic graph canvas first.
5. Add the learning-loop rail and activity stream.
6. Polish interaction states:
- hover.
- selected node.
- selected edge/path.
- empty/live-loading/offline.
7. Verify with local live data.
8. Capture screenshots at desktop and narrow widths.
9. Run:
- `pnpm build` or local `vite build`.
- `tsc` for shared/backend/frontend.
10. Open a PR. Do not push directly to `main`.
## 11. Acceptance Criteria
The redesign is acceptable only when:
- The default view is readable in 10 seconds.
- The graph is dynamic, not static columns.
- It includes metrics, graph, learning-loop rail, activity stream, selected-node
panel, and legend.
- The primary risk path is obvious.
- Labels do not overlap in the default view.
- Whole graph mode exists but can be visually denser.
- It uses real data from the materializer.
- It remains composed from light shadcn/ruixen primitives where possible.
- It builds successfully.
- It is verified after a hard refresh because the PWA can cache stale bundles.
## 12. What Not To Do
- Do not make a marketing page.
- Do not make a generic dashboard.
- Do not rewrite the backend materializer unless the UI needs a small additive
field.
- Do not return to the dark UI wholesale.
- Do not keep the static column layout as the final answer.
- Do not show raw full paths as always-on canvas labels.
- Do not show every edge at equal opacity.
- Do not hide the learning loop in copy only; it needs a visible rail or panel.
## 13. Demo Script The UI Should Support
The final UI should support this story:
1. Engineer A and Engineer B work in the same repo.
2. One has unpushed changes.
3. PodMan observes the overlap.
4. A collision node appears and pulses.
5. PodMan sends a sync PR / warning intervention.
6. The intervention diamond appears.
7. The team accepts.
8. The outcome appears in the activity stream.
9. A `learned_from` edge appears or pulses.
10. The learning-loop rail advances to Adapt.
That is the recursive self-improvement moment. Everything else is supporting
evidence.
## 14. Open Questions For The Implementer
- Should dynamic layout be `d3-force` or animated layered SVG?
- Should loop/activity be added to `PodGraph` or exposed as separate endpoints?
- Should the demo seed one intact accepted outcome chain?
- Should `Whole graph` be hidden behind an explicit "inspect full graph" affordance?
- Should mobile show a simplified activity-first version instead of the full graph?
Answer these in code comments or PR notes when implementing.
## 15. Final Reminder
The backend now has real graph glue. The UI needs to become a **live learning
observatory**, not a static graph dump.
Make the light version earn the same reaction as the dark Bauhaus mock:
```text
I can see what happened.
I can see what PodMan did.
I can see what it learned.
```
-730
View File
@@ -1,730 +0,0 @@
# PodMan - Canonical Master Plan
> Source of truth for PodMan product intent, current implementation truth, demo
> strategy, public interfaces, risks, sponsor story, and next build order.
>
> If this file conflicts with `README.md`, `docs/idea.md`, `docs/livekit.md`,
> `docs/gemini.md`, `docs/mongodb.md`, `docs/digitalocean.md`,
> `docs/demo-setup.md`, or `docs/superpowers/specs/*`, follow this file and
> treat the older docs as reference material to reconcile later.
---
## 1. Product thesis
**PodMan sees active work before it becomes visible to GitHub, remembers how the
team works, researches better paths in the background, and coordinates teammates
without being intrusive.**
GitHub knows pushed branches, PRs, issues, and comments. It cannot see the most
expensive coordination failures while they are still forming on laptops: two
engineers editing the same unpushed file, someone blocked on an endpoint a
teammate is nearly done with, duplicated work starting silently, or a team
walking into a dead-end implementation path.
PodMan puts engineers in a consented LiveKit pod, watches live IDE/screen
context, fuses that with scheduled local git reports, GitHub state, MongoDB team
memory, and background research, then routes only useful interventions through
Hermes. The default is a small visual card. Hermes can message teammates when
the team needs coordination. Voice is reserved for urgent escalation.
**One-line product definition:** PodMan is a non-intrusive, continual-learning
team assistant for active coding.
**One-line demo promise:** PodMan notices live work, finds a better path,
remembers a previous intervention, and escalates only when the team actually
needs it.
---
## 2. Product contract
### Inputs
- **Live IDE/screen context:** engineers join a LiveKit room and publish screen
share so the backend agent can sample real work in progress.
- **Scheduled local git state:** each laptop should report dirty files,
unpushed commits, branch, and latest commit about every minute. This is the
deterministic fallback for facts vision cannot reliably infer.
- **MongoDB team memory:** ownership, current tasks, blockers, repeated
mistakes, preferred tools, decisions, intervention history, and outcomes.
- **GitHub repo state:** public repo metadata, branches, PR artifacts, and
issue/PR state when it exists.
- **Background research signals:** tool, repo, skill, package, docs, and
dead-end evidence discovered while teammates are working.
### Outputs
- **Default:** small visual intervention card in the PodMan frontend.
- **Coordination:** Hermes message to the right teammate(s) or project channel.
- **Urgent escalation:** voice only when timing or risk justifies interruption.
- **Action path:** optional sync PR, research recommendation, summary, fix
suggestion, or teammate notification.
### Memory rules
- Remember team-level work patterns, not raw screen recordings.
- Store structured observations, collisions, interventions, outcomes, and pod
state.
- Add exact-signature recall before vector recall: normalized file, symbol,
engineer pair, event type, and accepted/dismissed outcome.
- Privacy must stay explicit: engineers consent by joining the pod and sharing
screen context; do not store raw screenshots, full recordings, or secrets.
### Non-goals
- Not a dashboard as the product center.
- Not a screenshot analyzer with no action loop.
- Not sponsor-padding; every sponsor technology must be load-bearing or clearly
marked as optional polish.
- Not a task manager, Slack clone, full auth system, or general surveillance
tool.
---
## 3. Track fit: Continual Learning
PodMan fits **Continual Learning** because the system gets more useful from team
history and intervention outcomes.
- **Team model:** observations build ownership, hotspot, blocker, tool, and
decision memory per pod.
- **Outcome loop:** accepted, dismissed, and confirmed interventions become
supervision for future thresholds and routing.
- **Session compounding:** a later similar situation should reference prior
memory, choose a better action sooner, or lower the noise level.
- **Visible demo proof:** the first intervention writes memory; the second
similar situation retrieves it and says, in effect, "I have seen this pattern
before."
The learning proof should not depend on Atlas Vector Search being finished.
Exact MongoDB recall is enough for the MVP learning beat.
---
## 4. Current implementation truth
Verified on `2026-06-27` from local repo inspection, authenticated `gh`, and
the current remote plan commit.
### GitHub state
- Repo: <https://github.com/karti-ai/podman>
- Visibility: public
- Default branch: `main`
- Current local branch: `main`
- Local branch state during this rewrite: behind `origin/main` by two commits
- Issues: none
- PRs: none
- `origin/main` latest relevant commits:
- `8271188 feat(frontend): live room view, beat connectivity test, session resume`
- `65a0791 docs(plan): audit server state + mark tasks 1-5 done, reflect actual arch`
### Working / started
- Monorepo packages exist: `frontend`, `backend`, `shared`, `database`, and
`infra`.
- Backend is split into two processes:
- API service in `backend/src/server.ts`.
- LiveKit agent worker in `backend/src/agent.ts`.
- Backend API exposes:
- `GET /health`
- `POST /api/token`
- `POST /api/sync-pr`
- `POST /api/outcome`
- `GET /api/memory/stats`
- `GET /api/pods`
- `POST /api/pods`
- `GET /api/pods/:id`
- `PATCH /api/pods/:id`
- `DELETE /api/pods/:id`
- `POST /api/pods/:id/members`
- `DELETE /api/pods/:id/members/:name`
- Remote API health check returned `{"ok":true}` at
`http://165.22.129.249:8787/health` during verification.
- The LiveKit agent uses `@livekit/rtc-node` to join as `podman-agent`, subscribe
to `TrackSource.SOURCE_SCREENSHARE`, sample frames near 1 fps, convert frames
to RGBA, and encode downscaled JPEGs with `sharp`.
- Gemini vision is wired in `backend/src/vision/gemini.ts` with JSON structured
output, response schema, low media resolution, and model ID from env.
- Collision detection exists and groups engineer contexts by normalized file,
then fires when 2+ engineers touch the same file and at least one unpushed or
dirty signal exists.
- Shared LiveKit data topic and wire messages exist:
- topic: `podman.intervention`
- messages: `COLLISION`, `VOICE_CUE`, `ACK`, `GIT_REPORT`
- MongoDB persistence groundwork exists for observations, collisions,
interventions, outcomes, and pods.
- Frontend has pod selection, pod join, post-join pod view, LiveKit join helper,
and dev-mode fallback.
- `origin/main` adds live room participants, active-speaker state, session
resume, a "Play beat" audio connectivity test, and a deliberate "Share my
screen" button that publishes with `Track.Source.ScreenShare`. Merge that
remote commit before doing more frontend work on the local checkout.
- DigitalOcean infra scaffolding exists:
- `infra/.do/app.yaml` is the split App Platform direction.
- `infra/app.yaml` is an older single-service backend spec and should be
treated as legacy until reconciled.
### Server snapshot
From the remote plan snapshot and health check on `2026-06-27`:
- Backend API: running on `http://165.22.129.249:8787` and `/health` returned
`{"ok":true}`.
- Frontend: reported running on `:81`; port `80` was already taken.
- Agent worker: reported not running; it still needs LiveKit credentials and
`pnpm --filter @podman/backend dev:agent`.
- Treat this as operational evidence, not architecture truth. Reverify before
demo.
### Partial / completed since the original audit
- `backend/src/voice/live.ts` now publishes a `VOICE_CUE` fallback and attempts
Gemini audio publication into LiveKit. The agent only calls it for critical
interventions so voice remains an urgent escalation path.
- Hermes now has a data-channel teammate message path via `HERMES_MESSAGE` on
the existing `podman.intervention` topic. This is the MVP notification bridge,
not a Slack/Discord integration.
- `backend/src/memory/vectors.ts` implements exact-signature recall first and
can use Voyage/Gemini embeddings with Atlas Vector Search when configured.
- Exact-signature recall now attaches prior interventions/outcomes and prefers
accepted real collisions, giving the learning beat deterministic MongoDB
proof before vector search.
- `backend/src/memory/policy.ts` now uses severity, per-pod cooldown, and prior
outcome history. It is still a simple policy, not a trained threshold model.
- `POST /api/sync-pr` now creates a visible Markdown sync artifact commit before
opening the PR.
- Frontend `PodView` renders intervention cards, Hermes messages, voice cues,
and the accepted sync PR artifact link.
- Browser screen publishing exists, but the active join path must be proven to
tag tracks as screen share so the backend agent can filter them correctly. The
`origin/main` screen-share button appears to address this; local code remains
behind until that commit is merged.
- `GIT_REPORT` exists in shared types and agent handling. `scripts/podman-agent.mjs`
is the finished per-laptop git sidecar — polls every 15 s, upserts git fields
to `engineer_states` collection. The backend agent now fuses those Mongo
git-state fields into live contexts before collision detection; direct
LiveKit `GIT_REPORT` publication from the sidecar remains optional.
- Background research recommendations are a product requirement and demo goal,
not an implemented research agent yet.
- Deployment reliability is partial; API health is reachable, but API/static
site/worker together must still be reverified before demo.
- Env docs now align on `gemini-3.5-flash` for vision and
`gemini-3.1-flash-tts-preview` for voice. The backend still preserves a Gemini
Live path for future available Live models.
### Not yet proven
- Real browser -> LiveKit room -> backend agent screen-frame capture end to end.
- Real Gemini inference from a live shared IDE frame using the stage key/model.
- Real data-channel intervention card rendering in the active frontend.
- Hermes message routing to teammates.
- Voice escalation heard by participants through LiveKit.
- A meaningful real sync PR flow with correct GitHub scopes and artifact.
- Atlas Vector Search / Voyage recall path.
- DigitalOcean static site + API service + LiveKit agent worker all running
together.
- Background research recommendation that is both timely and evidence-backed.
---
## 5. Architecture to build toward
```
Engineer browser PWA
- joins a pod room
- publishes screen share and optional mic
- receives intervention cards and voice
|
v
LiveKit room
- one room per pod
- screen-share tracks are the live work signal
- small reliable data packets carry interventions
|
v
PodMan backend agent worker
- @livekit/rtc-node room participant
- screen-track subscription
- frame throttle and JPEG encode
- Gemini structured vision
- scheduled GIT_REPORT fusion
- GitHub state fusion
- collision, blocker, duplicate-work, and dead-end detection
- MongoDB memory recall and policy
|
v
Hermes action layer
- visual card routing
- teammate messages
- urgent voice escalation
- optional research/action/sync PR workflows
|
v
Backend API + MongoDB + GitHub
- token minting, pod CRUD, outcomes, memory stats
- observations, collisions, interventions, outcomes, pod memory
- public repo state and PR artifacts
```
The backend must remain split:
- **API service:** routable HTTP process with `/api/*` endpoints and health
checks.
- **Agent worker:** outbound LiveKit participant with no HTTP health-check port
requirement.
This split matters for DigitalOcean App Platform: the LiveKit agent should be a
worker, not a web service that App Platform expects to health-check over HTTP.
---
## 6. Public interfaces to preserve
Do not rename or reshape these without updating frontend, backend, docs, and demo
scripts together.
### Backend HTTP
- `GET /health`
- `POST /api/token`
- `POST /api/sync-pr`
- `POST /api/outcome`
- `GET /api/memory/stats`
- `GET /api/pods`
- `POST /api/pods`
- `GET /api/pods/:id`
- `PATCH /api/pods/:id`
- `DELETE /api/pods/:id`
- `POST /api/pods/:id/members`
- `DELETE /api/pods/:id/members/:name`
### LiveKit data channel
- Topic: `podman.intervention`
- Core messages:
- `COLLISION`: agent -> PWA; contains `collision` and `intervention`.
- `ACK`: PWA -> agent/API; intervention response.
- `GIT_REPORT`: local git sidecar -> agent; dirty/unpushed ground truth.
- `VOICE_CUE`: text cue/fallback for voice escalation.
### Required environment
```bash
LIVEKIT_URL=
LIVEKIT_API_KEY=
LIVEKIT_API_SECRET=
GEMINI_API_KEY=
GEMINI_VISION_MODEL=
GEMINI_LIVE_MODEL=
GITHUB_TOKEN=
GITHUB_REPO=karti-ai/podman
MONGODB_URI=
VOYAGE_API_KEY=
POD_ROOM=demo-pod
PORT=8787
VITE_BACKEND_URL=http://localhost:8787
VITE_LIVEKIT_URL=
```
Keep all non-`VITE_` secrets server-side.
---
## 7. Critical implementation callouts
### LiveKit
- Screen share is a video track. The backend agent should consume raw screen
frames through `@livekit/rtc-node`.
- The agent must filter screen share, not webcam:
`pub.source === TrackSource.SOURCE_SCREENSHARE`.
- Frontend publishing must tag the track as screen share; otherwise the agent can
miss it.
- Throttle aggressively. Screens can arrive near video frame rate; Gemini should
receive sampled frames only.
- Keep reliable data packets small. Use them for intervention metadata, not
screenshots, large diffs, or research dumps. Treat reliable payloads as
roughly 15 KiB max.
- A historical closed `livekit/node-sdks` issue reported high memory use when
consuming video; run memory checks during agent frame tests and stop if the
loop leaks.
### Gemini
- Use structured output for vision: JSON mime type plus response schema.
- Use low media resolution for ambient screen watching; reserve higher
resolution for debugging or targeted inspection.
- Never expose `GEMINI_API_KEY` to the browser.
- Gemini Live API is still a risk for the first demo path. Use card + Hermes
message first; add browser TTS or pre-generated voice fallback before relying
on Gemini Live for stage audio.
- Keep model IDs in env so preview/availability changes do not require code
changes.
### MongoDB
- Local MongoDB is fine for dev CRUD and memory counts.
- Atlas or Atlas Local is needed for the sponsor-grade Vector Search story.
- Build exact-signature recall first:
normalized file + symbol + engineer pair + event type + outcome.
- Writes from the agent should be best-effort. Mongo hiccups should degrade
memory, not kill live detection.
- Do not store raw screenshots or recordings.
### GitHub
- The repo is public and currently has no issue/PR backlog, so do not make the
plan issue-driven yet.
- GitHub cannot see local dirty files or unpushed commits. That is still a core
product moat.
- Sync PRs should use deterministic GitHub REST/Octokit flows, not browser
automation.
- Verify token scopes and demo repo permissions before stage time.
### DigitalOcean
- Use App Platform as:
- static site for frontend,
- HTTP service for API,
- worker for the LiveKit agent.
- Do not model the agent worker as a health-checked HTTP service.
- Keep a local and recorded fallback even if deployment works; venue network is a
stage risk.
### Hermes
- Treat Hermes as the action and messaging layer, not as a replacement for the
current implemented backend agent until code changes make that real.
- Hermes should choose the least intrusive channel:
card -> message -> voice.
- Hermes can own research summaries, teammate notification, sync PR initiation,
and urgent escalation once those workflows exist.
---
## 8. Build ladder
Do not mark a rung done until it is proven in logs, UI, or a visible external
artifact.
### P0 - make the live loop undeniable
1. **Preserve and reconcile the plan**
- Merge local `docs/PLAN.md` with `origin/main:docs/PLAN.md`.
- Keep both the broad product thesis and concrete server/current-state facts.
- After the docs are safe, merge or rebase the two newer `origin/main` commits
before implementing frontend work.
2. **Browser publish proof**
- Start backend API and frontend.
- Join a real LiveKit room from the browser.
- Confirm the browser publishes a screen-share track with the correct source.
3. **Agent frame proof**
- Start `pnpm --filter @podman/backend dev:agent`.
- Confirm room join, screen-track subscription, frame sampling, and JPEG
encode logs.
- Watch process memory while consuming frames.
4. **Gemini vision proof**
- Send one live sampled IDE frame to Gemini.
- Log parsed JSON with `currentFile`, `currentSymbol`, `activity`,
`hasUnpushedChanges`, and `confidence`.
- Add a confidence/logging gate if noisy frames cause bad reads.
5. **Scheduled git truth** ✅ partial
- `scripts/podman-agent.mjs` polls every 15 s: `git status --short`,
`git diff --stat HEAD`, `git log --oneline -1`, `git branch --show-current`.
- Upserts `changedFiles`, `diffStat`, `recentCommit`, `branch`, `gitUpdatedAt`
to `engineer_states` collection in MongoDB (upsert by `podId::name` key).
- **Still needed:** fuse `engineer_states` git fields into the collision
detector, and/or publish `GIT_REPORT` data channel messages so the agent
worker can incorporate git truth into vision-based decisions.
6. **Intervention card + Hermes notification**
- Publish a real intervention on `podman.intervention`.
- Render it as a small card in the frontend.
- Route a Hermes message to the affected teammate(s) or project channel once
the bridge exists.
7. **Background research recommendation**
- When the team is heading into a poor tool/repo/skill choice or dead end,
produce a recommendation card with short evidence.
- Minimum evidence: why it matters, what to use instead, and who should act.
8. **Learning proof**
- First intervention writes observation/collision/recommendation/outcome
memory.
- Second similar situation retrieves exact prior memory and changes the
message: "I have seen this pattern before."
9. **Urgency routing**
- Default to card.
- Escalate to Hermes message when coordination involves other teammates.
- Escalate to voice only when urgent.
10. **Action artifact**
- If demo uses same-file collision, click the card to open a real draft sync
PR or visible GitHub artifact.
- If demo uses research recommendation, show the accepted recommendation and
memory outcome instead.
11. **Deployment or fallback proof**
- Prove API/static/worker deployment together, or explicitly run local with a
recorded backup.
- Keep backup video on a separate device.
### P1 - polish the money moment
- Add visible live inference captions in the PWA.
- Add a small memory stats panel backed by `/api/memory/stats`.
- Add browser-side TTS or pre-generated voice fallback for urgent interventions.
- Add Hermes notification bridge once the target channel is chosen.
- Improve research cards with compatibility, install effort, docs quality, repo
health, and security/trust signals.
### P2 - sponsor and scale polish
- Implement Voyage embedding + Atlas Vector Search recall.
- Improve policy learning from outcomes.
- Deploy DigitalOcean static site + API service + worker as the submission path.
- Add optional GitHub issue/PR backlog integration after issues/PRs actually
exist.
### Cut if behind
- Webcam grid.
- Mic transcription.
- Full auth/accounts.
- Slack/Linear/Jira integrations unless Hermes requires one immediately.
- Complex dashboards.
- Server-published audio if browser/pre-generated voice proves escalation.
- Vector Search if exact Mongo recall demonstrates the learning beat.
---
## 9. Critical 3-minute demo script
**Rule:** open on one active IDE, not a grid. PodMan is an agent, not a
dashboard.
1. **0:00 - Set the scene**
- One engineer is actively coding in the IDE.
- The presenter says: "This work is not pushed yet. GitHub cannot see it."
2. **0:20 - Show the live signal**
- Show a compact caption: current file, inferred task, git dirty/unpushed
state.
- Show that PodMan is watching consented screen context, not stored
recordings.
3. **0:40 - Introduce the better-tool moment**
- A teammate starts down a weak path: wrong package, dead repo, bad API,
duplicated effort, or risky implementation.
- PodMan has been researching in the background.
4. **1:05 - Money moment**
- PodMan shows a small card:
"This path is likely a dead end. Use X instead; it matches our stack and is
actively maintained."
- The card names the affected teammate and the suggested action.
5. **1:25 - Hermes coordination**
- Hermes notifies the right teammate(s), not the whole room.
- No voice yet unless the situation is urgent.
6. **1:50 - Learning beat**
- A similar issue appears.
- PodMan references memory:
"I have seen this pattern before. Last time the team accepted the X
recommendation."
- Show `/api/memory/stats` or the visible memory indicator.
7. **2:20 - Urgency escalation**
- Raise the severity with a same-file collision, blocking dependency, failing
test, or imminent bad push.
- Hermes escalates to voice only now.
8. **2:40 - Close**
- Show the public repo, deployed/local URL, and memory stats.
- Closing line: "PodMan coordinates work while it is still happening."
### Reliable fallback demo
If the research recommendation is not reliable by stage time, use the same-file
collision fallback:
1. Two engineers open the same visible file.
2. `GIT_REPORT` or vision marks one as dirty/unpushed.
3. Agent publishes `COLLISION` on `podman.intervention`.
4. Frontend renders the card.
5. The card opens a sync PR artifact.
6. A second similar collision retrieves prior memory.
---
## 10. Sponsor strategy
### Gemini
Gemini must be load-bearing for the vision loop:
- live IDE/screen frame -> structured work context,
- optional message/recommendation generation,
- optional Live voice only after card/Hermes routing is stable.
Do not overclaim voice if it is using browser/pre-generated TTS. Say plainly that
it is the reliability fallback.
### LiveKit
LiveKit is the real-time spine:
- engineers join one pod room,
- screen-share tracks carry active work context,
- PodMan joins as a participant,
- data packets carry interventions,
- voice can be added as urgent escalation.
Pitch line: "Unpushed work is invisible to GitHub, so real-time presence is the
only way to coordinate before the push."
### MongoDB + Voyage
MongoDB is the learning proof:
- observations, collisions, recommendations, interventions, and outcomes persist,
- prior memory changes a later intervention,
- exact recall is the MVP,
- Voyage + Atlas Vector Search is the stronger sponsor-grade version after exact
recall works.
### DigitalOcean
DigitalOcean earns its place when:
- frontend runs as a static site,
- API runs as an HTTP service,
- LiveKit agent runs as a worker,
- public URL is shown in submission or demo.
Local fallback is acceptable for stage reliability, but the submission should
include the deployment URL if possible.
---
## 11. Risks and mitigations
| Risk | Mitigation |
| -------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------- |
| Looks like a dashboard | Keep the UI quiet. Hero is card/message/action, not a grid. |
| Looks like a screenshot analyzer | Always show screen signal + git truth + memory + action. |
| Interrupts too much | Default to cards, escalate to Hermes messages, reserve voice for urgency. |
| Overclaims implemented features | Mark voice, Hermes bridge, vectors, adaptive policy, research agent, real sync PR, and DO worker deploy incomplete until proven. |
| Vision misses unpushed state | Use scheduled `GIT_REPORT` for deterministic dirty/unpushed truth. |
| Research recommendation lacks evidence | Show only concise evidence: stack fit, repo/tool health, install effort, docs/trust signal. |
| No visible learning | Build exact Mongo recall before vector search. |
| LiveKit frame loop leaks memory | Monitor agent memory during video consumption; throttle hard. |
| GitHub issue/PR backlog absent | Do not invent issue-driven backlog; repo currently has no issues or PRs. |
| Venue network failure | Rehearse on hotspot and keep recorded backup. |
| DO worker deploy hangs | Deploy agent as worker, not health-checked service. |
---
## 12. Documentation reconciliation tasks
After this plan is accepted, update the supporting docs so they stop conflicting
with this file:
- `README.md`: replace POST-screenshot-first language with LiveKit screen-track
agent architecture and Hermes action-layer wording.
- `docs/idea.md`: broaden from blocker/dependency voice demo to card/message/
urgent-voice coordination plus research and memory.
- `docs/livekit.md`: remove "Hermes does NOT subscribe to engineer screen
tracks"; current architecture uses backend agent screen subscription.
- `docs/gemini.md`: keep structured vision, but mark Gemini Live as P1 and avoid
claiming voice is implemented.
- `docs/mongodb.md`: align collection names with current code
(`observations`, `collisions`, `interventions`, `outcomes`, `pods`) and add
exact-signature recall.
- `docs/digitalocean.md`: split API service and agent worker; do not deploy the
worker as a health-checked HTTP service; mark `infra/app.yaml` legacy or
reconcile it with `infra/.do/app.yaml`.
- `docs/demo-setup.md`: update the script to include better-tool research,
learning recall, Hermes notification, and urgency-based voice.
---
## 13. Acceptance checklist
Before saying PodMan is demo-ready:
- [ ] `pnpm format:check` passes or all failures are documented as unrelated.
- [ ] `pnpm typecheck` passes.
- [ ] Browser joins a real LiveKit room.
- [ ] Browser publishes a screen-share track with the correct source.
- [ ] Backend agent subscribes to the screen-share track.
- [ ] Agent logs at least one parsed Gemini context from a real IDE screen.
- [x] Local git report supplies dirty/unpushed truth on a schedule (`scripts/podman-agent.mjs` — 15 s poll → MongoDB `engineer_states`). Agent fusion still needed.
- [x] Frontend renders a real intervention card.
- [x] Hermes notification path works for teammate messages over the LiveKit data
channel.
- [ ] Voice is heard only for urgent escalation or a fallback is declared.
- [x] Outcome ACK writes to MongoDB and updates intervention status.
- [x] `/api/memory/stats` shows counts increasing.
- [x] Second similar situation uses prior exact memory in the message.
- [ ] Research recommendation card is evidence-backed, or fallback collision demo
is used.
- [x] Sync PR action creates a visible GitHub artifact if used in demo.
- [ ] DigitalOcean deployment or local fallback is rehearsed.
- [ ] Backup recording is ready on a separate device.
---
## 14. Evidence appendix
### Repo and GitHub state
- Public repo: <https://github.com/karti-ai/podman>
- Verified with authenticated `gh` on `2026-06-27`.
- Default branch: `main`.
- No GitHub issues or PRs existed at verification time.
### Hackathon / event
- AI Engineer World's Fair: <https://www.ai.engineer/worldsfair/2026>
- Cerebral Valley hackathon page:
<https://cerebralvalley.ai/e/aiewf-hackathon-2026>
### LiveKit
- Screen share docs: <https://docs.livekit.io/transport/media/screenshare/>
- Data packets docs: <https://docs.livekit.io/transport/data/packets/>
- Node SDK reference: <https://docs.livekit.io/reference/client-sdk-node/>
- Node SDK releases: <https://github.com/livekit/node-sdks/releases>
- Node SDK issue risk: <https://github.com/livekit/node-sdks/issues/444>
### Gemini
- Structured output:
<https://ai.google.dev/gemini-api/docs/structured-output>
- Media resolution: <https://ai.google.dev/gemini-api/docs/media-resolution>
- Live API: <https://ai.google.dev/gemini-api/docs/live-api>
### DigitalOcean
- App Platform app spec:
<https://docs.digitalocean.com/products/app-platform/reference/app-spec/>
### MongoDB
- Vector Search index type:
<https://www.mongodb.com/docs/vector-search/index/vector-search-type/>
- Node driver Atlas Vector Search:
<https://www.mongodb.com/docs/drivers/node/current/atlas-vector-search/>
-89
View File
@@ -1,89 +0,0 @@
# Agent Learning Plan
Status: draft
Goal: ship a visible recursive self-improvement loop without overbuilding
## Must-Have
1. Store agent runs.
2. Store trace summaries.
3. Store active and candidate strategy versions.
4. Attach verifier or outcome evidence.
5. Show one strategy improvement in the demo narrative.
## Build Order
### R1: Trace the run
Write one `agent_runs` record for an important coordination decision and append
trace events for:
- observation
- recall
- prediction
- intervention
- outcome
- adaptation
### R2: Version the strategy
Create an active strategy version for one of:
- collision detector threshold
- intervention routing
- graph discovery filter
- card wording prompt
### R3: Score the outcome
Use the simplest verifier:
- accepted real collision = useful
- dismissed = noisy
- no response after cooldown = uncertain
### R4: Propose a narrow change
Examples:
- "For this exact signature, prefer sync PR card."
- "For dismissed docs-only overlaps, suppress voice escalation."
- "For repeated auth.ts collisions, raise severity."
### R5: Promote or reject
Promote only when evidence is strong enough. Otherwise keep the candidate as
rejected or open.
## Demo Path
1. Show baseline strategy.
2. Trigger a collision.
3. Accept or dismiss the intervention.
4. Store outcome.
5. Show a candidate strategy update.
6. Promote it.
7. Trigger a similar event.
8. Show changed behavior.
## Nice-to-Have
- Strategy comparison panel.
- Model-generated prompt patch with verifier.
- Vector recall over strategy history.
- Rollback UI.
## Cut
- Full autonomous code rewriting.
- Multi-agent strategy debates.
- Long-term benchmark suite.
- Training a model.
## Acceptance Criteria
- The demo can point to a MongoDB record proving the agent changed behavior.
- The changed behavior is visible.
- The strategy has a parent and evidence.
- Rejected or failed changes are not deleted.
-84
View File
@@ -1,84 +0,0 @@
# Agent Learning Policy
Status: draft
Scope: guardrails for recursive self-improvement
## Prime Rule
PodMan may improve its agent behavior only when the improvement is narrow,
evidence-backed, versioned, and reversible.
## Allowed Learning
PodMan may learn:
- Which prompt version produces clearer interventions.
- Which detector threshold reduces false positives.
- Which routing channel gets accepted without being intrusive.
- Which verifier best predicts user acceptance.
- Which graph-discovery rule produces cleaner risk paths.
## Disallowed Learning
PodMan must not:
- Promote a strategy because the model says it is better.
- Rewrite broad system behavior from one example.
- Hide failures, dismissals, or rejected candidates.
- Learn from raw screenshots, secrets, or private terminal content.
- Turn voice into the default route.
- Create irreversible actions without human approval.
## Promotion Rules
A candidate strategy can become active only when all are true:
1. It has a parent strategy version.
2. It describes one concrete behavior change.
3. It has a verifier plan.
4. It has evidence from a run, outcome, or test.
5. It improves or fixes the target metric.
6. It does not increase user interruption without payoff.
## Rejection Rules
Reject and retain the candidate when:
- The verifier regresses.
- The change is too broad.
- The evidence is missing.
- The candidate conflicts with privacy rules.
- The candidate makes the demo less stable.
## Evidence Strength
| Evidence | Strength | Use |
| --- | --- | --- |
| Model opinion | Weak | Proposal only |
| Trace observation | Medium | Candidate rationale |
| Human accepted outcome | Strong | Promotion candidate |
| Human dismissed outcome | Strong | Suppression or rejection |
| Automated verifier | Strong | Promotion or rejection |
| Repeated accepted exact signature | Strong | Policy confidence increase |
## Versioning Rules
- Strategy versions are immutable after promotion or rejection.
- There is one active version per `podId + kind`.
- A rollback activates the previous version; it does not edit history.
- Parent-child lineage must be preserved.
## Safety Rules
- Store summaries, not raw sensitive content.
- Prefer deterministic checks over model judgment.
- Use exact MongoDB recall before vector recall.
- Ask for approval before changing code or data with external effects.
- Treat hackathon demo stability as a hard constraint.
## Demo Honesty
Seeded strategy versions are acceptable when labeled as demo-backed. Do not claim
a strategy was learned live unless a run and outcome actually created the
promotion evidence.
-74
View File
@@ -1,74 +0,0 @@
# Agent Learning Prompt
Use this prompt for an agent responsible for improving PodMan's own behavior.
## Prompt
You are PodMan's agent-learning evaluator.
Your job is to inspect a completed agent run, identify one narrow improvement,
define how to verify it, and decide whether to propose, promote, or reject a
strategy change.
You must not claim improvement without evidence. You must not propose broad
rewrites. Keep every change small, reversible, and tied to a run or outcome.
## Inputs
- Current active strategy version.
- Agent run summary.
- Trace events.
- Intervention outcome.
- Verifier result.
- Recent false positives or accepted events.
- Current demo constraints.
## Procedure
1. Identify the target behavior.
2. Identify the failure or success evidence.
3. Decide whether a strategy change is warranted.
4. Propose one narrow change.
5. Define the verifier.
6. Decide status: no change, candidate, promote, reject.
7. Write a short explanation suitable for the Team memory activity stream.
## Output Format
```text
Target
- Strategy kind:
- Active version:
- Behavior under review:
Evidence
- Run:
- Outcome:
- Verifier:
- Confidence:
Decision
- Status:
- Proposed change:
- Why this is narrow:
- Risk:
Verifier
- Metric:
- Passing condition:
- Failing condition:
Memory Write
- Collection:
- Record summary:
- Graph/activity summary:
```
## Hard Rules
- Exact outcomes beat model opinion.
- Rejected candidates stay in memory.
- No raw screenshots or secrets.
- No broad policy change from one weak signal.
- No voice-first behavior.
-185
View File
@@ -1,185 +0,0 @@
# Agent Learning Spec
Status: draft
Scope: how PodMan agents improve their own prompts, policies, detectors, and routing behavior
Owner: agent learning / recursive self-improvement
## Purpose
Agent learning is the recursive self-improvement layer. It is not the same as
team memory. Team memory learns about engineers and work. Agent learning learns
which agent strategies produce better outcomes.
The demo claim:
1. PodMan tries a coordination strategy.
2. The run is traced in MongoDB.
3. A verifier or human outcome scores it.
4. Gemini or another agent proposes a narrow strategy change.
5. The new strategy is versioned.
6. A later run uses the improved strategy and shows a better result.
## Core Objects
### Agent run
One attempt to execute a goal.
```text
agent_runs
runId
podId
goal
trigger
strategyVersionId
status
startedAt
completedAt
score
verifierSummary
inputRefs
outputRefs
```
Allowed `status` values:
```text
running, succeeded, failed, improved, regressed, abandoned
```
### Trace event
Append-only event log for a run.
```text
agent_trace_events
runId
podId
step
phase
eventType
inputSummary
outputSummary
toolName
error
metrics
createdAt
```
### Strategy version
Versioned prompt, detector rule, policy, verifier, or routing strategy.
```text
strategy_versions
strategyVersionId
podId
kind
name
parentVersionId
status
summary
promptText
policy
verifier
metrics
createdAt
promotedAt
```
Allowed `kind` values:
```text
prompt, policy, detector, verifier, routing
```
Allowed `status` values:
```text
candidate, active, retired, rejected
```
### Learning proposal
A candidate change before promotion.
```text
learning_proposals
proposalId
podId
sourceRunId
targetKind
parentVersionId
proposedChange
rationale
verifierPlan
status
createdAt
resolvedAt
```
Allowed `status` values:
```text
open, accepted, rejected, superseded
```
## MongoDB Indexes
| Collection | Index | Purpose |
| --- | --- | --- |
| `agent_runs` | `{ podId: 1, startedAt: -1 }` | Recent run history |
| `agent_runs` | `{ podId: 1, strategyVersionId: 1 }` | Compare strategy performance |
| `agent_trace_events` | `{ runId: 1, step: 1 }` | Reconstruct run |
| `strategy_versions` | `{ podId: 1, kind: 1, status: 1 }` | Find active strategy |
| `strategy_versions` | `{ podId: 1, createdAt: -1 }` | Version history |
| `learning_proposals` | `{ podId: 1, status: 1 }` | Open candidate changes |
## Learning Loop
```text
observe run -> score run -> propose change -> test candidate -> promote or reject
```
Agent learning must always connect these records:
```text
agent_run -> trace_events -> verifier result -> learning_proposal -> strategy_version
```
## Verifier Contract
Every promoted strategy needs a verifier signal.
Allowed verifier types:
- Human accepted or dismissed outcome.
- Test pass or fail result.
- Reduced false positive rate.
- Reduced intervention count with same or better accepted outcomes.
- Faster successful run.
- Better graph discovery precision.
- Explicit demo operator approval.
Self-evaluation alone is not enough to promote a strategy.
## Relationship to Team Graph
Agent learning can appear in the Team memory graph as activity and loop status,
but it should not clutter the main risk graph by default.
Graph discovery may show:
- `agent_run` activity in the stream.
- `strategy_versions` count in the learning loop.
- A selected-node detail saying a policy changed because a prior outcome was
dismissed or accepted.
## Acceptance Criteria
- Every strategy change has a parent.
- Every promoted strategy cites evidence.
- Rejected strategies are retained with a reason.
- Agent traces are append-only.
- The system can answer: "What changed, why, and did it help?"
+7 -3
View File
@@ -1,8 +1,9 @@
# Continual-Learning Graph Spec
> Owner: graph data + visualization. Status: demo-backed (live `team_model` reads land later).
> Owner: graph data + visualization. Status: demo-backed / active.
> Satisfies the documentation-first gate for the `backend/src/graph/*` and
> `frontend/src/components/GraphView.tsx` files.
> `frontend/src/components/GraphView.tsx` files. This file is the canonical
> graph spec.
## What this is (and is NOT)
@@ -26,7 +27,9 @@ The graph lives in two places, both keyed by `podId`:
{ podId, graph: PodGraph, updatedAt }
```
`GET /api/pods/:podId/graph` returns `team_model.graph`, or a demo graph when none exists yet.
`GET /api/pods/:podId/graph` returns the live materialized graph first, then
`team_model.graph`, then a labeled demo graph when neither live nor seeded
data exists.
2. **Normalized (for traversal):** the same nodes/edges are mirrored into two collections so
the model can be walked with MongoDB `$graphLookup` (the graph-database pattern):
@@ -95,6 +98,7 @@ Additive routes in `backend/src/server.ts` (shared file — additive only).
| `collisions` | **collision** nodes; `collides` (eng→col) + `touches` (file→col) |
| `interventions` | **intervention** nodes; `warns` (col→intervention) |
| `outcomes` | `learned_from` (intervention→owner) on accepted; flips nodes to `learned` |
| `suppressions` | `suppressed` activity beats — a dismissed signature recurred and PodMan stayed quiet (negative-feedback made visible; written at repeat time by the agent) |
Metrics (learned owners / open risk paths / accept rate) are live counts.
-69
View File
@@ -1,69 +0,0 @@
# Continual Learning Plan
Status: draft
Goal: prove PodMan learns from outcomes in the hackathon demo
## Must-Have Demo Loop
1. Observe two engineers touching the same file.
2. Store the observation and git state in MongoDB.
3. Predict a collision.
4. Send a card or Hermes message.
5. Record accept or dismiss outcome.
6. Adapt `team_model`.
7. Show the learned graph edge or changed future behavior.
## Build Order
### R1: Make exact recall reliable
- Normalize file paths.
- Build stable memory signatures.
- Look up prior accepted and dismissed outcomes.
- Prefer exact recall over vector recall.
### R2: Make outcomes update memory
- Accepted real collision creates or strengthens ownership.
- Accepted real collision creates `learned_from`.
- Dismissed outcome lowers confidence or suppresses route.
### R3: Expose loop data to the graph
- Add optional loop snapshot.
- Add optional activity stream.
- Keep existing `PodGraph` fields stable.
### R4: Show the observatory
- Render observe/store/predict/outcome/adapt.
- Show recent activity.
- Make selected-node detail explain why memory changed.
### R5: Prepare a clean demo chain
- Ensure one collision -> intervention -> accepted outcome exists.
- Ensure repeated signature recalls prior memory.
- Verify graph shows learned ownership.
## Nice-to-Have
- Atlas Vector Search over memory summaries.
- Confidence scoring per ownership edge.
- Per-file memory timeline.
- Strategy promotion tied to outcomes.
## Cut
- Raw screenshot storage.
- Full autonomous training.
- Broad dashboard metrics.
- Multi-pod learning generalization.
## Acceptance Criteria
- A judge can see what changed in memory.
- The second similar event behaves differently.
- Exact MongoDB records prove the loop.
- The graph remains legible with real data.
-97
View File
@@ -1,97 +0,0 @@
# Continual Learning Policy
Status: draft
Scope: what PodMan may learn about a team
## Prime Rule
PodMan learns coordination patterns, not personal surveillance profiles.
## Allowed Memory
PodMan may store:
- File and symbol ownership.
- Active file overlap.
- Repeated collision signatures.
- Intervention history.
- Accepted and dismissed outcomes.
- Routing preferences by event type and severity.
- Summaries of decisions relevant to future coordination.
## Forbidden Memory
PodMan must not store:
- Raw screenshots.
- Screen recordings.
- Secrets or credentials.
- Full terminal logs.
- Personal performance judgments.
- Private content unrelated to the coding task.
## Evidence Policy
| Evidence | Can predict? | Can adapt memory? |
| --- | --- | --- |
| Vision only | Yes, low confidence | No |
| Git watcher | Yes | No, unless repeated |
| GitHub state | Yes | No, unless verified |
| Accepted real outcome | Yes | Yes |
| Dismissed outcome | Yes, for suppression | Yes, as negative signal |
| Verifier result | Yes | Yes |
## Intervention Policy
Use the least intrusive channel:
1. Watch quietly.
2. Card.
3. Hermes message.
4. Voice.
Voice is only for urgent, high-confidence, time-sensitive risks.
## Adaptation Policy
Allowed adaptations:
- Add learned ownership after accepted real outcome.
- Raise confidence for repeated accepted signatures.
- Lower confidence for dismissed signatures.
- Prefer the previously accepted intervention kind.
- Suppress repeated low-value warnings.
Disallowed adaptations:
- Broad threshold changes from one example.
- Treating vector similarity as proof.
- Hiding dismissals.
- Making interruption more aggressive without evidence.
## Retention Policy
Keep:
- Outcomes.
- Signatures.
- Team model memory.
- Strategy metrics.
Summarize or expire:
- Old observations.
- Low-confidence vision-only events.
- Detailed trace text.
Delete immediately:
- Secrets.
- Accidental raw sensitive captures.
## Demo Policy
Seeded data is acceptable only if the demo script is honest about it. Live
learning requires a live or staged outcome write that visibly updates the graph
or future decision.
-87
View File
@@ -1,87 +0,0 @@
# Continual Learning Prompt
Use this prompt for the agent that decides what PodMan should remember from a
coordination event.
## Prompt
You are PodMan's continual-learning memory agent.
Your job is to inspect observations, collisions, interventions, and outcomes,
then decide what team memory should be updated. You must separate observed
facts, inferred risks, human outcomes, and durable learned memory.
Do not claim something was learned unless an accepted real outcome, verifier, or
human label supports it.
## Inputs
- Pod id.
- Recent engineer states.
- Recent observations.
- Candidate collision.
- Prior exact-signature memory.
- Intervention record.
- Outcome record.
- Current team model.
## Procedure
1. Normalize file and symbol.
2. Build exact signature.
3. Check prior accepted and dismissed outcomes.
4. Classify the current event.
5. Decide whether memory should change.
6. Emit the graph impact.
7. Write a short explanation.
## Output Format
```text
Event
- Signature:
- Engineers:
- File:
- Symbol:
- Evidence:
Prior Memory
- Accepted matches:
- Dismissed matches:
- Ownership:
Decision
- Memory action:
- Confidence:
- Reason:
Graph Impact
- Nodes:
- Edges:
- Activity text:
Safety
- Sensitive data present:
- Redaction needed:
```
## Memory Actions
Allowed actions:
- no_change
- strengthen_signature
- weaken_signature
- create_learned_owner
- update_route_preference
- suppress_signature
- request_human_label
## Hard Rules
- Exact recall before vector recall.
- Dismissals are learning signals.
- `learned_from` requires accepted real outcome.
- Store summaries, not raw screen content.
- Prefer less intrusive future behavior when uncertain.
-217
View File
@@ -1,217 +0,0 @@
# Continual Learning Spec
Status: draft
Scope: how PodMan learns team memory from live work and outcomes
Owner: continual learning / Team memory
## Purpose
Continual learning is the product proof that PodMan gets more useful from use.
It learns team-level coordination memory: ownership, repeated collisions,
accepted interventions, dismissed noise, and preferred routing.
The visible loop:
```text
observe -> store -> predict -> outcome -> adapt
```
## Source Collections
### `engineer_states`
Latest per-engineer state from vision and local git.
Key fields:
- `podId`
- `name`
- `currentFile`
- `changedFiles`
- `branch`
- `confidence`
- `visionUpdatedAt`
- `gitUpdatedAt`
- `updatedAt`
### `observations`
Structured perception events.
Key fields:
- `podId`
- `engineerId`
- `currentFile`
- `symbol`
- `activity`
- `confidence`
- `observedAt`
### `collisions`
Predicted risk events.
Key fields:
- `id`
- `podId`
- `file`
- `symbol`
- `engineers`
- `severity`
- `status`
- `memorySignature`
- `detectedAt`
### `interventions`
Actions PodMan sent or suggested.
Key fields:
- `id`
- `podId`
- `collisionId`
- `kind`
- `channel`
- `message`
- `suggestedAction`
- `createdAt`
### `outcomes`
Human or verifier supervision.
Key fields:
- `id`
- `podId`
- `interventionId`
- `collisionId`
- `accepted`
- `wasRealCollision`
- `learnedOwner`
- `recordedAt`
### `team_model`
Durable pod memory.
Key fields:
- `podId`
- `graph`
- `ownership`
- `collisionSignatures`
- `interventionPolicy`
- `updatedAt`
### `memory_vectors`
Optional semantic recall. Exact recall comes first.
Key fields:
- `podId`
- `sourceKind`
- `sourceId`
- `text`
- `embedding`
- `embeddingModel`
- `tags`
## Learning Rules
### Observe
Write structured evidence from vision, git, GitHub, and agent traces.
### Store
Persist source records and materialized summaries. Do not store raw screenshots
or recordings.
### Predict
Create a collision when multiple engineers converge on the same normalized file
or symbol and at least one signal shows active or unpushed work.
### Outcome
Record whether the intervention was accepted, dismissed, real, or false.
### Adapt
Only accepted real outcomes can create `learned_from` graph edges. Dismissals
adapt suppression, routing, or confidence.
## Exact Signature
Use deterministic signatures:
```text
podId:eventType:normalizedFile:symbol:sortedEngineers
```
Rules:
- Sort engineer names.
- Normalize file paths.
- Use `*` for missing symbol.
- Never include timestamps.
## UI-Facing Loop Snapshot
The graph response may include:
```text
loop
activeStep
steps[]
key
label
value
detail
status
```
Step mapping:
| Step | Source |
| --- | --- |
| Observe | recent observations and git updates |
| Store | team model, graph records, memory vectors |
| Predict | open collisions |
| Outcome | accepted and dismissed outcomes |
| Adapt | learned owners, learned edges, strategy changes |
## Activity Stream
The graph response may include:
```text
activity[]
id
at
kind
title
detail
nodeId
edgeId
```
Allowed `kind` values:
```text
editing, collision, intervention, outcome, learned, agent
```
## Acceptance Criteria
- The system can show one accepted outcome changing future memory.
- Exact recall works without vector search.
- The Team memory graph can explain the learning loop.
- Dismissals and false positives are retained.
- The demo does not rely on raw screenshots or hidden state.
+56
View File
@@ -0,0 +1,56 @@
# PodMan — Demo Runbook (final, read-off)
> One page. Reflects the **current live state** (`podman.live`) as of the final hour.
> Companion to `docs/demo.md` (full 4-min script) — this is the de-risked version.
## ⛔ Read first — last-hour changes that affect the demo
1. **Suppression is DISABLED for the demo** (`shouldIntervene` only filters `info`).
Every real collision now surfaces — good (no accidental silencing). **But the
`docs/demo.md` 2:10 beat "it learned to stay QUIET / no card fires" will NOT
work.** Do the **escalate** half only (below). Do not promise silence on stage.
2. **Do NOT redeploy prod.** The live build has demo code not on `main`
(`userPodContext`); redeploying `main` would remove it *and* surface stale
"SUPPRESSED" beats. If the auto-deploy timer is on, stop it before the run:
`systemctl stop podman-hermes-sync-deploy.timer`.
3. **Never debug on stage.** Narrate, fall back to the backup recording, keep moving.
## ✅ Pre-flight (do once, before you start)
- [ ] `curl https://podman.live/health``{"ok":true}`
- [ ] Deploy timer stopped (above); confirm exactly one agent: `systemctl status podman-platform-agent`
- [ ] Browser open on the pod; **click "Enable audio"** in-room (TTS needs it)
- [ ] Backup recording cued on a second device
- [ ] Two screen-shares live (the two "engineers")
## 🎬 The 3-minute path (the coherent, working story)
| Time | Do | Say (land the value) |
|---|---|---|
| **0:00** | Pod view, two live screen tiles | "Coordination is the bottleneck now, not coding. PodMan watches everyone's work live — no one has to ask 'what are you on?'" |
| **0:30** | Point at tiles; show activity stream filling | "Real screen-shares over **LiveKit**; each frame → **Gemini Vision** returns file/symbol/activity — a perception layer, not a chatbot." |
| **1:05** | Two engineers edit the **same file**, unpushed | "GitHub can't see this — nothing's pushed. We fuse live screen context with **local git truth**." → **collision card** appears. |
| **1:40** | **Accept** the card | Real **GitHub sync-PR** is created + outcome recorded. "One tap closes the loop." |
| **2:10** | **Trigger the same collision again** | Card now says **"Seen before."** and escalates straight to the **Gemini TTS voice** cue. "It recalled the prior event from **MongoDB Atlas vector search** and escalated — that's the learning." |
| **2:35** | Cut to **GraphView** | "This graph is materialized **live from MongoDB right now** — and here's the **`learned_from` edge**: PodMan learned who owns this file from the accepted outcome. Not a mock." |
| **3:00** | Stack + close | "All on **DigitalOcean**, systemd-supervised; ambient score is **Gemini Lyria**. Coordination awareness, collisions caught before they cost an afternoon, memory that sharpens each session." |
## 🗣️ Sponsor coverage (say each ≥ once)
- **Gemini** — Vision (0:30), TTS voice (2:10), Lyria (3:00) — add Live conversation if time allows
- **LiveKit** — "real screen-shares over LiveKit" (0:30), TTS over LiveKit (2:10)
- **MongoDB** — "Atlas vector-search recall" + "graph materialized live from MongoDB" (2:10, 2:35)
- **DigitalOcean** — "all on DigitalOcean, systemd-supervised" (3:00)
## 🚑 If it breaks
| Failure | Recovery |
|---|---|
| Voice doesn't fire | Cut to the card, say the line aloud — cards are the default path |
| Collision won't trigger | Use the backup recording for that beat; keep narrating |
| Graph looks empty/odd | Reload once; if still off, narrate from the recording |
| Agent silent | Suspect a duplicate `podman-hermes` process, not the UI — but **don't fix on stage** |
## Optional (only if ahead of time): the five-minute-meeting killer
Open the live voice conversation; ask *"PodMan, what is everyone working on, and where is the collision detector?"* → answers via real tool calls (**Gemini Live API**). Skip entirely if tight on time or if it drops.
---
**The spine:** live awareness → collision caught pre-push → accept → **"Seen before." + voice + `learned_from` edge**. That's the win. Suppression/"stays quiet" is intentionally out this run.
-108
View File
@@ -1,108 +0,0 @@
# Demo Setup
Pre-stage checklist for the 3-minute live demo. Do this on all 3 laptops before walking on stage.
---
## Before demo day
- [ ] `demo-pod` room created in LiveKit Cloud dashboard
- [ ] Hermes deployed on DO (or confirmed running locally as fallback)
- [ ] MongoDB Atlas cluster running, `MONGODB_URI` set in Hermes env
- [ ] All `.env` vars populated and verified via `GET /health` returning `{ ok: true }`
- [ ] Record a backup video of the full demo working end-to-end
- [ ] Rehearse the demo script 3× with real audio
---
## Laptop setup (all 3 machines)
### Editor settings
- Font size: **18pt or larger** — Gemini Vision must read file names and code
- Single editor window — no split panes, no overlapping terminals
- File tab visible with full file name shown (not truncated)
- Light or dark theme is fine — avoid low-contrast themes
### Browser
- Chrome (best `getDisplayMedia` support)
- PWA tab open and joined to `demo-pod`
- Earbuds / headphones plugged in and tested
- Volume: medium — PodMan voice should be clearly audible but not startle
### Screen layout
- Editor takes 2/3 of screen
- Terminal takes bottom 1/3 (always visible)
- No other windows on top
---
## Demo file setup
Pre-create these files in the demo repo before the demo:
**Alice's machine:**
- Open `auth/middleware.ts` — has visible function stubs
- Terminal shows nothing running initially, then `Server running on :3001` at the right moment
**Bob's machine:**
- Open `frontend/login.tsx` — has visible form component code
- Terminal idle
**Carol's machine:**
- Open `frontend/integration.ts` or similar
- Terminal shows: `curl http://localhost:3001/auth``curl: (7) Failed to connect`
---
## Demo script timing
| Time | Action | Who |
| ----- | ----------------------------------------------- | ----------- |
| 0:00 | All three join `demo-pod` | All |
| 0:05 | PodMan greets by voice | Hermes auto |
| 0:20 | Alice opens `auth/middleware.ts`, starts typing | Alice |
| 0:45 | Bob opens `frontend/login.tsx` | Bob |
| 0:50 | Carol runs `curl` command, sees error | Carol |
| ~1:20 | BLOCKER_DETECTED nudge fires | Hermes auto |
| 1:50 | Alice starts her server (`node server.js`) | Alice |
| ~2:00 | DEPENDENCY_READY nudge fires | Hermes auto |
| 2:20 | Optional: show session 2 ownership warm-start | Presenter |
| 2:45 | Close | Presenter |
---
## Gemini Vision reliability tips
- Keep font at 18pt+ throughout the demo — do not zoom out
- Avoid opening file picker dialogs or overlapping modals during the demo
- File names in editor tabs must be fully visible (not `auth/middle...`)
- If Hermes logs show `confidence < 0.6` frames: bump font size, ensure file tab is clear
- Terminal output must be on a single line — avoid long stack traces during demo
---
## Cooldown note
Hermes has a 3-minute cooldown between nudges per pod. For the demo, if you need to trigger a second event quickly:
Option 1: restart Hermes between the two demo scenarios (resets cooldown state)
Option 2: set `NUDGE_COOLDOWN_MS=0` via env var during demo (add this override to Hermes)
---
## Fallback plan
If any system fails on stage:
1. **Hermes unreachable:** switch to local (`pnpm --filter backend dev`) — PWA auto-falls back to `localhost:8787`
2. **Gemini Vision low confidence:** presenter narrates what PodMan "saw" while playing the backup video
3. **LiveKit audio not working:** play backup video — show the nudge text cards on screen instead
4. **Full system failure:** play the backup recording, narrate the demo live
Always have the backup video on a separate device, not the same laptop running Hermes.
+156
View File
@@ -0,0 +1,156 @@
# PodMan — 4-Minute Demo Script
**Theme:** Continual Learning. **Hard limit:** 4:00. Practice to land at 3:45.
**The one-line story:** writing code isn't the bottleneck anymore — _coordinating
who's writing what_ is. PodMan is a pair programmer for the whole team: it watches
every member's work in real time, gives everyone live status without anyone having
to interrupt anyone, and learns your team's dynamics so it nudges less and helps
more over time.
**The hook to land:** a "quick five-minute question" actually costs ~25 minutes of
lost focus — for two people. PodMan removes the reason to ask. Multiply the saved
recovery time across every teammate, every day, and that is the value.
---
## The script (4:00)
### 0:000:30 — The problem + hook
> "AI made writing code easy. The thing still slowing teams down is coordination
> — checking each other's work, re-planning collisions, and the constant 'what
> are you working on?' A five-minute question really costs both people 25 minutes
> of lost focus. PodMan is a pair programmer for the whole team: it watches
> everyone's work live, so anyone can see another's status without interrupting
> them — and it learns your team as it goes."
_On screen:_ the pod view, two teammates joined, screen-share tiles live.
### 0:301:05 — Real-time team awareness (LiveKit + Gemini Vision)
- Point at the two live screen tiles. "These are real screen shares over
**LiveKit**. Our agent subscribes to the tracks and samples frames."
- "Each frame goes to **Gemini Vision**, which returns structured context — file,
symbol, activity — not a chatbot, a perception layer."
- Show the live activity stream filling in (Signals vs Reasoning sections).
- Land the value: "This is the part that replaces 'what are you working on?' —
every teammate's current work is just _visible_, in real time. Nobody had to
ask."
_Built-by-us callout:_ `backend/src/vision/gemini.ts`, the LiveKit agent worker.
### 1:051:40 — The catch (detection + first intervention)
- Have alice and bob both edit the **same file** with unpushed changes.
- "Normally nobody notices until merge time. GitHub can't see this — nothing's
pushed. Our detector fuses live screen context with **local git truth** from a
watcher on each laptop."
- A collision card appears: _"alice + bob both on detector.ts (unpushed)."_
- Let the **Gemini TTS** urgent voice fire once over LiveKit: _"alice and bob are
both editing detector.ts. Please sync before pushing."_
- Land the value: "That's a merge conflict and a wasted afternoon caught before it
happened — and neither of them had to be tracking the other."
_Built-by-us callout:_ `collision/detector.ts`, `action/hermes.ts`,
`voice/live.ts`.
### 1:402:10 — Cross-channel overlap (research + code)
- Keep alice editing `livekit.py`.
- Have bob share a browser tab on LiveKit docs/SDK pages.
- A collaboration nudge appears: _"🤝 bob is researching LiveKit agents
(docs.livekit.io) while alice edits livekit.py — sync up before duplicating
effort."_
- Land the value: "This is not a merge conflict. PodMan caught duplicated effort
across channels — code on one screen, research on another — and nudged the team
before two people solved the same problem twice."
_Built-by-us callout:_ `vision/gemini.ts`, `collision/research.ts`,
`memory/vectors.ts`.
### 2:102:50 — Continual learning (the theme — the money shot)
This is the differentiator. Two beats, both from pre-seeded memory:
1. **It learned to stay quiet.** Trigger a pattern that was dismissed as a false
alarm earlier. "Last session a teammate marked this kind of alert as not a
real conflict. Watch — PodMan stays silent. No nagging." (No card fires.)
2. **It learned to escalate.** Trigger the real-conflict pattern that was
accepted before. The card now says **"Seen before."** and goes straight to
the spoken urgent cue.
- "The only input was one accept/dismiss tap. No retraining, no labeling. This is
**MongoDB Atlas vector search** recalling similar past events plus a policy
that adapts on the recalled outcome."
- Optional: show `/api/memory/stats` counts climbing — accumulated experience.
_Built-by-us callout:_ `memory/vectors.ts` ($vectorSearch), `memory/policy.ts`
(outcome-conditioned gate), `memory/store.ts`.
### 2:503:30 — The five-minute meeting, killed (Gemini Live API)
- Frame it: "Instead of breaking a teammate's focus to ask what they're up to,
you ask PodMan."
- Open the live voice conversation. Ask out loud: _"PodMan, what is everyone
working on, and where is the collision detector implemented?"_
- It answers with **real tool calls**`search_repo`, git history, current
collisions — not guesses.
- "This is the **Gemini Live API**, streaming speech-to-speech over LiveKit, with
custom function tools we wrote so it grounds every answer in the actual repo
and live state. That's the status sync, answered in seconds, with zero recovery
tax on anyone else."
_Built-by-us callout:_ `agents/podman-live-conversation/agent.py`.
### 3:303:50 — Stack + close
- "All on **DigitalOcean** — static frontend, API, and agent workers, supervised
by systemd. The ambient score is **Gemini Lyria** generated per pod through the
Interactions API."
- Close: "Engineering ability stopped being the bottleneck — coordination is.
PodMan gives a whole team real-time awareness without the interruptions, catches
collisions before they cost an afternoon, and learns each team's dynamics so it
helps more over time. Saved focus, multiplied across every teammate. That's
continual learning, shipped."
### 3:504:00 — Buffer / Q&A handoff
---
## Sponsor-prize coverage (say each at least once)
| Prize | Spoken moment | Segment |
| ---------------- | -------------------------------------------------------------------------------- | ---------------------- |
| **Gemini** | Vision perception, Live API agent w/ tools, TTS voice, Lyria score | 0:30, 1:05, 2:50, 3:30 |
| **LiveKit** | "real screen shares over LiveKit", agent subscribes, TTS audio track, live voice | 0:30, 1:05, 3:30 |
| **MongoDB** | "Atlas vector search recalling past events" | 1:50 |
| **DigitalOcean** | "all on DigitalOcean, systemd-supervised workers" | 3:30 |
---
## If something breaks (live recovery)
| Failure | Recovery |
| ----------------------- | ----------------------------------------------------------------------- |
| Voice doesn't fire | Cut to the card; say the line aloud; cards are the default path anyway. |
| Live conversation drops | Skip 2:503:30; lean longer on the learning beat. |
| Collision won't trigger | Use the backup recording for that beat; keep narrating. |
| Agent flapping | Pre-checked — but if so, `systemctl restart podman-platform-agent`. |
**Rule:** never debug on stage. Narrate, fall back to recording, keep moving.
---
## Tight timing summary
| Time | Beat |
| ---- | ---------------------------------------------------------- |
| 0:00 | Problem (coordination cost) + hook + original-work line |
| 0:30 | Real-time team awareness — LiveKit + Gemini Vision |
| 1:05 | The catch — collision caught before merge |
| 1:40 | Cross-channel overlap — research + code nudge |
| 2:10 | **Continual learning — quiet + escalate** |
| 2:50 | The five-minute meeting, killed — Gemini Live conversation |
| 3:30 | DigitalOcean + Lyria + close |
| 3:50 | Buffer |
+189
View File
@@ -0,0 +1,189 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<title>PodMan — Demo Runbook</title>
<style>
:root{
--bg:#0b0f17; --panel:#121826; --panel2:#0e1421; --ink:#e6edf6; --muted:#9fb0c3;
--line:#23304a; --blue:#5b9bff; --violet:#a78bfa; --red:#ff6b6b; --amber:#f4b942;
--green:#43d39e; --slate:#8aa0b8;
}
*{box-sizing:border-box}
body{margin:0;background:radial-gradient(1200px 600px at 70% -10%,#15203a 0,var(--bg) 55%);
color:var(--ink);font:15px/1.55 -apple-system,BlinkMacSystemFont,"Segoe UI",Roboto,Helvetica,Arial,sans-serif;}
.wrap{max-width:960px;margin:0 auto;padding:32px 22px 80px}
header{border-bottom:1px solid var(--line);padding-bottom:18px;margin-bottom:26px}
h1{font-size:30px;margin:0 0 6px;letter-spacing:-.3px}
h1 .dot{color:var(--green)}
.sub{color:var(--muted);font-size:15px;max-width:680px}
h2{font-size:13px;text-transform:uppercase;letter-spacing:.14em;color:var(--blue);
margin:34px 0 12px;border-left:3px solid var(--blue);padding-left:10px}
pre{background:var(--panel2);border:1px solid var(--line);border-radius:10px;
padding:16px 18px;overflow:auto;color:#cfe0f5;font:12.5px/1.45 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace}
.grid{display:grid;gap:14px}
.card{background:linear-gradient(180deg,var(--panel),var(--panel2));border:1px solid var(--line);
border-radius:12px;padding:16px 18px}
table{width:100%;border-collapse:collapse;font-size:14px}
th,td{text-align:left;padding:9px 10px;border-bottom:1px solid var(--line);vertical-align:top}
th{color:var(--muted);font-weight:600;font-size:12px;text-transform:uppercase;letter-spacing:.06em}
td.t{white-space:nowrap;color:var(--violet);font-variant-numeric:tabular-nums;font-weight:600}
.say{color:var(--green)}
.pill{display:inline-block;padding:1px 8px;border-radius:999px;font-size:11px;font-weight:700;letter-spacing:.04em}
.p-blue{background:rgba(91,155,255,.16);color:var(--blue)} .p-violet{background:rgba(167,139,250,.16);color:var(--violet)}
.p-green{background:rgba(67,211,158,.16);color:var(--green)} .p-amber{background:rgba(244,185,66,.16);color:var(--amber)}
.warn{background:rgba(255,107,107,.08);border:1px solid rgba(255,107,107,.4);border-radius:12px;padding:14px 18px}
.warn h3{margin:0 0 8px;color:var(--red);font-size:14px;letter-spacing:.02em}
ol{margin:0;padding-left:20px} ol li{margin:6px 0}
code{background:#0a1220;border:1px solid var(--line);border-radius:5px;padding:1px 6px;color:#bcd4f2;font-size:12.5px}
.muted{color:var(--muted)} .spine{color:var(--green);font-weight:600}
footer{margin-top:40px;color:var(--muted);font-size:12.5px;border-top:1px solid var(--line);padding-top:14px}
</style>
</head>
<body>
<div class="wrap">
<header>
<h1>PodMan <span class="dot"></span> Demo Runbook</h1>
<div class="sub">A pair programmer for the whole team. It watches everyone's work live, catches
collisions <em>before the push</em> (when GitHub still can't see them), and learns the team's
dynamics so it nudges less and helps more over time.</div>
</header>
<div class="card">
<strong>The 10-second story:</strong> <span class="muted">Coding isn't the bottleneck anymore —
<em>coordination</em> is. PodMan gives a team real-time awareness with zero interruptions, catches
the pre-push collision that costs an afternoon, and the Team-Memory Graph proves it learned —
materialized live from MongoDB, not a mock.</span>
</div>
<h2>How it works — architecture</h2>
<pre>
░ PodMan data flow ░
┌──────────────┐ screen-share track ┌────────────────────────────┐
│ Engineers │ ──────────────────────▶ │ LiveKit Cloud │
│ (browsers) │ ◀───────────────────────│ room = "demo-pod" │
└──────┬───────┘ cards · voice · cues └─────────────┬──────────────┘
│ (data channel) │ frames @ ~1 fps
│ live graph (poll 5s) │ + local git truth
│ ▼
│ ┌────────────────────────────────────┐
│ │ HERMES AGENT (Node, podman-hermes)│
│ │ frame ─▶ GEMINI VISION (file/ │
│ │ symbol/activity) │
│ │ fuse with git watcher (unpushed!) │
│ │ ─▶ detect same-file collision │
│ │ ─▶ Gemini TTS voice over LiveKit │
│ └──────────────────┬───────────────────┘
│ │ read / write
┌──────┴────────────┐ materialize live ┌───────────▼──────────────────┐
│ TEAM-MEMORY │ ◀──────────────────────│ MongoDB Atlas │
│ GRAPH (SVG) │ graph/live.ts │ observations · collisions │
│ owns/edits/ │ │ interventions · outcomes │
│ collides/ │ │ team_model · graph_nodes/edges│
│ learned_from │ │ + $vectorSearch recall │
└───────────────────┘ └───────────────────────────────┘
Infra: DigitalOcean droplet · systemd-supervised API + agent · Caddy static frontend
Models: Gemini (Vision · TTS · Live voice · embeddings · Lyria ambient score)
</pre>
<h2>How it works — the continual-learning loop</h2>
<pre>
observe ──▶ store ──▶ predict ──▶ outcome ──▶ adapt
│ │ │ │ │
Gemini MongoDB same-file accept / learned_from
Vision + records overlap dismiss edge + owner
git truth detected (1 tap) persists
▲ │
└───────────────── recall ($vectorSearch) ◀─────────┘
"Seen before." → escalates on a repeat
The proof: one accept/dismiss tap — no retraining, no labels — changes the
NEXT run. Run 2 of the same collision says "Seen before." and a violet
learned_from edge appears on the graph. That edge is read straight from Atlas.
</pre>
<h2>The 3-minute demo (read off this)</h2>
<table>
<thead><tr><th>Time</th><th>Do</th><th>Say — land the value</th></tr></thead>
<tbody>
<tr><td class="t">0:00</td><td>Pod view, two live screen tiles</td>
<td class="say">"Coordination is the bottleneck now, not coding. PodMan watches everyone's work live — nobody has to ask 'what are you on?'"</td></tr>
<tr><td class="t">0:30</td><td>Point at tiles; activity stream fills</td>
<td class="say">"Real screen-shares over <b>LiveKit</b>; each frame → <b>Gemini Vision</b> → file/symbol/activity. A perception layer, not a chatbot."</td></tr>
<tr><td class="t">1:05</td><td>Two engineers edit the <b>same file</b>, unpushed</td>
<td class="say">"GitHub can't see this — nothing's pushed. We fuse live screen context with <b>local git truth</b>." → collision card appears.</td></tr>
<tr><td class="t">1:40</td><td><b>Accept</b> the card</td>
<td class="say">Real <b>GitHub sync-PR</b> created + outcome recorded. "One tap closes the loop."</td></tr>
<tr><td class="t">2:10</td><td><b>Trigger the same collision again</b></td>
<td class="say">Card now says <b>"Seen before."</b> and escalates to the <b>Gemini TTS voice</b>. "It recalled the prior event from <b>MongoDB Atlas vector search</b> and escalated — that's the learning."</td></tr>
<tr><td class="t">2:35</td><td>Cut to <b>GraphView</b></td>
<td class="say">"This graph is materialized <b>live from MongoDB right now</b> — and here's the <b>learned_from edge</b>: PodMan learned who owns this file. Not a mock."</td></tr>
<tr><td class="t">3:00</td><td>Stack + close</td>
<td class="say">"All on <b>DigitalOcean</b>, systemd-supervised; ambient score is <b>Gemini Lyria</b>. Awareness, collisions caught pre-push, memory that sharpens each session."</td></tr>
</tbody>
</table>
<h2>Sponsor coverage — say each ≥ once</h2>
<table>
<thead><tr><th>Prize</th><th>Spoken moment</th><th>Beat</th></tr></thead>
<tbody>
<tr><td><span class="pill p-blue">Gemini</span></td><td>Vision perception · TTS voice · Lyria score (Live conversation if time)</td><td class="muted">0:30 · 2:10 · 3:00</td></tr>
<tr><td><span class="pill p-violet">LiveKit</span></td><td>"real screen-shares over LiveKit"; TTS audio over LiveKit</td><td class="muted">0:30 · 2:10</td></tr>
<tr><td><span class="pill p-green">MongoDB</span></td><td>"Atlas vector-search recall" · "graph materialized live from MongoDB"</td><td class="muted">2:10 · 2:35</td></tr>
<tr><td><span class="pill p-amber">DigitalOcean</span></td><td>"all on DigitalOcean, systemd-supervised workers"</td><td class="muted">3:00</td></tr>
</tbody>
</table>
<h2 style="color:var(--red);border-color:var(--red)">⛔ Last-hour caveats — read before you run</h2>
<div class="warn">
<h3>1 · Suppression is DISABLED for the demo</h3>
<div class="muted">Every collision now surfaces (good — no accidental silencing). The old
<em>"it learned to stay QUIET / no card fires"</em> beat will <b>not</b> work. Use the
<b>escalate</b> beat instead: trigger the same collision again → <b>"Seen before." + voice +
learned_from edge</b>. Don't promise silence on stage.</div>
</div>
<div class="warn" style="margin-top:12px">
<h3>2 · Do NOT redeploy prod</h3>
<div class="muted">The live build has demo code not on <code>main</code> (e.g. <code>userPodContext</code>);
redeploying would remove it <em>and</em> surface stale "SUPPRESSED" beats. If the auto-deploy timer is on,
stop it: <code>systemctl stop podman-hermes-sync-deploy.timer</code>.</div>
</div>
<div class="warn" style="margin-top:12px">
<h3>3 · Never debug on stage</h3>
<div class="muted">Narrate, fall back to the backup recording, keep moving.</div>
</div>
<h2>If it breaks — live recovery</h2>
<table>
<thead><tr><th>Failure</th><th>Recovery</th></tr></thead>
<tbody>
<tr><td>Voice doesn't fire</td><td>Cut to the card, say the line aloud — cards are the default path. (Click "Enable audio" in-room first!)</td></tr>
<tr><td>Collision won't trigger</td><td>Use the backup recording for that beat; keep narrating.</td></tr>
<tr><td>Graph looks empty / odd</td><td>Reload once; if still off, narrate from the recording.</td></tr>
<tr><td>Agent silent</td><td>Suspect a duplicate <code>podman-hermes</code> process, not the UI — but don't fix on stage.</td></tr>
</tbody>
</table>
<h2>Pre-flight checklist</h2>
<div class="card">
<ol>
<li><code>curl https://podman.live/health</code><code>{"ok":true}</code></li>
<li>Deploy timer stopped; exactly one agent — <code>systemctl status podman-platform-agent</code></li>
<li>Browser on the pod; <b>click "Enable audio"</b> (TTS needs it)</li>
<li>Backup recording cued on a second device</li>
<li>Two screen-shares live (the two "engineers")</li>
</ol>
</div>
<footer>
<div class="spine">The spine: live awareness → collision caught pre-push → accept → "Seen before." + voice + learned_from edge → "this graph is live from MongoDB."</div>
<div style="margin-top:6px">Runs on the <code>demo</code> branch (ahead of <code>main</code>). Suppression intentionally off this run. · Generated for the final demo.</div>
</footer>
</div>
</body>
</html>
+18 -4
View File
@@ -58,6 +58,13 @@ The mirror at `infra/.do/app.yaml` is kept identical for DO UI/import workflows.
- No HTTP route and no HTTP health check
- Default room: `POD_ROOM=demo-pod`
### Worker: live conversation agent (Python)
- Source: `agents/podman-live-conversation/`
- The Gemini Live voice agent (`gemini-3.1-flash-live-preview`), run with `uv`.
- On the droplet it runs as `podman-live-conversation-agent.service`
(`infra/systemd/`). It is separate from the Node services and the TS agent.
---
## Required Runtime Environment
@@ -66,10 +73,15 @@ The mirror at `infra/.do/app.yaml` is kept identical for DO UI/import workflows.
LIVEKIT_URL=wss://your-livekit-server.livekit.cloud
LIVEKIT_API_KEY=...
LIVEKIT_API_SECRET=...
LIVEKIT_CONVERSATION_AGENT_NAME=podman-live-conversation
GEMINI_API_KEY=...
GEMINI_API_KEY=... # GOOGLE_API_KEY also accepted
GEMINI_VISION_MODEL=gemini-2.0-flash
GEMINI_LIVE_MODEL=gemini-3.1-flash-tts-preview
GEMINI_LIVE_MODEL=gemini-3.1-flash-tts-preview # TTS voice
GEMINI_CONVERSATION_MODEL=gemini-3.1-flash-live-preview
GEMINI_EMBEDDING_MODEL=gemini-embedding-001
GEMINI_TTS_VOICE=Charon
# GEMINI_MUSIC_MODEL=lyria-3-clip-preview # optional override
GITHUB_TOKEN=...
GITHUB_REPO=karti-ai/podman
@@ -82,8 +94,10 @@ PORT=8787
POD_ROOM=demo-pod
```
`VOYAGE_API_KEY` is optional for local/demo fallback. Without it, Mongo exact
signature recall still works; Atlas Vector Search recall is skipped.
`VOYAGE_API_KEY` is optional. Without it, Gemini embeddings provide vector
recall; without any embedding provider, recall degrades to exact signature
matching. The Lyria background score uses the Gemini Interactions API and the
same `GEMINI_API_KEY`.
---
+87 -96
View File
@@ -1,137 +1,128 @@
# Gemini Integration Spec
PodMan uses Gemini for two distinct jobs: **vision** (understanding screens) and **voice** (speaking nudges).
Status: active / matches code.
PodMan uses Gemini for five jobs, all through the `@google/genai` SDK
(`GoogleGenAI`) with a single `GEMINI_API_KEY` (`GOOGLE_API_KEY` /
`GOOGLE_GENERATIVE_AI_API_KEY` also accepted):
1. **Vision** — turn screen frames into structured work context.
2. **Embeddings** — vector recall over past coordination events.
3. **TTS voice** — spoken urgent escalations over LiveKit.
4. **Live conversation** — a real-time voice agent teammates talk to.
5. **GenMedia (Lyria)** — a per-pod background score.
Collision detection and intervention text are **deterministic in code**, not
Gemini calls. PodMan does not ask Gemini "is this a conflict?" — that is decided
by `backend/src/collision/detector.ts` from fused vision + git truth. This is a
deliberate reliability choice for the live demo.
---
## 1. Vision — Screen Understanding
## 1. Vision — screen understanding
**Model:** `gemini-2.0-flash` (fast, cheap, strong multimodal)
**Model:** `GEMINI_VISION_MODEL` (default `gemini-2.0-flash`)
**Code:** `backend/src/vision/gemini.ts``analyzeFrame()`
**Trigger:** every 30s per active engineer, when Hermes receives a `POST /ingest` frame
**Trigger:** the LiveKit agent samples a JPEG frame from each engineer's
screen-share track (not an HTTP upload — frames arrive over LiveKit).
**Input:** base64-encoded JPEG, max 1280×720, ~5080KB after compression
**Input:** a single base64 JPEG, sampled at low media resolution.
**Prompt:**
```
You are analyzing a software engineer's screen during a coding session.
Extract the following JSON. If you cannot determine a field with confidence above 0.7, set it to null.
**Output:** structured JSON via `responseJsonSchema` (no markdown parsing):
```ts
{
"currentFile": "string | null", // active file visible in editor tab or title bar
"inferredTask": "string | null", // 1 sentence: what the engineer appears to be doing
"terminalVisible": true | false, // is a terminal or CLI panel visible
"recentTerminalOutput": "string | null", // last meaningful line of terminal output if visible
"confidence": 0.01.0 // your overall confidence in this extraction
mode: 'editing' | 'research', // browser/docs/SDK research vs editor work
currentFile: string, // open file path, e.g. src/auth/session.ts
currentSymbol: string, // function/class under the cursor
activity: string, // editing | reading | debugging | terminal | PR review
hasUnpushedChanges: boolean, // dirty git gutter / modified markers visible
researchTopic: string, // e.g. "LiveKit agents setup", for research mode
researchSource: string, // source domain, e.g. "docs.livekit.io"
confidence: number // 0..1
}
Respond with valid JSON only. No explanation. No markdown.
```
**Confidence gate:** if `confidence < 0.6`, Hermes discards the frame — no state update, no event detection triggered.
When a frame shows a browser/docs/SDK page instead of an editor, Gemini Vision
classifies it as `mode: "research"` and extracts the topic/source. That feeds the
cross-channel overlap detector: one teammate researching LiveKit docs while
another edits `livekit.py` becomes a collaboration nudge, not a merge-conflict
alert. Editor/IDE frames remain `mode: "editing"` and use the existing file,
symbol, activity, and dirty-change fields.
**Rate limit:** 1 call per engineer per 30s. With 3 engineers = 6 calls/min ≈ $0.002/min at Flash pricing.
**Latency/cost levers (in code):**
**Demo setup requirement:** editors must have large font (18pt+), single window, file name clearly visible in tab. This is the primary reliability lever.
- `thinkingConfig: { thinkingBudget: 0 }` — minimal thinking for the ambient loop.
- `mediaResolution: MEDIA_RESOLUTION_LOW` — smaller image tokens.
- Missing `confidence` defaults to `0.5`.
**Demo reliability:** large editor font, single window, visible file tab. This is
the primary lever for clean reads.
---
## 2. Event Detection — Coordination Awareness
## 2. Embeddings — semantic recall
**Model:** `gemini-2.0-flash` (text only, fast)
**Model:** `GEMINI_EMBEDDING_MODEL` (default `gemini-embedding-001`, 768 dims)
**Code:** `backend/src/memory/vectors.ts`
**Trigger:** after every successful state write to MongoDB, Hermes runs event detection over all active engineer contexts.
Each collision is embedded into a short memory text (`file`, `symbol`,
`engineers`, `severity`, unpushed flag) and stored on the `collisions` document.
On a new collision PodMan embeds the query and runs MongoDB Atlas `$vectorSearch`
(index `collision_embedding`) to recall similar past events and their outcomes.
**Input:** JSON snapshot of all engineers' current states + ownership map
**Prompt:**
```
You are a team coordination agent. Below is the current state of each engineer on the team.
Engineer states:
{{engineerStates}}
Ownership map (who owns which files):
{{ownershipMap}}
Detect if any of these coordination events are occurring:
- DEPENDENCY_READY: an engineer who was blocked or waiting now has what they need because another engineer completed relevant work
- BLOCKER_DETECTED: an engineer appears stuck (same file, error in terminal, no progress) and another teammate could help
- DUPLICATE_WORK: two or more engineers are working on the same file simultaneously
If an event is detected, respond with:
{
"event": "DEPENDENCY_READY" | "BLOCKER_DETECTED" | "DUPLICATE_WORK" | null,
"involvedEngineers": ["engineerId", ...],
"file": "string | null",
"reason": "1 sentence explanation"
}
If no event, respond with { "event": null }.
Respond with valid JSON only.
```
**Provider order:** Voyage (`VOYAGE_API_KEY`, `voyage-4-lite`) is tried first when
present; Gemini embeddings are the fallback. Without either, recall degrades to
exact signature/file matching — the demo still works.
---
## 3. Nudge Generation — Voice Message
## 3. TTS voice — urgent escalation over LiveKit
**Model:** `gemini-2.0-flash` (text only)
**Model:** `GEMINI_LIVE_MODEL` (default `gemini-3.1-flash-tts-preview`)
**Default voice:** `GEMINI_TTS_VOICE` (default `Charon`)
**Code:** `backend/src/voice/live.ts``speak()` / `speakInRoom()`
**Trigger:** when event detection returns a non-null event
**Input:** event type + engineer names + file + reason
**Prompt:**
```
You are PodMan, a friendly AI teammate. Generate a short spoken message (12 sentences max) to notify the team about this coordination event.
Event: {{eventType}}
Engineers involved: {{engineerNames}}
File: {{file}}
Context: {{reason}}
Rules:
- Use first names only
- Be direct and specific
- Do not use filler words
- Sound natural when spoken aloud
- Do not start with "Hey" or "Attention"
Respond with the message text only.
```
**Example output:**
> "Carol — Alice just got the auth endpoint running. You're clear to integrate."
Flow: a short, natural voice line is generated for a critical collision, returned
as audio, and published as a LiveKit microphone-source audio track. The track is
held for the audio duration plus tail/hold so browsers do not cut playout short.
Browser audio must be unlocked by a user gesture first. The frontend always
renders the `VOICE_CUE` text as a fallback. See `docs/livekit.md` for delivery.
---
## 4. Voice Output — Gemini TTS via LiveKit
## 4. Live conversation — real-time voice agent
**Model:** `gemini-3.1-flash-tts-preview`
**Model:** `GEMINI_CONVERSATION_MODEL` (default `gemini-3.1-flash-live-preview`)
**Code:** `agents/podman-live-conversation/agent.py` (Python LiveKit Agents,
`google.realtime.RealtimeModel`)
**Integration:** Hermes generates Gemini TTS audio and publishes it as a LiveKit audio track. The code still preserves a Gemini Live path for future available Live models.
A teammate can start a live, streaming speech-to-speech session with PodMan. The
agent answers using **function tools** rather than guessing, including:
**Flow:**
- `get_active_pod_context`, `get_recent_changes`, `search_team_memory`
- `search_repo`, `repo_recent_commits`, `repo_find_commits` (repo + git history)
- `record_conversation_note`
- `delegate_to_hermes`, `abort_active_hermes_job` (hands work to the async Hermes
job runner — see `docs/hermes.md`)
1. Nudge message text generated (step 3)
2. Hermes passes text to Gemini Live via LiveKit Agents
3. Gemini Live streams audio back in real-time
4. LiveKit publishes audio into the pod room
5. All participants hear it through their audio output
Started/stopped via `POST /api/pods/:id/live-conversation/start` and `.../stop`.
**Why Gemini Live (not plain TTS):**
---
- Streams audio directly — no intermediate WAV file conversion
- Latency ~300500ms from text to first audio packet
- Natural-sounding voice
- Strong prize story: Gemini Live 2.5 is the headline model
## 5. GenMedia — Lyria background score
**Model:** `lyria-3-clip-preview` (override with `GEMINI_MUSIC_MODEL`)
**Endpoint:** Gemini **Interactions API** (`/v1beta/interactions`)
**Code:** `backend/src/voice/music.ts`
A pod-specific ~30s clip is generated through the Interactions API, cached in
MongoDB, and served via `GET /api/pods/:id/music` to play as ambient room audio.
---
## Cooldown
Per-pod cooldown of **3 minutes** between nudges. Prevents spam if multiple events fire simultaneously. Implemented in Hermes, not in Gemini.
Per-pod cooldown (`NUDGE_COOLDOWN_MS`, default 180000 ms / 3 min) gates repeated
interventions. Implemented in `backend/src/memory/policy.ts`, not in Gemini.
-69
View File
@@ -1,69 +0,0 @@
# Graph Discovery Plan
Status: draft
Goal: make MongoDB graph discovery visible as a dynamic learning observatory
## Must-Have
1. Keep live materializer as source of graph truth.
2. Add optional loop and activity fields.
3. Build a dynamic graph layout.
4. Default to risk path.
5. Make selected-node detail explain the story.
## Build Order
### R1: Stabilize discovered graph
- Keep file and engineer noise filters.
- Keep collision collapse.
- Keep priority for accepted-outcome paths.
- Keep graph size capped.
### R2: Add observatory data
- Compute learning-loop snapshot.
- Compute activity stream.
- Preserve current graph contract.
### R3: Improve path selection
- Pick one primary risk path.
- Include learned path when present.
- Dim unrelated collisions and repeated interventions.
### R4: Render dynamically
- Use `d3-force` or animated layered layout.
- Make nodes draggable.
- Curve or bundle edges.
- Animate `learned_from`.
### R5: Verify with real data
- Fetch live `demo-pod` graph.
- Confirm labels do not collide badly.
- Confirm red edges do not dominate.
- Confirm activity and loop explain the graph.
## Nice-to-Have
- Reachability panel using `$graphLookup`.
- Hover path previews.
- Edge bundling by file or collision.
- Time scrubber for graph snapshots.
## Cut
- Generic analytics dashboard.
- Large graph database migration.
- Rendering every historical event.
- Static fixed-column final layout.
## Acceptance Criteria
- Risk path is obvious in 10 seconds.
- Learned path is visible when data exists.
- Whole graph mode exists but is not the default.
- The graph remains backed by MongoDB, not hardcoded mock data.
-83
View File
@@ -1,83 +0,0 @@
# Graph Discovery Policy
Status: draft
Scope: graph hygiene, evidence thresholds, and UI truthfulness
## Prime Rule
The graph must be sparse enough to explain the learning loop and truthful enough
to audit from MongoDB.
## Node Policy
Create nodes only when they add explanation value.
Allowed:
- Current engineers.
- Real files.
- Current or recent collisions.
- Interventions tied to surviving collisions.
- Learned ownership paths.
Avoid:
- Test engineers.
- Scratch files.
- URLs or environment values misread as files.
- Repeated identical intervention diamonds.
- Orphan nodes with no story value.
## Edge Policy
Edges need evidence.
| Edge | Required evidence |
| --- | --- |
| `editing` | observation or git state |
| `touches` | file involved in collision |
| `collides` | collision prediction |
| `warns` | intervention record |
| `learned_from` | accepted real outcome |
| `owns` | learned or configured ownership |
## De-Hairball Policy
Default mode must not show every relationship equally.
Rules:
- Default to risk path.
- Collapse repeated collision signatures.
- Cap files and collisions.
- Dim non-risk edges.
- Bundle or curve dense edges.
- Hide low-priority labels until hover or select.
- Prefer selected-node explanation over labels everywhere.
## Truthfulness Policy
- Do not show `learned_from` for orphaned or dismissed outcomes.
- Do not label vector similarity as learned memory.
- Do not show demo seed as live learning unless labeled.
- Do not hide false positives from activity or memory.
## Privacy Policy
Graph labels should not expose secrets, raw terminal output, or sensitive file
contents. File paths are acceptable when they are repo paths and not secret
values.
## Visual Policy
Semantic colors stay stable:
- Engineer: blue.
- File: slate.
- Feature: amber.
- Collision: red.
- Intervention: violet.
- Learned: violet dashed edge.
Chrome should use the app's light shadcn tokens.
-81
View File
@@ -1,81 +0,0 @@
# Graph Discovery Prompt
Use this prompt for an agent that materializes or reviews PodMan's Team memory
graph.
## Prompt
You are PodMan's graph discovery agent.
Your job is to turn MongoDB records into a sparse, truthful graph that explains
the continual-learning loop. Do not maximize node count. Maximize legibility and
evidence.
The default output should show the risk path and learned path, not every
possible edge.
## Inputs
- Pod id.
- Pod roster.
- Recent engineer states.
- Recent observations.
- Collisions.
- Interventions.
- Outcomes.
- Team model.
- Existing graph nodes and edges.
## Procedure
1. Normalize file paths.
2. Remove noise.
3. Create engineer and file nodes.
4. Collapse repeated collisions by signature.
5. Preserve accepted-outcome paths.
6. Create intervention nodes for surviving collisions.
7. Create learned edges only from accepted real outcomes.
8. Select the primary risk path.
9. Build activity and loop summaries.
10. Explain selected-node stories.
## Output Format
```text
Graph Summary
- Pod:
- Nodes:
- Edges:
- Primary risk path:
- Learned path:
Discovery Decisions
- Collapsed:
- Dropped as noise:
- Preserved because learned:
Loop
- Observe:
- Store:
- Predict:
- Outcome:
- Adapt:
Activity
- Recent events:
Risks
- Missing evidence:
- Potential hairball:
- Demo caveat:
```
## Hard Rules
- No `learned_from` without accepted real outcome.
- No raw screenshots or secrets in labels.
- Do not rewrite the backend materializer unless explicitly asked.
- Prefer additive graph fields.
- Default to risk path.
- Keep whole graph optional.
-146
View File
@@ -1,146 +0,0 @@
# Graph Discovery Spec
Status: draft
Scope: how PodMan discovers graph nodes, edges, risk paths, and learning paths from MongoDB
Owner: graph discovery / Team memory observatory
## Purpose
Graph discovery turns MongoDB memory into a legible Team memory graph. It is not
only layout. It decides which relationships matter, which path is highlighted,
and which evidence explains the graph.
The graph must answer:
1. Who is working?
2. Which files or symbols overlap?
3. Where is the risk?
4. What did PodMan do?
5. What outcome changed memory?
## Source Data
Graph discovery reads:
- `pods`
- `engineer_states`
- `observations`
- `collisions`
- `interventions`
- `outcomes`
- `team_model`
- `graph_nodes`
- `graph_edges`
- optional `memory_vectors`
- optional `agent_runs`
- optional `strategy_versions`
## UI Graph Contract
```text
PodGraph
podId
generatedAt
nodes
edges
metrics
loop?
activity?
```
Node kinds:
```text
engineer, feature, file, collision, intervention
```
Edge kinds:
```text
owns, editing, touches, collides, warns, learned_from
```
## Discovery Rules
### Engineer nodes
Create from pod roster, recent observations, git state, or collision membership.
### File nodes
Create only from normalized real file paths. Reject noise such as URLs, env
values, scratch names, and non-file strings.
### Collision nodes
Create from distinct collision signatures. Collapse repeats. Prioritize
collisions referenced by accepted outcomes.
### Intervention nodes
Create one visible intervention per surviving collision unless whole-graph mode
explicitly expands history.
### Learned paths
Create `learned_from` only when an accepted real outcome links an intervention
to a durable memory update.
## Path Modes
### Risk path
Default mode. Highlight the clearest current chain:
```text
engineer -> file -> collision -> intervention -> learned owner
```
Dim unrelated graph material.
### Learning edges
Highlight `learned_from`, `owns`, and the outcomes that produced them.
### Whole graph
Show all materialized nodes and edges with de-emphasized non-critical edges.
## MongoDB Traversal
Use `graph_edges` for reachability:
```text
source -> target -> next target
```
Primary traversal questions:
- What risks does this engineer reach?
- Which files feed this collision?
- Which intervention came from this collision?
- Which learned owner came from this intervention?
## Metrics
Minimum metrics:
- Learned owners.
- Open risk paths.
- Accept rate.
Optional metrics:
- Observations.
- Interventions.
- Memory vectors.
- Strategy versions.
## Acceptance Criteria
- Default graph is not a hairball.
- Every visible learned edge has outcome evidence.
- Every selected node can explain why it matters.
- Activity stream matches graph events.
- Graph can be rebuilt from MongoDB source records.
+108
View File
@@ -0,0 +1,108 @@
# Hermes Spec
Status: active / matches code.
"Hermes" is PodMan's **action layer** — the part that turns a detected problem
into something a teammate sees, hears, or gets done. It spans three things:
1. **Interventions** — cards, messages, and urgent voice in the pod room.
2. **Async jobs** — longer tasks delegated from the live conversation agent.
3. **Ops watchdog** — keeps the production services healthy.
The LiveKit identity for the main agent is `podman-hermes`.
---
## 1. Interventions
**Code:** `backend/src/agent/podman.ts`, `backend/src/action/hermes.ts`,
`backend/src/voice/live.ts`.
When the agent detects a collision, it runs the learning loop (recall → policy
gate; see `docs/cont_learning.md`) and then publishes the **least intrusive**
intervention that fits:
- **Card / message** — a data-channel packet on the `podman.intervention` topic
(`publishHermesIntervention` / `publishHermesMessage`). Default path.
- **Urgent voice** — only for `critical` collisions. `speak()` generates Gemini
TTS audio and publishes it as a LiveKit audio track.
- **Research overlap nudge** — a collaboration card when one engineer is editing
a file while another is researching the same topic in docs/browser context.
This uses `suggestedAction.kind = "ping_teammate"` and is spoken once for the
demo beat, but it is explicitly **not** a merge conflict.
Intervention text is short and deterministic (template, not an LLM call):
`Conflict: alice + bob both on detector.ts (unpushed). Seen before.` The spoken
line is phrased for natural TTS prosody. Each intervention is persisted to the
`interventions` collection; the teammate's accept/dismiss returns via
`POST /api/outcome`.
Research-overlap text is also deterministic:
`🤝 bob is researching LiveKit agents (docs.livekit.io) while alice edits livekit.py — sync up before duplicating effort.`
A per-pod cooldown (`NUDGE_COOLDOWN_MS`, default 3 min) and a single-shot
"active conflict" guard prevent repeat nagging; a conflict re-arms once it
resolves.
---
## 2. Async Hermes jobs
**Code:** `backend/src/hermes/jobs.ts`. **Storage:** `hermes_jobs` +
`hermes_job_events` (see `docs/mongodb.md`).
The live conversation agent can hand a longer task to Hermes via its
`delegate_to_hermes` tool. Lifecycle:
```
queued → running → (waiting_for_confirmation) → completed | aborted | failed
```
`createHermesJob()` records the job, emits an `accepted` event, and kicks off
`runHermesJob()` in the background. The runner gathers context and runs scoped,
read-mostly steps based on the prompt and success criteria:
- always: `git status --short --branch`, `git diff --stat`
- if the ask mentions GitHub: a repo reachability check via the GitHub API
- if it mentions Mongo/memory/telemetry: collection counts
- if it mentions build/test/typecheck/broken: `pnpm typecheck`
**Confirmation gate:** if `riskLevel === 'deploy_allowed'` and
`requiresConfirmation`, the job parks at `waiting_for_confirmation` instead of
acting. **Abort:** `abortHermesJob()` signals the runner's `AbortController`.
Every step appends a `hermes_job_event` (redacted + truncated), which is both
stored and published live to the room as a `HERMES_JOB_EVENT` data message from a
short-lived `podman-hermes-job-*` identity. The conversation UI streams these via
`GET /api/.../hermes-job/events/stream`.
**Endpoints:** `POST /api/internal/hermes/jobs`,
`GET /api/internal/hermes/jobs/:jobId`, `.../abort`, `.../events`,
`.../events/stream`, plus the pod-scoped `.../live-conversation/:sessionId/hermes-job`.
---
## 3. Ops watchdog
**Code:** `scripts/hermes-watchdog.mjs`, `scripts/hermes-sync-deploy.mjs`,
`scripts/hermes-notify.mjs`. **Detail:** `docs/digitalocean.md`.
systemd supervises the app processes; Hermes owns the loop around them:
- `pnpm hermes:watchdog` checks systemd services, public routes, `/health`,
`/api/pods`, and `pnpm deploy:doctor`. Failures trigger targeted restarts.
- `podman-hermes-watchdog.timer` runs it every 5 minutes.
- `podman-hermes-sync-deploy.timer` polls `origin/main` every 2 minutes and, on a
clean tree, fast-forwards, builds, publishes `frontend/dist`, restarts
API/agent/Caddy, and runs the strict watchdog.
- Reports go to `/var/log/podman/hermes-watchdog-latest.json`; set
`PODMAN_ALERT_WEBHOOK_URL` to forward failures to Discord/Slack/webhook.
---
## What Hermes is NOT
- Not an autonomous code-writing agent. Job steps are scoped, read-mostly checks;
deploy-level actions require explicit confirmation.
- Not a second collision detector. Detection is deterministic
(`collision/detector.ts`); Hermes only acts on the result.
-93
View File
@@ -1,93 +0,0 @@
# PodMan — Idea
## One-line value prop
PodMan is a real-time AI team coordination agent that watches consented work signals, maintains live project memory, and proactively notifies collaborators when dependencies, blockers, or handoffs emerge — before anyone has to ask.
---
## Problem
Teams working on the same project lose time because progress is fragmented across people, editors, terminals, and half-finished messages. Coordination gaps — a completed endpoint, a resolved blocker, two engineers duplicating work — are discovered too late, causing idle time, broken handoffs, and missed dependencies.
Slack doesn't help. Stand-ups are too slow. GitHub only knows pushed state.
---
## Solution
PodMan is an ambient AI agent that:
1. Watches each engineer's screen via periodic snapshots (consented, browser-native)
2. Extracts structured context using Gemini Vision — current file, inferred task, terminal state
3. Maintains a shared live model of the team in MongoDB Atlas — who is doing what, who owns which files
4. Detects coordination events: dependency ready, blocker detected, duplicate work
5. Speaks proactively into the team's LiveKit room — engineers hear PodMan through their earbuds without leaving their editor
**The AI's job is not to chat. It is to notice what teammates miss and say so, exactly when it matters.**
---
## Target user
Small software teams: hackathon squads, startup engineering teams, student dev teams collaborating in real time on a shared codebase.
---
## Core AI job
- Maintain per-person live context (file, task, terminal)
- Infer shared project state (who owns what, what's blocked, what's ready)
- Detect 3 coordination event types:
- `DEPENDENCY_READY` — engineer A was waiting on work engineer B just completed
- `BLOCKER_DETECTED` — engineer appears stuck; another teammate can unblock
- `DUPLICATE_WORK` — 2+ engineers working on the same file simultaneously
- Generate a 12 sentence proactive voice nudge
- Deliver it into the LiveKit room via Gemini Live 2.5
---
## How it fits the Continual Learning track
PodMan builds an **ownership map** in MongoDB that persists across sessions:
- Session 1: PodMan needs 35 minutes of screen observations to infer who owns what
- Session 2+: PodMan already knows. First nudge fires in under 30 seconds.
The system gets demonstrably more useful the more it is used, with no user configuration required. That is the track definition met exactly.
---
## Architecture (one paragraph)
Each engineer opens a browser PWA on their laptop. The PWA captures a screen frame every 30 seconds via `getDisplayMedia` and POSTs it to Hermes, the server-side orchestrator running on DigitalOcean. Hermes calls Gemini Vision to extract structured context, writes it to MongoDB Atlas, updates the ownership map, and runs event detection across all active engineers. When a coordination event fires, Hermes generates a short spoken message and publishes it as audio into the team's LiveKit room via Gemini Live 2.5. Engineers hear PodMan through their earbuds. No Slack. No tab switching. No interruption to the editor flow.
---
## Demo wow moment
> Alice is building the auth endpoint. Carol is visibly blocked — her terminal shows `connection refused`. PodMan detects the blocker and says aloud: "Carol, looks like you're waiting on auth. Alice is actively building it — hang tight."
>
> Two minutes later, Alice's server starts. PodMan says: "Carol, Bob — Alice just got the auth endpoint running. You're clear to integrate."
>
> Nobody asked. Nobody pinged anyone on Slack. PodMan just knew.
---
## What PodMan is NOT
- Not a chat interface
- Not a dashboard product
- Not raw surveillance — engineers consent by joining the room and sharing their screen
- Not a task manager
- Not a GitHub integration (v1)
---
## Prize alignment
| Prize | How PodMan earns it |
| --------------------- | ----------------------------------------------------------------------------------------------------- |
| Best Gemini 3.5 / 2.5 | Gemini Vision for screen understanding + Gemini Live 2.5 for voice output |
| Best LiveKit | LiveKit is the real-time backbone for room presence and voice delivery — load-bearing, not decorative |
| Best DigitalOcean | Hermes deployed on DigitalOcean App Platform; MongoDB Atlas on DO-adjacent infrastructure |
+57 -57
View File
@@ -1,15 +1,24 @@
# LiveKit Integration Spec
LiveKit is the real-time backbone for PodMan. It handles room presence and voice delivery. It is load-bearing — not decorative.
Status: active / matches code.
LiveKit is the real-time backbone for PodMan. It carries the **screen-share
perception input**, the **intervention data channel**, and **all room audio**
(Gemini TTS escalations, the Lyria score, and the live conversation agent). It is
load-bearing, not decorative.
---
## Room structure
- One LiveKit room per project pod: `room = podId`
- Engineers join as named participants (e.g. `alice`, `bob`)
- Hermes joins as `podman-hermes`
- All participants stay connected for the duration of the session
- One LiveKit room per pod: `room = podId`.
- Engineers join as named participants (e.g. `alice`, `bob`).
- PodMan runs **multiple agent identities** in/around a room:
- `podman-hermes` — the main vision + intervention agent (`@livekit/rtc-node`).
- `podman-live-conversation` — the Gemini Live voice agent (Python).
- short-lived `podman-hermes-job-*` publishers for async job events.
- A fixed identity matters: a second `podman-hermes` evicts the first and they
flap, dropping interventions. systemd keeps exactly one alive in production.
---
@@ -17,88 +26,79 @@ LiveKit is the real-time backbone for PodMan. It handles room presence and voice
**Joining:**
1. PWA calls `POST /pods/:podId/token` → receives `{ token, url }`
2. LiveKit client connects to the room with the token
3. PWA publishes screen track via `getDisplayMedia`
4. PWA sets mic enabled for ambient presence
1. PWA calls `POST /api/token` with `{ podId, identity }` `{ token, url }`.
2. LiveKit client connects with the token.
3. PWA publishes the screen track via `getDisplayMedia`.
4. PWA enables mic for ambient presence (used by the conversation agent).
**Receiving:**
- LiveKit client automatically receives Hermes audio track
- No special subscription needed — LiveKit delivers audio to all participants
- PWA also listens for data channel messages from Hermes for UI card updates
- Subscribes to remote agent audio tracks (TTS, Lyria, conversation) and attaches
them to a hidden audio sink.
- Browser autoplay restrictions apply: the PWA calls `room.startAudio()` from a
user gesture (`Enable audio`, `Test PodMan voice`, `Share screen`, first room
click).
- Listens on the data channel for cards and `VOICE_CUE` fallback text.
**Data channel listener (PWA):**
```ts
room.on(RoomEvent.DataReceived, (payload, participant) => {
if (participant?.identity !== 'podman-hermes') return;
const nudge = JSON.parse(new TextDecoder().decode(payload));
// nudge: { type, message, involvedEngineers, file, sentAt }
appendNudgeToFeed(nudge);
if (!participant?.identity.startsWith('podman-')) return;
const msg = JSON.parse(new TextDecoder().decode(payload));
// msg.type: COLLISION | ACK | GIT_REPORT | VOICE_CUE | HERMES_JOB_EVENT
appendInterventionToFeed(msg);
});
```
All data messages share the `podman.intervention` topic (`DATA_TOPIC`).
---
## Hermes side (LiveKit Agent)
## Agent side (`podman-hermes`)
**Framework:** LiveKit Agents (Node.js)
**Framework:** `@livekit/rtc-node`. **Code:** `backend/src/agent/podman.ts`,
`backend/src/action/hermes.ts`, `backend/src/voice/live.ts`.
**Startup:**
1. Subscribes to engineers' screen-share tracks and samples frames for Gemini
Vision.
2. Detects collisions, gates them through the learning policy, then publishes a
card/message on the data channel.
3. For critical collisions, generates Gemini TTS audio and publishes it as a
microphone-source audio track, held for the audio duration plus a tail/hold
window so subscribers finish playout. Voice publishing logs frame count,
estimated duration, queued playout, and hold time for diagnostics.
1. Hermes mints its own token via the same `createPodToken` function with `identity: 'podman-hermes'`
2. Connects to the configured room as `podman-hermes`
3. Registers as a LiveKit Agent with Gemini Live 2.5 as voice provider
---
**Voice delivery:**
## Live conversation agent (`podman-live-conversation`)
1. Nudge message text is ready (from Gemini text generation)
2. Hermes passes text to Gemini Live 2.5 via LiveKit Agents voice pipeline
3. Audio streams into the room in real-time
4. All participants hear it
**Framework:** LiveKit Agents for Python (`AgentSession`, `function_tool`,
`google.realtime.RealtimeModel`). **Code:**
`agents/podman-live-conversation/agent.py`.
**Data channel message (sent alongside audio):**
```ts
const nudge = {
type: 'DEPENDENCY_READY' | 'BLOCKER_DETECTED' | 'DUPLICATE_WORK',
message: string, // the spoken text
involvedEngineers: string[],
file: string | null,
sentAt: string, // ISO timestamp
};
room.localParticipant.publishData(
new TextEncoder().encode(JSON.stringify(nudge)),
{ reliable: true }
);
```
Joins the pod room on demand (`POST /api/pods/:id/live-conversation/start`),
streams speech-to-speech with Gemini Live, and answers using repo/git/memory
function tools. It can delegate long tasks to the async Hermes job runner and
narrate progress. See `docs/hermes.md`.
---
## Token endpoint
Already implemented at `POST /api/token`.
Hermes uses the same endpoint. Grants:
`POST /api/token` mints room tokens for engineers and agents alike. Grants:
- `roomJoin: true`
- `canPublish: true` (for audio track)
- `canPublishData: true` (for data channel)
- `canPublish: true` (audio + screen)
- `canPublishData: true` (data channel)
- `canSubscribe: true`
---
## Gemini voice model
- Model ID: `gemini-3.1-flash-tts-preview`
- Hermes generates Gemini TTS audio and publishes it as a LiveKit audio track.
- The backend keeps a Gemini Live path for future model availability, but the verified deployment path uses TTS.
Short-lived job publishers use `canSubscribe: false`.
---
## What LiveKit does NOT do in PodMan
- No video tracks from Hermes
- No mic transcription (not needed for v1)
- No SFU mixing — standard room behavior is sufficient
- No video tracks published by agents.
- No mic transcription outside the live conversation agent.
- No custom SFU mixing — standard room behavior is sufficient.
+187 -114
View File
@@ -1,143 +1,216 @@
# MongoDB Atlas Integration Spec
MongoDB Atlas is PodMan's shared memory. It stores live engineer state, the ownership map that enables continual learning, coordination events, and nudge history.
Status: demo-backed / active
MongoDB Atlas is PodMan's shared memory. It stores live work observations,
collision predictions (with vector embeddings for recall), interventions,
outcomes, latest engineer state, the materialized Team memory graph, and async
Hermes job runs.
See also:
- [`docs/cont_learning.md`](cont_learning.md) for outcome-backed team memory,
graph materialization, and `$graphLookup` traversal.
---
## Collections
## Current Collections
### `engineer_states`
Latest context per engineer. Two writers, one collection — vision pipeline upserts vision fields, git watcher script upserts git fields independently. Hermes reads the merged document for event detection.
Latest context per engineer. The local git watcher writes git fields; the vision
pipeline may write screen-derived fields. Each writer updates only its own
fields so MongoDB upserts merge cleanly.
```ts
{
_id: string, // engineerId (stable across sessions)
podId: string,
name: string, // display name
Key fields:
// --- Vision fields (written by Hermes via POST /ingest) ---
currentFile: string | null, // active file inferred from screen
inferredTask: string | null, // what engineer appears to be doing
terminalVisible: boolean,
recentTerminalOutput: string | null,
confidence: number, // Gemini Vision confidence (01)
visionUpdatedAt: Date,
- `podId`
- `name`
- `currentFile`
- `inferredTask`
- `confidence`
- `changedFiles`
- `diffStat`
- `recentCommit`
- `branch`
- `visionUpdatedAt`
- `gitUpdatedAt`
- `updatedAt`
// --- Git fields (written directly by scripts/podman-agent.mjs) ---
changedFiles: string[], // files with uncommitted changes (git status)
diffStat: string | null, // e.g. "auth/middleware.ts | 24 +++++"
recentCommit: string | null, // most recent commit message
branch: string | null, // current branch name
gitUpdatedAt: Date,
Primary use: deterministic dirty/unpushed truth for collision detection and
graph discovery.
// --- Shared ---
updatedAt: Date // most recent write from either source
}
```
### `observations`
**Index:** `{ podId: 1, updatedAt: -1 }`
Structured perception events from consented screen context and agent inference.
**Two writers, no conflict:** vision upsert uses `$set` on vision fields only; git upsert uses `$set` on git fields only. MongoDB upsert semantics merge them cleanly.
Key fields:
**Usage:** Hermes reads all documents for a given `podId` after each update to run event detection. Both vision and git context are available in the same document — `changedFiles` provides ground truth, `currentFile` provides screen context.
- `podId`
- `engineerId`
- `currentFile`
- `symbol`
- `activity`
- `confidence`
- `observedAt`
Primary use: observe/store proof and active editing edges in the Team memory
graph.
### `collisions`
Predicted coordination risks, with memory enrichment for recall.
Key fields:
- `id`
- `podId`
- `file`
- `symbol`
- `engineers`
- `severity`
- `overlapKind` — optional; `file`/undefined for same-file collisions,
`research` for code-edit ↔ research overlaps.
- `researchTopic`, `researchSource`, `researcher`, `editor` — optional fields
present only for research overlaps.
- `memorySignature`
- `githubState`
- `detectedAt`
- `memoryText` — short text embedded for recall
- `embedding` — vector (Voyage `voyage-4-lite` or Gemini `gemini-embedding-001`)
- `embeddingProvider``voyage` | `gemini`
Vector index `collision_embedding` (Atlas Vector Search) powers `$vectorSearch`
recall in `backend/src/memory/vectors.ts`. When Atlas vector search is
unavailable, recall falls back to app-side cosine, then exact signature/file
matching.
Primary use: collision cards, vector + signature recall, and graph risk paths.
### `interventions`
Actions PodMan sent or suggested.
Key fields:
- `id`
- `podId`
- `collisionId`
- `kind`
- `message`
- `suggestedAction`
- `status`
- `createdAt`
Primary use: closing the loop from prediction to a visible card, Hermes message,
or urgent voice cue.
### `outcomes`
Human or verifier supervision recorded through `POST /api/outcome`.
Key fields:
- `podId`
- `interventionId`
- `collisionId`
- `accepted`
- `wasRealCollision`
- `recordedAt`
Primary use: accepted and dismissed outcomes drive exact recall, suppression,
and learned graph paths.
### `suppressions`
Durable negative-feedback proof: one record per *suppressed repeat* — a
previously-dismissed collision signature recurred and PodMan stayed quiet.
Written at repeat time by `backend/src/agent/podman.ts` (once per recurrence,
re-armed on resolution), materialized as `suppressed` activity by
`backend/src/graph/live.ts`.
Key fields:
- `id`
- `podId`
- `collisionId`
- `file`
- `engineers`
- `priorInterventionId` — the dismissed intervention this repeat matched
- `priorDismissedAt`
- `suppressedAt` — the repeat time (drives recency in the activity stream)
Index `{ podId: 1, suppressedAt: -1 }`. Counted in `/api/memory/stats`.
**Preserve in any DB cleanup** — this is visible learning evidence, not noise.
### `team_model`
Durable per-pod summary memory.
Key fields:
- `podId`
- `ownership`
- `hotspots`
- `graph`
- `updatedAt`
Primary use: stable Team memory, including seeded `graph` snapshots used after
live materialization and before demo fallback.
### `graph_nodes` and `graph_edges`
Normalized mirror of the Team memory graph for MongoDB traversal.
Indexes:
- `graph_nodes`: `{ podId: 1, id: 1 }` unique
- `graph_edges`: `{ podId: 1, source: 1 }`
Primary use: `GET /api/pods/:podId/graph/reach/:id` with `$graphLookup`.
### `hermes_jobs` and `hermes_job_events`
Async Hermes task runs delegated from the live conversation agent (see
`docs/hermes.md`).
- `hermes_jobs` — one doc per job (`id` unique; `{ sessionId, status, updatedAt }`
index). Fields: `id`, `podId`, `sessionId`, `prompt`, `contextScope`,
`riskLevel`, `successCriteria`, `status`, `finalSummary`, timestamps.
- `hermes_job_events` — append-only step log (`{ jobId, createdAt }` index):
`accepted`, `heartbeat`, `step_started`, `step_output`, `needs_confirmation`,
`step_completed`, `completed`, `aborted`, `failed`. Output is redacted +
truncated before storage and mirrored to the room over LiveKit.
Primary use: durable, replayable record of what Hermes did, streamed live to the
conversation UI.
---
### `ownership_map`
## Graph Truth Order
Tracks who works on which files. Built up over the session. **Persists across sessions** — this is the continual learning artifact.
`GET /api/pods/:podId/graph` follows this order:
```ts
{
_id: string, // `${podId}:${file}`
podId: string,
file: string,
primaryOwner: string, // engineerId with most recent activity on this file
contributors: string[], // all engineerIds observed on this file
observationCount: number, // total frames where this file was seen
lastSeenAt: Date
}
```
1. Live graph from real collections.
2. Seeded graph from `team_model.graph` and mirrored graph records.
3. Demo fallback graph for stage safety.
**Index:** `{ podId: 1, file: 1 }` (unique)
**Upsert logic:**
- On each context update where `currentFile` is non-null:
- Increment `observationCount`
- Update `primaryOwner` to the engineer with the most recent `lastSeenAt` on this file
- Add engineerId to `contributors` if not present
- Update `lastSeenAt`
**Continual learning:** Hermes loads this collection on startup for the pod. If history exists, it pre-populates the in-memory ownership cache before the first frame arrives.
Seeded and fallback graphs are acceptable for demos only when labeled honestly.
---
### `events`
## Demo Proof Path
Every coordination event detected by Hermes.
```ts
{
_id: ObjectId,
podId: string,
type: 'DEPENDENCY_READY' | 'BLOCKER_DETECTED' | 'DUPLICATE_WORK',
involvedEngineers: string[],
file: string | null,
reason: string, // 1-sentence explanation from Gemini
nudgeSent: boolean, // false if suppressed by cooldown
detectedAt: Date
}
```
**Index:** `{ podId: 1, detectedAt: -1 }`
Observe screen/git state -> detect collision -> send intervention -> accept or
dismiss outcome -> recall similar event -> show changed graph or changed
behavior.
---
### `nudges`
## What MongoDB Does Not Store
Every voice nudge sent to the room.
```ts
{
_id: ObjectId,
podId: string,
eventId: ObjectId, // ref to events collection
targetEngineers: string[],
message: string, // the spoken text
sentAt: Date
}
```
**Index:** `{ podId: 1, sentAt: -1 }`
**Cooldown check:** before sending a nudge, Hermes queries this collection for any nudge in the last 3 minutes for the same `podId`. If found, suppresses the new nudge and marks the event as `nudgeSent: false`.
---
## Hermes startup sequence
```
1. Connect to Atlas using MONGODB_URI
2. Load ownership_map for this podId
3. Build in-memory cache: Map<file, { primaryOwner, contributors }>
4. Begin accepting /ingest requests
```
---
## Atlas configuration
- **Cluster tier:** M0 (free) is sufficient for hackathon scale
- **Region:** same as DigitalOcean deployment (e.g. NYC1)
- **Auth:** connection string in `MONGODB_URI` env var
- **Collections created automatically** on first write (no schema migration needed)
---
## What MongoDB does NOT store
- Raw screenshot frames (too large — frames are processed in-memory by Hermes and discarded)
- Full Gemini response objects (only extracted fields are stored)
- Session recordings
- Raw screenshot frames.
- Screen recordings.
- Secrets or credentials.
- Full terminal logs.
- Full Gemini response objects beyond extracted fields needed for memory.
+166
View File
@@ -0,0 +1,166 @@
# Build spec: cross-channel overlap (code-edit ↔ research nudge)
> Self-contained spec for an implementing agent. Everything needed to build is
> here: current-state anchors, exact edits, code sketches, verification. Read
> the referenced files before editing.
## Context / why
Today PodMan only fires when **two engineers touch the same file** (basename
match) with unpushed changes — `backend/src/collision/detector.ts:36`. That makes
the system feel like a one-file trick. The wow we want: detect a **code↔research
overlap** — one teammate is *editing* `livekit.py` while another is *researching*
the same topic in a browser (LiveKit docs/SDK). PodMan nudges:
> 🤝 bob is deep in LiveKit docs while you edit livekit.py — sync up before duplicating effort.
This reframes PodMan from merge-conflict detector to a **team-coordination agent
that catches duplicated effort / knowledge overlap** — stronger continual-learning
story, distinct demo beat.
### Locked decisions
- **Signal capture:** Gemini vision on the existing LiveKit screenshare. When a
teammate shares a browser/docs window, vision classifies it `research` and
extracts `{researchTopic, researchSource}`. No browser extension, no new client
surface. Reuses `backend/src/agent.ts` + `backend/src/vision/gemini.ts`.
- **Matching:** semantic embeddings (reuse `embed()` + `cosine()` in
`backend/src/memory/vectors.ts`). Deterministic stem/keyword fallback fires when
an embed call returns `null`, so the demo path never depends on a live vector call.
- **Framing:** collaboration nudge (`ping_teammate`), spoken once for the beat.
## Current-state anchors (read these first)
- `backend/src/collision/detector.ts:14-20``fileKey()` stem/basename logic to mirror.
- `backend/src/collision/detector.ts:36-89``detectCollisions()` (same-file path; leave unchanged).
- `backend/src/agent/podman.ts:85-116``onScreenFrame()` orchestration; `:124-126` `conflictKey()`; `:128-183` `handle()`.
- `backend/src/agent/podman.ts:38-46``engineersOverlapOnFile()` (git ground-truth; research must skip it).
- `backend/src/vision/gemini.ts:7-70``SCHEMA`, prompt, and `EngineerContext` mapping.
- `backend/src/memory/vectors.ts:52-66` `cosine()`, `:68-70` private `embed()`.
- `backend/src/action/hermes.ts:46-62``publishHermesIntervention()` (speaks only if `voiceLine` passed).
- `frontend/src/livekit/useInterventions.ts:36-48` — data-channel handler (renders `intervention.message` verbatim).
- `shared/src/engineer.ts:6-23` `EngineerContext`; `shared/src/collision.ts:7-28` `Collision`.
**Off-spec gate:** nothing in `docs/` covers cross-channel overlap. Update specs
in step 1 **before** code (repo documentation-first rule).
## Concurrency / safety (several people build at once)
- Shared-contract edits are **additive optional fields only** — no signature
changes. Safe.
- `backend/src/agent/podman.ts` is hot: two in-place edits (`conflictKey`,
`handle` branch). `git pull --rebase` before pushing.
- **No required frontend change** — nudge text rides in `intervention.message`,
already rendered by `frontend/src/components/PodView.tsx`.
## Implementation steps
### 1. Specs first
- `docs/gemini.md` — vision classifies `editing` vs `research`; extracts `researchTopic`/`researchSource`.
- `docs/hermes.md` — new intervention type **research overlap** (collaboration nudge, `ping_teammate`, spoken once); explicitly NOT a merge conflict.
- `docs/mongodb.md` — new optional `Collision` fields.
- `docs/demo.md` — insert ~30s beat after the same-file collision.
### 2. Shared contract (additive, optional)
`shared/src/engineer.ts` — add to `EngineerContext`:
```ts
mode?: 'editing' | 'research';
researchTopic?: string;
researchSource?: string; // domain, e.g. "docs.livekit.io"
```
`shared/src/collision.ts` — add to `Collision`:
```ts
overlapKind?: 'file' | 'research'; // undefined = file (preserves current behavior)
researchTopic?: string;
researchSource?: string;
researcher?: string; // engineer doing research
editor?: string; // engineer editing the file
```
### 3. Vision: classify editing vs research
`backend/src/vision/gemini.ts` — extend `SCHEMA` with `mode`, `researchTopic`,
`researchSource` (+ propertyOrdering). Update the prompt so a browser/docs/SDK
frame returns `mode:'research'` + topic + source domain, else `mode:'editing'`
with existing IDE fields. Map new fields into the returned `EngineerContext`.
Keep `thinkingConfig.thinkingBudget:0` and `MEDIA_RESOLUTION_LOW`.
### 4. Semantic matcher helper (reuse embed/cosine)
`backend/src/memory/vectors.ts` — add and export:
```ts
export async function semanticSimilarity(a: string, b: string): Promise<number | null> {
const [va, vb] = await Promise.all([embed(a, 'query'), embed(b, 'document')]);
if (!va || !vb) return null;
return cosine(va, vb);
}
```
### 5. New detector (additive file)
`backend/src/env.ts` — add `RESEARCH_OVERLAP_THRESHOLD` (default `0.6`).
`backend/src/collision/research.ts` (new):
```ts
export interface ResearchOpts { similarity?: (a: string, b: string) => Promise<number | null>; threshold?: number; }
export async function detectResearchOverlaps(
contexts: EngineerContext[],
gitStates: Map<string, GitState> | undefined,
opts: ResearchOpts = {},
): Promise<Collision[]>;
```
Logic:
- **researchers** = contexts with `mode==='research'` && `researchTopic`.
- **editor files** = ground truth from `gitStates[*].changedFiles` (reliable) plus
any `mode==='editing'` `currentFile`; each tagged with its engineer.
- For each distinct (researcher, editor) pair on a file, score
`similarity("${topic} ${source}", "<file stem words> <symbol> <activity>")`
(default `similarity = semanticSimilarity`). Fire when `score >= threshold`
(default `RESEARCH_OVERLAP_THRESHOLD`).
- **Fallback:** if `score === null`, deterministic stem/token overlap (mirror
`fileKey()` stemming) — guarantees `livekit``livekit.py` fires offline.
- Dedupe to best-scoring file per researcher; require distinct engineers.
- Emit `Collision`: `id: col_research_<stem>_<Date.now()>`, `engineers:[editor, researcher]`,
`file:<editor file>`, `severity:'warn'`, `overlapKind:'research'`, plus
`researcher`, `editor`, `researchTopic`, `researchSource`.
### 6. Wire into orchestrator — `backend/src/agent/podman.ts`
- In `onScreenFrame`, after `detectCollisions(...)` (line ~100):
```ts
const research = await detectResearchOverlaps([...this.contexts.values()], gitStates);
const collisions = [...fileCollisions, ...research];
```
(concat **before** the re-arm + handle loops at `:104` / `:111` / `:115`).
- `conflictKey()` (`:124`) — namespace by overlap kind so research and file
overlaps on the same file don't share an edge-trigger key:
```ts
return `${collision.overlapKind ?? 'file'}:${comparableBasename(collision.file)}`;
```
- `gitOverlap` loop (`:111-113`) — guard: only call `engineersOverlapOnFile` when
`collision.overlapKind !== 'research'` (researcher won't have the file dirty;
leave `gitOverlap` undefined for research).
- `handle()` (`:128`) — branch on `overlapKind === 'research'`:
- `message`: `` `🤝 ${researcher} is researching ${researchTopic}` + (researchSource ? ` (${researchSource})` : '') + ` while ${editor} edits ${shortFile} — sync up before duplicating effort.` ``
- `voiceLine`: `` `${researcher} is researching ${researchTopic} while ${editor} works on ${shortFile}. Worth a quick sync.` ``
- `suggestedAction.kind = 'ping_teammate'`.
- Pass `voiceLine` to `publishHermesIntervention` **regardless of severity** so
the beat is spoken once. (For file collisions keep existing
`severity === 'critical' ? voiceLine : undefined`.)
- Existing `recallSimilar` / `shouldIntervene` gate stays unchanged.
### 7. Frontend (nice-to-have, cut if behind)
`frontend/src/livekit/useInterventions.ts` — capture `msg.collision.overlapKind`;
show a 🤝 badge on the card in `PodView.tsx`. Core path needs nothing.
## Verification
1. **Build:** `pnpm -r build` (shared builds first — order matters).
2. **Detector (offline, deterministic):** harness calling `detectResearchOverlaps`
with stubbed `opts.similarity`:
- researcher `{mode:'research', researchTopic:'LiveKit agent init', researchSource:'docs.livekit.io'}`
+ gitState with `livekit.py` dirty for another engineer → exactly one
`overlapKind:'research'` collision naming both.
- single engineer both sides → no overlap.
- `similarity` returns `null` → keyword fallback still fires on `livekit`.
3. **Live:** run the agent (systemd on the box, or local); one participant shares a
browser tab on LiveKit docs, another keeps `livekit.py` dirty (git sidecar
running) → 🤝 nudge card + one-time spoken cue. Fires once (edge-triggered),
re-arms after the browser closes.
4. **Regression:** same-file collision still fires, unaffected (separate
`conflictKey` namespace).
+299
View File
@@ -0,0 +1,299 @@
# Build spec: Work History "Coordination ROI" band
> Self-contained spec for an implementing agent. Everything needed to build is
> here: current-state anchors, exact edits, code sketches, verification. Read
> the referenced files before editing. Changes are **additive** — one optional
> field on a shared type, two new backend queries, one new presentational
> component. No route signature changes, no DB writes, no new deps.
## Context / why
The Work History dialog (opens when you click a teammate's "History" button)
today shows three raw counters (files / screen logs / git changes), a Recent
files bar list, and a Timeline. It is pure **activity volume** — it summarizes
nothing and shows none of PodMan's actual value.
We add a **Coordination ROI band** at the top of the dialog that answers "what
did PodMan save this person?": estimated rework-hours saved by clashes Hermes
caught early, plus hard defensible counts. This is a *summary* layer — it must
NOT re-stream the pod activity feed ("Team Memory"). Existing Recent files +
Timeline sections stay exactly as they are, below the new band.
### Locked decisions
- **Primary story = saved time / ROI.** One number leads: `~Xh Ym rework saved`.
- **Transparent heuristic.** The hours are a labeled estimate (`~`, `(est.)`)
with an info tooltip showing the per-clash breakdown. Technical judges can see
the model; we never claim precision.
- **Backend extension required** (collisions/interventions are not in the
current `MemberWorkHistory` payload). Field is optional → old frontends and
pods with no collisions degrade gracefully (band hidden).
- **Cut:** editing/research focus donut. ROI is the single story; a donut
dilutes it.
- **Credit split** across involved engineers so a 2-person clash does not
double-count the pod total.
## Heuristic (this is the contract — implement exactly)
For each `collisions` doc in the window where the member is involved AND an
`interventions` doc exists for it (Hermes actually surfaced it):
1. **Eligibility (must be "real"):** count the collision only if
`gitOverlap === true` OR `severity === 'critical'`. Otherwise skip.
2. **Weight by kind/severity:**
| condition | minutesEach |
|---|---|
| `overlapKind === 'research'` | 10 |
| `severity === 'critical'` (same-file) | 45 |
| `severity === 'warn'` (same-file) | 20 |
| `severity === 'info'` (same-file) | 10 |
(Evaluate research first; a research overlap is always 10 regardless of severity.)
3. **Member share:** `minutesEach / max(1, engineers.length)`. Sum all shares →
`savedMinutes` (round to nearest integer).
4. **breakdown[]:** group eligible collisions by label
(`"critical same-file"`, `"warn same-file"`, `"research overlap"`, etc.),
each entry `{ label, count, minutesEach }` — for the tooltip.
Member "involved" = `engineers` array contains the member (case-insensitive),
OR `researcher`/`editor` equals the member (case-insensitive). Reuse the
existing `sameMember()` helper logic.
Hard counts (no estimation):
- `clashesCaught` = number of eligible collisions (the integer behind the hours).
- `filesDeconflicted` = distinct `collision.file` values across eligible collisions.
- `conflictFreeCommits` / `totalCommits`: see "commits" note below.
### Commits note (keep simple, defensible)
There is no per-commit log in the window. Use the git ground-truth we have:
`totalCommits` = count of distinct `changedFiles` for the member from
`engineer_states` (already loaded as `gitState.changedFiles.length`).
`conflictFreeCommits` = `totalCommits - filesDeconflicted` (clamped ≥ 0). This
reads as "files in flight that never hit a clash". Label it in the UI as
**"conflict-free files"**, not commits, so the wording matches the data. (The
band copy below already says "files".)
## Current-state anchors (read these first)
- `shared/src/member-history.ts:24-36``MemberWorkHistory` interface (add `roi?`).
- `backend/src/activity/member-history.ts:73-174``getMemberWorkHistory()`.
- `:83-94` — the `Promise.all` that loads `observations` + `engineer_states`. Add collisions/interventions here.
- `:45-47``sameMember()` helper to reuse for involvement check.
- `:124``gitState` already in scope; `:169` uses `gitState.changedFiles.length`.
- `:161-173` — the returned object (add `roi`).
- `backend/src/memory/db.ts:47-48` — canonical collection names `collisions`, `interventions` (typed `Collision`, `Intervention`).
- `backend/src/activity/store.ts:167-178` — reference query shape for both collections.
- `shared/src/collision.ts:7-36``Collision` (`engineers`, `severity`, `overlapKind`, `gitOverlap`, `researcher`, `editor`, `file`, `detectedAt`).
- `shared/src/intervention.ts:8-18``Intervention` (`collisionId`, `podId`).
- `frontend/src/components/PodView.tsx:33-35` — type imports from `@podman/shared`.
- `frontend/src/components/PodView.tsx:1035-1041` — the `history && !loading && !error` block; the 3-stat grid is the insert point (band goes ABOVE it).
- `frontend/src/components/PodView.tsx:1093-1100``HistoryStat` (style reference for new sub-stats).
- `frontend/src/components/PodView.tsx``timeLabel()` exists for relative time; reuse if needed.
- Route is unchanged: `backend/src/server.ts:502-506` already returns whatever `getMemberWorkHistory` produces.
## Edit 1 — shared type (`shared/src/member-history.ts`)
Add to the `MemberWorkHistory` interface (after `timeline`):
```ts
/** Coordination ROI summary — clashes Hermes caught for this member. Optional
* so pods with no collisions / older payloads render without the band. */
roi?: MemberWorkHistoryRoi;
}
export interface MemberWorkHistoryRoi {
/** Estimated rework minutes saved (heuristic, labeled "~/est." in UI). */
savedMinutes: number;
/** Eligible collisions caught early (hard count). */
clashesCaught: number;
/** Distinct files that hit an eligible clash. */
filesDeconflicted: number;
/** Member files in flight that never hit a clash. */
conflictFreeFiles: number;
/** Total member files in flight (git changedFiles). */
totalFiles: number;
/** Per-kind breakdown for the tooltip. */
breakdown: { label: string; count: number; minutesEach: number }[];
}
```
`shared/src/index.ts:56-59` already re-exports the member-history types via a
`export type { ... }` block — add `MemberWorkHistoryRoi` to that list.
## Edit 2 — backend (`backend/src/activity/member-history.ts`)
1. Import the types: `import type { Collision } from '@podman/shared';` (Intervention only needs `collisionId`, can stay untyped or import too).
2. In the `Promise.all` at `:83`, add two queries (windowed, podId-scoped):
```ts
db.collection<Collision>('collisions')
.find({ podId, detectedAt: { $gte: since } }, { projection: { _id: 0 } })
.sort({ detectedAt: -1 })
.limit(200)
.toArray(),
db.collection<{ collisionId: string }>('interventions')
.find({ podId }, { projection: { collisionId: 1, _id: 0 } })
.toArray(),
```
3. After building `fileRows`, compute `roi` with a new local helper
`computeRoi(member, collisions, interventions, gitState)`:
```ts
function computeRoi(
member: string,
collisions: Collision[],
interventionCollisionIds: Set<string>,
changedFileCount: number,
): MemberWorkHistoryRoi {
const involved = (c: Collision) =>
c.engineers?.some((e) => sameMember(e, member)) ||
sameMember(c.researcher, member) ||
sameMember(c.editor, member);
const eligible = collisions.filter(
(c) =>
involved(c) &&
interventionCollisionIds.has(c.id) &&
(c.gitOverlap === true || c.severity === 'critical'),
);
const weightOf = (c: Collision): { label: string; minutes: number } => {
if (c.overlapKind === 'research') return { label: 'research overlap', minutes: 10 };
if (c.severity === 'critical') return { label: 'critical same-file', minutes: 45 };
if (c.severity === 'warn') return { label: 'warn same-file', minutes: 20 };
return { label: 'info same-file', minutes: 10 };
};
let savedMinutes = 0;
const groups = new Map<string, { count: number; minutesEach: number }>();
for (const c of eligible) {
const { label, minutes } = weightOf(c);
savedMinutes += minutes / Math.max(1, c.engineers?.length ?? 1);
const g = groups.get(label) ?? { count: 0, minutesEach: minutes };
g.count += 1;
groups.set(label, g);
}
const filesDeconflicted = new Set(eligible.map((c) => c.file)).size;
const totalFiles = changedFileCount;
return {
savedMinutes: Math.round(savedMinutes),
clashesCaught: eligible.length,
filesDeconflicted,
conflictFreeFiles: Math.max(0, totalFiles - filesDeconflicted),
totalFiles,
breakdown: [...groups.entries()].map(([label, g]) => ({
label,
count: g.count,
minutesEach: g.minutesEach,
})),
};
}
```
4. Build the intervention id set and attach to the return:
```ts
const interventionIds = new Set(interventions.map((i) => i.collisionId));
const roi = computeRoi(member, collisions, interventionIds, gitState?.changedFiles?.length ?? 0);
```
Add `roi` to the returned object at `:161-173`. Always return it (it self-zeroes
when there are no clashes); the frontend decides whether to show the band based
on `clashesCaught`/`savedMinutes`.
> Note `Collision.id` survives the `{ _id: 0 }` projection — `id` is a real field
> (`shared/src/collision.ts:8`), distinct from Mongo `_id`. Do not project it out.
## Edit 3 — frontend (`frontend/src/components/PodView.tsx`)
1. Add `MemberWorkHistoryRoi` to the `@podman/shared` type import (`:33-35`).
2. Insert `<RoiBand roi={history.roi} />` immediately inside the
`history && !loading && !error` block, BEFORE the `grid ... sm:grid-cols-3`
stat grid (`:1037`).
3. New presentational component (place near `HistoryStat`, `:1093`):
```tsx
function formatSaved(minutes: number): string {
if (minutes < 60) return `~${minutes}m`;
const h = Math.floor(minutes / 60);
const m = minutes % 60;
return m ? `~${h}h ${m}m` : `~${h}h`;
}
function RoiBand({ roi }: { roi?: MemberWorkHistoryRoi }) {
if (!roi || roi.clashesCaught === 0) return null; // no clashes → no band
const conflictFree = roi.totalFiles
? Math.round((roi.conflictFreeFiles / roi.totalFiles) * 100)
: 100;
return (
<section className="rounded-lg border bg-primary/5 p-4">
<div className="flex flex-wrap items-end justify-between gap-3">
<div>
<div className="flex items-center gap-1.5">
<p className="font-mono text-2xl font-semibold">{formatSaved(roi.savedMinutes)}</p>
<span className="text-sm text-muted-foreground">rework saved</span>
<RoiTooltip roi={roi} />
</div>
<p className="mt-0.5 text-xs text-muted-foreground">
estimated · clashes caught pre-commit
</p>
</div>
<div className="text-right">
<p className="font-mono text-lg font-medium">{roi.clashesCaught}</p>
<p className="text-xs text-muted-foreground">clashes caught early</p>
</div>
</div>
<div className="mt-3 h-2 overflow-hidden rounded-full bg-muted">
<div className="h-full rounded-full bg-primary" style={{ width: `${conflictFree}%` }} />
</div>
<p className="mt-1.5 text-[0.68rem] text-muted-foreground">
conflict-free: {roi.conflictFreeFiles} of {roi.totalFiles} files ·{' '}
{roi.filesDeconflicted} auto-deconflicted
</p>
</section>
);
}
```
4. Tooltip — reuse the existing Tooltip primitives already imported in this file
(`Tooltip`, `TooltipTrigger`, `TooltipContent` — see the History button at
`:982-988`). `RoiTooltip` renders an `ⓘ`/`InfoIcon` trigger; content lists
`breakdown` rows as `{count} × {minutesEach}m {label}` plus a footer line
`"split across engineers · est. only"`. Keep it a few lines.
## Verification
```bash
pnpm -r build # @podman/shared first, then backend + frontend typecheck
```
Manual / demo check (against a pod that has collisions, e.g. demo-pod):
1. Open the app, click a teammate who was in a clash → History.
2. Band shows at top: `~Xh Ym rework saved`, clashes-caught integer, conflict-free bar.
3. Hover `ⓘ` → breakdown rows match the heuristic table.
4. Click a teammate with **no** clashes → band absent, rest of dialog unchanged.
5. Existing Recent files + Timeline sections render unchanged below the band.
Quick data sanity (optional, on the box or via mongosh):
```js
db.collisions.find({ podId: "demo-pod" }).count() // > 0 for a real band
db.interventions.find({ podId: "demo-pod" }).count() // surfaced clashes
```
## Coordination / merge risk (team is concurrent)
- Touches two shared hot files: `shared/src/member-history.ts` and
`frontend/src/components/PodView.tsx`. Both edits are **additive** (one
optional field + one new component + one insert line). Low conflict risk, but
announce before pushing.
- `git pull --rebase origin main` immediately before push. Never force-push main.
- No changes to `/api/...` route signatures, no new deps, no Mongo writes.
## Deploy (after merge, on the box — see CLAUDE.md ops)
```bash
cd /root/podman && git pull && pnpm -r build
rm -rf /var/www/podman/* && cp -r frontend/dist/* /var/www/podman/
systemctl restart podman-platform-api podman-platform-agent
```
(Frontend-visible change + backend payload change → both the static build and
the API service must be refreshed.)
@@ -1,135 +0,0 @@
---
name: podman-design
description: Full system design for PodMan — real-time AI team coordination agent using Gemini Vision, Gemini Live 2.5, LiveKit, and MongoDB Atlas
metadata:
type: project
---
# PodMan — System Design
## Concept
PodMan is a real-time AI team coordination agent for software teams. Engineers join a LiveKit room with earbuds. Each engineer's browser PWA captures their screen every 30s and sends it to Hermes (server-side orchestrator on DigitalOcean). Hermes uses Gemini Vision to extract structured context per engineer, detects coordination events, and speaks proactive nudges into the room via Gemini Live 2.5 through LiveKit. MongoDB Atlas stores team state and an ownership map that persists across sessions.
**Track:** Continual Learning — the ownership map makes PodMan faster and smarter each session with no user configuration.
---
## Architecture
```
┌──────────────── Engineer laptop (Browser PWA) ──────────────────┐
│ getDisplayMedia → frame every 30s │
│ HTTP POST /ingest → { screenshot, engineerId, podId } │
│ LiveKit room joined → receives voice audio from Hermes │
│ Earbuds: hears PodMan proactive nudges │
└──────────────────────────────────────────────────────────────────┘
│ POST /ingest
┌────────────────── HERMES (DigitalOcean) ─────────────────────────┐
│ 1. Receive frame → Gemini Vision → EngineerContext │
│ 2. Write context to MongoDB (per-user state) │
│ 3. Update ownership map (file → engineer) │
│ 4. Run event detector over all active contexts │
│ 5. If event detected → Gemini generates voice message │
│ 6. Push audio into LiveKit room via Gemini Live 2.5 │
└──────────────────────────────────────────────────────────────────┘
│ read/write
MongoDB Atlas
(engineer_states, ownership_map,
events, nudges)
```
---
## Components
### PWA (local agent)
- Joins LiveKit room via existing `joinPod` flow
- Captures frame every 30s via `getDisplayMedia`, compresses to JPEG (1280×720, quality 0.7)
- POSTs `{ engineerId, podId, screenshotBase64, capturedAt }` to `POST /ingest`
- Receives Hermes audio track (automatic via LiveKit)
- Listens for data channel messages → renders nudge feed
- Two screens: join screen (built), active session screen (to build)
### Hermes (orchestrator)
- Express server + LiveKit Agent on DigitalOcean
- `POST /ingest`: receives frame, queues for vision
- Vision pipeline: Gemini 2.0 Flash → `EngineerContext`
- Confidence gate: discard frames with confidence < 0.6
- State writer: upsert `engineer_states` + `ownership_map` in MongoDB
- Event detector: Gemini text prompt over all active states
- Nudge generator: Gemini text → 12 sentence spoken message
- Voice publisher: Gemini Live 2.5 via LiveKit Agents → audio into room
- Data channel: sends structured nudge payload alongside audio
- Cooldown: 3 min between nudges per pod
### Gemini usage
- **Vision:** `gemini-2.0-flash` — screen → `{ currentFile, inferredTask, terminalVisible, recentTerminalOutput, confidence }`
- **Event detection:** `gemini-2.0-flash` — all engineer states → `{ event, involvedEngineers, file, reason }`
- **Nudge generation:** `gemini-2.0-flash` — event → spoken message text
- **Voice:** `gemini-3.1-flash-tts-preview` via LiveKit audio publication — text → audio
### MongoDB Atlas (4 collections)
- `engineer_states`: latest context per engineer, upserted each ingest
- `ownership_map`: file → primaryOwner + contributors, persists across sessions (continual learning)
- `events`: all detected coordination events
- `nudges`: all voice nudges sent + cooldown history
### LiveKit
- One room per pod
- Engineers publish screen track (used client-side for capture — Hermes does not subscribe)
- Hermes joins as `podman-hermes`, publishes audio + data channel messages
- Engineers receive audio automatically
---
## Event types
| Event | Trigger | Example nudge |
| ------------------ | -------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| `BLOCKER_DETECTED` | Engineer stuck (error in terminal, same file N frames) + teammate can help | "Carol, looks like you're waiting on auth. Alice is actively building it — hang tight." |
| `DEPENDENCY_READY` | Engineer A completes work that Engineer B was waiting on | "Carol, Bob — Alice just got the auth endpoint running. You're clear to integrate." |
| `DUPLICATE_WORK` | 2+ engineers on same file simultaneously | "Alice and Bob — you're both in login.tsx. Coordinate before pushing." |
---
## Continual learning story
The `ownership_map` collection persists across sessions. On Hermes startup:
1. Load ownership map for this pod from Atlas
2. Build in-memory cache: `Map<file, { primaryOwner, contributors }>`
3. Event detection uses priors immediately — no ramp-up phase
**Demo:** Session 1 takes 3 min to first nudge. Session 2 fires in < 30 seconds. That is the learning, visible on stage.
---
## Demo flow (3 min)
1. **(0:00)** Three engineers join pod. PodMan greets by voice.
2. **(0:20)** Alice opens `auth/middleware.ts`. Hermes infers ownership.
3. **(0:45)** Bob opens `frontend/login.tsx`. Carol's terminal shows connection refused.
4. **(1:20) BLOCKER_DETECTED:** "Carol, looks like you're waiting on auth. Alice is actively building it — hang tight."
5. **(2:00) DEPENDENCY_READY:** "Carol, Bob — Alice just got the auth endpoint running. You're clear to integrate."
6. **(2:20)** Optional: session 2 warm-start comparison.
7. **(2:45)** Close: "PodMan — the teammate that sees what Slack can't."
---
## Key risks
| Risk | Mitigation |
| -------------------------------------------- | ------------------------------------------------- |
| Gemini Vision accuracy | Large font, single editor window, confidence gate |
| Gemini Live 2.5 + LiveKit Agents integration | Build together hour 57, have TTS fallback |
| Frame POST latency | JPEG compression, target < 500ms |
| Event false positives | 3-min cooldown, pre-staged demo |
| DO deploy failure | Hermes runs local, PWA defaults to localhost:8787 |
+7 -1
View File
@@ -3,7 +3,13 @@ import tseslint from 'typescript-eslint';
export default tseslint.config(
{
ignores: ['**/dist/**', '**/build/**', '**/node_modules/**', '**/*.config.*'],
ignores: [
'**/dist/**',
'**/build/**',
'**/node_modules/**',
'**/.venv/**',
'**/*.config.*',
],
},
js.configs.recommended,
...tseslint.configs.recommended,
+21
View File
@@ -30,6 +30,27 @@
content="Ambient AI teammate that catches merge collisions before anyone pushes."
/>
<meta name="twitter:image" content="https://www.podman.live/og.png" />
<script>
// Tear down any service worker we shipped earlier (vite-plugin-pwa).
// PodMan runs no PWA in production: we deploy continuously and SW-driven
// reloads break the live demo. This unregisters leftover workers and
// clears their caches on every load so previously-affected browsers
// self-heal — no manual DevTools needed. It does NOT reload the page and
// is a no-op once nothing is registered.
if ('serviceWorker' in navigator) {
navigator.serviceWorker
.getRegistrations()
.then((regs) => regs.forEach((r) => r.unregister()))
.catch(() => {});
if (window.caches && caches.keys) {
caches
.keys()
.then((keys) => keys.forEach((k) => caches.delete(k)))
.catch(() => {});
}
}
</script>
</head>
<body>
<div id="root"></div>
+2
View File
@@ -11,6 +11,8 @@
},
"dependencies": {
"@base-ui/react": "^1.6.0",
"@clerk/react": "^6.11.1",
"@clerk/ui": "^1.23.0",
"@fontsource-variable/geist": "^5.2.9",
"@podman/shared": "workspace:*",
"@shadcn/react": "^0.1.0",
+191 -106
View File
@@ -1,15 +1,18 @@
import { useEffect, useState } from 'react';
import {
Show,
SignUp,
SignInButton,
SignUpButton,
UserButton,
useAuth,
useUser,
} from '@clerk/react';
import type { Room } from 'livekit-client';
import {
AlertCircleIcon,
BrainCircuitIcon,
CircleDotIcon,
RadioTowerIcon,
RefreshCwIcon,
SparklesIcon,
UsersIcon,
WifiIcon,
ShieldCheckIcon,
} from 'lucide-react';
import type { Pod, PodInput } from '@podman/shared';
import { joinPod } from './lib/pod.js';
@@ -19,7 +22,6 @@ import { CreatePodForm } from './components/CreatePodForm.js';
import { PodView } from './components/PodView.js';
import { GraphView } from './components/GraphView.js';
import { Alert, AlertDescription, AlertTitle } from '@/components/ui/alert';
import { Badge } from '@/components/ui/badge';
import { Button } from '@/components/ui/button';
import {
Card,
@@ -40,15 +42,39 @@ import {
import { Skeleton } from '@/components/ui/skeleton';
const SESSION_KEY = 'podman.session';
const fmt = new Intl.NumberFormat('en', { notation: 'compact' });
function pathPodId(): string | null {
const [segment] = window.location.pathname.split('/').filter(Boolean);
return segment ? decodeURIComponent(segment) : null;
}
function setPodPath(podId: string, replace = false): void {
const next = `/${encodeURIComponent(podId)}`;
if (window.location.pathname === next) return;
window.history[replace ? 'replaceState' : 'pushState']({}, '', next);
}
function setHomePath(): void {
if (window.location.pathname === '/') return;
window.history.pushState({}, '', '/');
}
function replacePath(path: string): void {
window.history.replaceState({}, '', path || '/');
}
function firstNameFrom(value: string | null | undefined): string {
return value?.trim().split(/\s+/).filter(Boolean)[0] ?? '';
}
export default function App() {
const { getToken, isLoaded, isSignedIn } = useAuth();
const { user } = useUser();
const [pods, setPods] = useState<Pod[]>([]);
const [loading, setLoading] = useState(true);
const [pending, setPending] = useState<Set<string>>(new Set());
const [error, setError] = useState<string | null>(null);
const [presence, setPresence] = useState<Record<string, string[]>>({});
const [memory, setMemory] = useState<api.MemoryStats | null>(null);
const [joinedPodId, setJoinedPodId] = useState<string | null>(null);
const [member, setMember] = useState('');
const [devMode, setDevMode] = useState(false);
@@ -57,18 +83,31 @@ export default function App() {
const [graphPodId, setGraphPodId] = useState<string | null>(null);
const joinedPod = joinedPodId ? (pods.find((p) => p.id === joinedPodId) ?? null) : null;
const userEmail = user?.primaryEmailAddress?.emailAddress;
const defaultMemberName =
user?.firstName?.trim() ||
firstNameFrom(user?.fullName) ||
firstNameFrom(userEmail?.split('@')[0]);
const currentUserProfile = {
displayName: defaultMemberName,
email: userEmail,
imageUrl: user?.imageUrl,
};
useEffect(() => {
api.setAuthTokenGetter(isSignedIn ? getToken : null);
return () => api.setAuthTokenGetter(null);
}, [getToken, isSignedIn]);
async function refresh() {
setLoading(true);
try {
const [nextPods, nextPresence, nextMemory] = await Promise.all([
const [nextPods, nextPresence] = await Promise.all([
api.listPods(),
api.getPresence().catch(() => presence),
api.getMemoryStats().catch(() => memory),
]);
setPods(nextPods);
setPresence(nextPresence);
setMemory(nextMemory);
setError(null);
} catch (e) {
setError((e as Error).message);
@@ -85,31 +124,44 @@ export default function App() {
return n;
});
async function connectToPod(podId: string, who: string) {
const result = await joinPod(podId, who, who);
async function connectToPod(podId: string, who: string, replaceRoute = false) {
const previousPath = window.location.pathname;
setPodPath(podId, replaceRoute);
try {
const result = await joinPod(
podId,
who,
who,
isSignedIn ? getToken : undefined,
currentUserProfile,
);
setRoom(result.room);
setDevMode(result.mode === 'dev');
setMember(who);
setJoinedPodId(podId);
sessionStorage.setItem(SESSION_KEY, JSON.stringify({ podId, member: who }));
} catch (e) {
replacePath(previousPath);
throw e;
}
}
useEffect(() => {
if (!isSignedIn) {
setLoading(false);
return;
}
void refresh();
}, []);
}, [isSignedIn]);
useEffect(() => {
if (joinedPodId) return;
if (!isSignedIn || joinedPodId) return;
let alive = true;
const tick = async () => {
try {
const [p, m] = await Promise.all([
api.getPresence(),
api.getMemoryStats().catch(() => memory),
]);
const p = await api.getPresence();
if (alive) {
setPresence(p);
setMemory(m);
}
} catch {
/* presence is best-effort */
@@ -121,11 +173,15 @@ export default function App() {
alive = false;
window.clearInterval(id);
};
}, [joinedPodId]);
}, [isSignedIn, joinedPodId]);
useEffect(() => {
if (!isSignedIn) return;
const routedPodId = pathPodId();
const raw = sessionStorage.getItem(SESSION_KEY);
if (!raw) return;
if (!raw) {
return;
}
let saved: { podId: string; member: string };
try {
saved = JSON.parse(raw);
@@ -133,17 +189,43 @@ export default function App() {
sessionStorage.removeItem(SESSION_KEY);
return;
}
const podId = routedPodId ?? saved.podId;
setRestoring(true);
void (async () => {
try {
await connectToPod(saved.podId, saved.member);
await connectToPod(podId, saved.member, !!routedPodId);
} catch {
sessionStorage.removeItem(SESSION_KEY);
} finally {
setRestoring(false);
}
})();
}, []);
}, [isSignedIn]);
useEffect(() => {
const onPopState = () => {
if (!isSignedIn) return;
const routedPodId = pathPodId();
if (!routedPodId) {
room?.disconnect();
setRoom(null);
setJoinedPodId(null);
return;
}
const raw = sessionStorage.getItem(SESSION_KEY);
if (!raw) return;
try {
const saved = JSON.parse(raw) as { member: string };
if (saved.member && routedPodId !== joinedPodId) {
void connectToPod(routedPodId, saved.member, true);
}
} catch {
sessionStorage.removeItem(SESSION_KEY);
}
};
window.addEventListener('popstate', onPopState);
return () => window.removeEventListener('popstate', onPopState);
}, [isSignedIn, joinedPodId, room]);
async function run(key: string, fn: () => Promise<void>) {
startPending(key);
@@ -163,7 +245,7 @@ export default function App() {
startPending('new');
setError(null);
try {
const created = await api.createPod(input);
const created = await api.createPod(input, currentUserProfile);
setPods((cur) => [...cur, created]);
} catch (e) {
setError((e as Error).message);
@@ -179,9 +261,10 @@ export default function App() {
run(id, async () => {
await api.deletePod(id);
setPods((cur) => cur.filter((x) => x.id !== id));
if (pathPodId() === id) setHomePath();
});
const handleAddMember = (id: string, name: string) =>
run(id, async () => upsert(await api.addMember(id, name)));
run(id, async () => upsert(await api.addMember(id, name, currentUserProfile)));
const handleRemoveMember = (id: string, name: string) =>
run(id, async () => upsert(await api.removeMember(id, name)));
@@ -197,12 +280,15 @@ export default function App() {
}
}
async function handleAddAndJoin(pod: Pod, name: string) {
async function handleAddAndJoin(pod: Pod) {
startPending(pod.id);
setError(null);
try {
upsert(await api.addMember(pod.id, name));
await connectToPod(pod.id, name);
if (!defaultMemberName) {
throw new Error('Sign in with Clerk before joining a pod.');
}
upsert(await api.addMember(pod.id, defaultMemberName, currentUserProfile));
await connectToPod(pod.id, defaultMemberName);
} catch (e) {
setError((e as Error).message);
} finally {
@@ -215,22 +301,34 @@ export default function App() {
setRoom(null);
setJoinedPodId(null);
sessionStorage.removeItem(SESSION_KEY);
setHomePath();
void refresh();
}
const showReconnecting = restoring || (joinedPodId !== null && joinedPod === null);
const liveNames = Array.from(new Set(Object.values(presence).flat()));
const liveTotal = liveNames.length;
const totalMembers = pods.reduce((sum, p) => sum + p.members.length, 0);
const activeRooms = Object.values(presence).filter((names) => names.length > 0).length;
const podManOnline = liveNames.some((name) => name.toLowerCase() === 'podman');
const latestActivity = memory
? memory.observations + memory.collisions + memory.interventions + memory.outcomes
: 0;
if (!isLoaded) {
return (
<div className="grid min-h-screen place-items-center bg-background text-foreground">
<Skeleton className="h-12 w-64" />
</div>
);
}
if (!isSignedIn) {
return <AuthGate />;
}
if (joinedPod) {
return (
<PodView team={joinedPod} me={member} room={room} devMode={devMode} onLeave={handleLeave} />
<PodView
team={joinedPod}
me={member}
room={room}
devMode={devMode}
currentUserProfile={currentUserProfile}
onLeave={handleLeave}
/>
);
}
@@ -241,46 +339,35 @@ export default function App() {
return (
<div className="min-h-screen bg-background text-foreground">
<div className="mx-auto flex min-h-screen w-full max-w-[1440px] flex-col gap-6 px-4 py-4 sm:px-6 lg:px-8">
<header className="sticky top-0 z-10 -mx-4 flex flex-col gap-5 border-b bg-background/86 px-4 pb-5 pt-2 backdrop-blur-xl sm:-mx-6 sm:px-6 lg:-mx-8 lg:px-8">
<div className="flex flex-col gap-4 lg:flex-row lg:items-end lg:justify-between">
<header className="sticky top-0 z-10 -mx-4 border-b bg-background/86 px-4 pb-4 pt-2 backdrop-blur-xl sm:-mx-6 sm:px-6 lg:-mx-8 lg:px-8">
<div className="flex flex-col gap-4 sm:flex-row sm:items-center sm:justify-between">
<div className="flex min-w-0 items-center gap-3">
<div className="grid size-10 place-items-center rounded-lg bg-primary text-sm font-semibold text-primary-foreground shadow-sm">
PM
</div>
<div className="min-w-0">
<div className="flex flex-wrap items-center gap-2">
<h1 className="text-[1.95rem] font-semibold leading-none tracking-tight">
PodMan
</h1>
<Badge variant={podManOnline ? 'default' : 'secondary'}>
<CircleDotIcon data-icon="inline-start" />
{podManOnline ? 'online' : 'standby'}
</Badge>
</div>
<p className="text-sm text-muted-foreground">
Live engineering rooms, team memory, and intervention routing.
</p>
<h1 className="text-[1.95rem] font-semibold leading-none tracking-tight">PodMan</h1>
</div>
</div>
<div className="flex items-center gap-2">
<Badge variant="outline" className="h-8 rounded-lg px-3">
<ShieldCheckIcon data-icon="inline-start" />
Privacy-limited
</Badge>
<Show when="signed-out">
<SignInButton mode="modal">
<Button variant="outline">Sign in</Button>
</SignInButton>
<SignUpButton mode="modal">
<Button>Sign up</Button>
</SignUpButton>
</Show>
<Show when="signed-in">
<UserButton />
</Show>
<Button variant="outline" onClick={() => void refresh()} disabled={loading}>
<RefreshCwIcon data-icon="inline-start" />
Refresh
</Button>
</div>
</div>
<div className="grid gap-2 sm:grid-cols-2 lg:grid-cols-4">
<StatPill icon={WifiIcon} label="Live" value={fmt.format(liveTotal)} />
<StatPill icon={RadioTowerIcon} label="Rooms" value={fmt.format(activeRooms)} />
<StatPill icon={UsersIcon} label="Roster" value={fmt.format(totalMembers)} />
<StatPill icon={BrainCircuitIcon} label="Memory" value={fmt.format(latestActivity)} />
</div>
</header>
{error && (
@@ -314,9 +401,6 @@ export default function App() {
<p className="text-xs font-medium uppercase text-muted-foreground">Workspaces</p>
<h2 className="text-xl font-semibold tracking-tight">Active pods</h2>
</div>
<p className="max-w-xl text-sm leading-6 text-muted-foreground">
Join the room that matches your current workstream.
</p>
</div>
{loading ? (
@@ -332,6 +416,7 @@ export default function App() {
pod={pod}
busy={pending.has(pod.id)}
presence={presence[pod.id] ?? []}
currentUserProfile={currentUserProfile}
onJoin={handleJoin}
onAddAndJoin={handleAddAndJoin}
onAddMember={handleAddMember}
@@ -352,29 +437,23 @@ export default function App() {
<EmptyDescription>Create the first room for this team.</EmptyDescription>
</EmptyHeader>
<EmptyContent>
<CreatePodForm busy={pending.has('new')} onCreate={handleCreate} compact />
<CreatePodForm
busy={pending.has('new')}
defaultMemberName={defaultMemberName}
onCreate={handleCreate}
compact
/>
</EmptyContent>
</Empty>
)}
</section>
<aside className="flex flex-col gap-4">
<CreatePodForm busy={pending.has('new')} onCreate={handleCreate} />
<Card>
<CardHeader>
<CardTitle>Operating brief</CardTitle>
<CardDescription>Current coordination signals.</CardDescription>
</CardHeader>
<CardContent className="space-y-3">
<BriefLine label="Notification default" value="Card" />
<BriefLine label="Escalation" value="Hermes voice only when urgent" />
<BriefLine label="Memory events" value={fmt.format(latestActivity)} />
<BriefLine
label="Live people"
value={liveNames.length ? liveNames.join(', ') : 'None'}
<CreatePodForm
busy={pending.has('new')}
defaultMemberName={defaultMemberName}
onCreate={handleCreate}
/>
</CardContent>
</Card>
</aside>
</main>
)}
@@ -383,33 +462,39 @@ export default function App() {
);
}
function StatPill({
icon: Icon,
label,
value,
}: {
icon: React.ComponentType<{ className?: string }>;
label: string;
value: string;
}) {
function AuthGate() {
return (
<div className="flex min-h-16 items-center gap-3 rounded-lg border bg-card/90 px-3 py-2 shadow-sm">
<div className="grid size-8 place-items-center rounded-md bg-muted">
<Icon className="size-4 text-muted-foreground" />
<div className="min-h-screen bg-background text-foreground">
<div className="mx-auto flex min-h-screen w-full max-w-[1120px] flex-col px-4 py-4 sm:px-6 lg:px-8">
<header className="flex items-center justify-between border-b pb-4 pt-2">
<div className="flex min-w-0 items-center gap-3">
<div className="grid size-10 place-items-center rounded-lg bg-primary text-sm font-semibold text-primary-foreground shadow-sm">
PM
</div>
<div className="min-w-0">
<p className="text-xs font-medium uppercase text-muted-foreground">{label}</p>
<p className="text-base font-medium">{value}</p>
<h1 className="text-[1.95rem] font-semibold leading-none tracking-tight">PodMan</h1>
</div>
</div>
);
}
<SignInButton mode="modal">
<Button variant="outline">Sign in</Button>
</SignInButton>
</header>
function BriefLine({ label, value }: { label: string; value: string }) {
return (
<div className="flex items-start justify-between gap-4 rounded-md bg-muted/45 px-3 py-2">
<span className="text-sm text-muted-foreground">{label}</span>
<span className="max-w-44 text-right text-sm font-medium">{value}</span>
<main className="grid flex-1 items-center gap-8 py-8 lg:grid-cols-[minmax(0,1fr)_420px]">
<section className="max-w-xl">
<p className="text-xs font-medium uppercase text-muted-foreground">Team memory</p>
<h2 className="mt-2 text-3xl font-semibold tracking-tight">
Create your account to enter PodMan
</h2>
<p className="mt-3 text-base text-muted-foreground">
PodMan saves your context across pods so agents can learn from your work in every
room you join.
</p>
</section>
<div className="flex justify-center lg:justify-end">
<SignUp routing="hash" />
</div>
</main>
</div>
</div>
);
}
+4 -2
View File
@@ -17,17 +17,19 @@ import { cn } from '@/lib/utils';
export function CreatePodForm({
busy,
onCreate,
defaultMemberName = '',
compact = false,
}: {
busy: boolean;
onCreate: (input: PodInput) => Promise<void>;
defaultMemberName?: string;
compact?: boolean;
}) {
const [open, setOpen] = useState(compact);
const [name, setName] = useState('');
const [repo, setRepo] = useState('karti-ai/podman');
const [description, setDescription] = useState('');
const [firstMember, setFirstMember] = useState('');
const [firstMember, setFirstMember] = useState(defaultMemberName);
async function submit() {
if (!name.trim()) return;
@@ -40,7 +42,7 @@ export function CreatePodForm({
});
setName('');
setDescription('');
setFirstMember('');
setFirstMember(defaultMemberName);
if (!compact) setOpen(false);
} catch {
/* parent owns the visible error */
+1 -5
View File
@@ -169,11 +169,7 @@ export function GraphView({ podId, onClose }: { podId: string; onClose: () => vo
/>
</div>
{graph.loop?.length ? (
<LearningLoop stages={graph.loop} />
) : (
<div />
)}
{graph.loop?.steps?.length ? <LearningLoop loop={graph.loop} /> : <div />}
</div>
{/* Activity stream · selected node */}
+36 -52
View File
@@ -2,13 +2,18 @@ import { useState } from 'react';
import {
BrainCircuitIcon,
MoreHorizontalIcon,
PlusIcon,
Trash2Icon,
UserRoundIcon,
VideoIcon,
} from 'lucide-react';
import type { Pod, PodInput } from '@podman/shared';
import { Avatar, AvatarBadge, AvatarFallback, AvatarGroup } from '@/components/ui/avatar';
import {
Avatar,
AvatarBadge,
AvatarFallback,
AvatarGroup,
AvatarImage,
} from '@/components/ui/avatar';
import { Badge } from '@/components/ui/badge';
import { Button } from '@/components/ui/button';
import {
@@ -16,7 +21,6 @@ import {
CardAction,
CardContent,
CardDescription,
CardFooter,
CardHeader,
CardTitle,
} from '@/components/ui/card';
@@ -43,9 +47,10 @@ export function PodCard({
pod,
busy,
presence,
onJoin,
currentUserProfile,
onJoin: _onJoin,
onAddAndJoin,
onAddMember,
onAddMember: _onAddMember,
onRemoveMember: _onRemoveMember,
onUpdate,
onDelete,
@@ -54,15 +59,19 @@ export function PodCard({
pod: Pod;
busy: boolean;
presence: string[];
currentUserProfile?: {
displayName: string;
email?: string;
imageUrl?: string;
};
onJoin: (pod: Pod, member: string) => void;
onAddAndJoin: (pod: Pod, name: string) => void;
onAddAndJoin: (pod: Pod) => void;
onAddMember: (id: string, name: string) => void;
onRemoveMember: (id: string, name: string) => void;
onUpdate: (id: string, patch: PodInput) => void;
onDelete: (id: string) => void;
onOpenGraph: (id: string) => void;
}) {
const [newMember, setNewMember] = useState('');
const [editing, setEditing] = useState(false);
const [draft, setDraft] = useState<PodInput>({
name: pod.name,
@@ -71,8 +80,14 @@ export function PodCard({
});
const inRoom = (name: string) => presence.some((p) => p.toLowerCase() === name.toLowerCase());
const profileForMember = (name: string) =>
pod.memberProfiles?.[name] ??
([currentUserProfile?.displayName, currentUserProfile?.email]
.filter(Boolean)
.some((value) => value?.toLowerCase() === name.toLowerCase())
? currentUserProfile
: undefined);
const active = presence.length > 0;
const primaryMember = pod.members[0] ?? '';
function saveEdit() {
onUpdate(pod.id, {
@@ -83,12 +98,8 @@ export function PodCard({
setEditing(false);
}
function submitMember(join: boolean) {
const name = newMember.trim();
if (!name) return;
if (join) onAddAndJoin(pod, name);
else onAddMember(pod.id, name);
setNewMember('');
function join() {
onAddAndJoin(pod);
}
return (
@@ -145,12 +156,16 @@ export function PodCard({
<div className="flex items-center justify-between gap-4">
<AvatarGroup>
{pod.members.slice(0, 4).map((member) => (
<Avatar key={member} title={member}>
{pod.members.slice(0, 4).map((member) => {
const profile = profileForMember(member);
return (
<Avatar key={member} title={profile?.email ?? member}>
{profile?.imageUrl && <AvatarImage src={profile.imageUrl} alt={member} />}
<AvatarFallback>{initials(member)}</AvatarFallback>
{inRoom(member) && <AvatarBadge />}
</Avatar>
))}
);
})}
{pod.members.length > 4 && <span className="text-sm text-muted-foreground">+</span>}
</AvatarGroup>
<div className="flex items-center gap-1.5 text-sm text-muted-foreground">
@@ -159,45 +174,14 @@ export function PodCard({
</div>
</div>
<div className="flex gap-2">
<Input
placeholder="Your name"
value={newMember}
onChange={(e) => setNewMember(e.target.value)}
onKeyDown={(e) => {
if (e.key === 'Enter') submitMember(true);
}}
/>
<Button
variant="outline"
size="icon"
onClick={() => submitMember(false)}
disabled={busy || !newMember.trim()}
>
<PlusIcon />
<span className="sr-only">Add member</span>
<div className="flex justify-end">
<Button className="min-w-24" onClick={join} disabled={busy}>
<VideoIcon data-icon="inline-start" />
Join
</Button>
</div>
</div>
</CardContent>
<CardFooter className="justify-between gap-2">
<Button
variant="outline"
onClick={() => primaryMember && onJoin(pod, primaryMember)}
disabled={busy || !primaryMember}
>
<VideoIcon data-icon="inline-start" />
Join
</Button>
<Button
className="min-w-28"
onClick={() => submitMember(true)}
disabled={busy || !newMember.trim()}
>
Add and join
</Button>
</CardFooter>
</Card>
<Dialog open={editing} onOpenChange={setEditing}>
File diff suppressed because it is too large Load Diff
@@ -1,4 +1,4 @@
import type { ActivityEvent } from '@podman/shared';
import type { PodGraphActivity } from '@podman/shared';
import { ScrollArea } from '@/components/ui/scroll-area';
import { ACTIVITY_TAG } from './encoding.js';
@@ -9,7 +9,7 @@ function timeOf(at: string): string {
return Number.isFinite(t) ? fmtTime.format(t) : '--:--';
}
export function ActivityStream({ events }: { events: ActivityEvent[] }) {
export function ActivityStream({ events }: { events: PodGraphActivity[] }) {
return (
<div className="flex h-full flex-col">
<p className="mb-2 text-xs font-medium uppercase tracking-wide text-muted-foreground">
@@ -33,7 +33,14 @@ export function ActivityStream({ events }: { events: ActivityEvent[] }) {
>
{tag.label}
</span>
<span className="min-w-0 flex-1 leading-snug text-foreground/90">{e.text}</span>
<span className="min-w-0 flex-1 leading-snug">
<span className="text-foreground/90">{e.title}</span>
{e.detail && (
<span className="block text-xs leading-snug text-muted-foreground">
{e.detail}
</span>
)}
</span>
</li>
);
})}
+17 -11
View File
@@ -1,43 +1,49 @@
import type { LearningStage } from '@podman/shared';
import type { PodLearningLoop } from '@podman/shared';
import { BLUE } from './encoding.js';
/**
* The continual-learning loop rail: observe store predict outcome adapt.
* The active stage (most-recent activity) gets a pulsing accent bar + ring.
*/
export function LearningLoop({ stages }: { stages: LearningStage[] }) {
export function LearningLoop({ loop }: { loop: PodLearningLoop }) {
return (
<div className="space-y-1">
<p className="mb-2 text-xs font-medium uppercase tracking-wide text-muted-foreground">
Learning loop
</p>
{stages.map((s, i) => (
{loop.steps.map((s, i) => {
const active = s.status === 'active' || s.key === loop.activeStep;
return (
<div key={s.key}>
<div
className="relative overflow-hidden rounded-lg border bg-card py-2 pl-3.5 pr-3 shadow-sm transition-colors data-[active=true]:bg-accent/40"
data-active={s.active}
style={s.active ? { boxShadow: `inset 0 0 0 1px ${BLUE}55` } : undefined}
data-active={active}
style={active ? { boxShadow: `inset 0 0 0 1px ${BLUE}55` } : undefined}
>
<span
aria-hidden
className={`absolute inset-y-0 left-0 w-1 ${s.active ? 'pm-pulse' : ''}`}
style={{ background: s.active ? BLUE : 'var(--border)' }}
className={`absolute inset-y-0 left-0 w-1 ${active ? 'pm-pulse' : ''}`}
style={{ background: active ? BLUE : 'var(--border)' }}
/>
<div className="flex items-baseline justify-between gap-2">
<p className="text-[0.7rem] font-medium uppercase tracking-wide text-muted-foreground">
<span className="tabular-nums">{String(i + 1).padStart(2, '0')}</span> {s.title}
<span className="tabular-nums">{String(i + 1).padStart(2, '0')}</span> {s.label}
</p>
<p className="font-heading text-sm font-semibold tabular-nums">{s.value}</p>
</div>
<p className="mt-0.5 text-xs leading-snug text-muted-foreground">{s.detail}</p>
</div>
{i < stages.length - 1 && (
<p aria-hidden className="py-0.5 text-center text-xs leading-none text-muted-foreground/60">
{i < loop.steps.length - 1 && (
<p
aria-hidden
className="py-0.5 text-center text-xs leading-none text-muted-foreground/60"
>
</p>
)}
</div>
))}
);
})}
</div>
);
}
+6 -4
View File
@@ -4,7 +4,7 @@ import type {
PodGraphNode,
PodGraphEdge,
PodGraphNodeKind,
ActivityKind,
PodGraphActivityKind,
} from '@podman/shared';
/**
@@ -21,12 +21,14 @@ export const VIOLET = '#7c3aed';
export const GREEN = '#16a34a';
/** Tag color + short label per activity-stream kind. */
export const ACTIVITY_TAG: Record<ActivityKind, { color: string; label: string }> = {
export const ACTIVITY_TAG: Record<PodGraphActivityKind, { color: string; label: string }> = {
editing: { color: SLATE, label: 'EDITING' },
collision: { color: RED, label: 'COLLISION' },
warns: { color: AMBER, label: 'WARNS' },
intervention: { color: AMBER, label: 'NUDGE' },
outcome: { color: GREEN, label: 'OUTCOME' },
learned_from: { color: VIOLET, label: 'LEARNED' },
learned: { color: VIOLET, label: 'LEARNED' },
agent: { color: BLUE, label: 'AGENT' },
suppressed: { color: VIOLET, label: 'SUPPRESSED' },
};
export const KIND_COLOR: Record<PodGraphNodeKind, string> = {
+63
View File
@@ -0,0 +1,63 @@
import { useEffect, useMemo, useState } from 'react';
import type { PodActivityEvent } from '@podman/shared';
import { getPodActivity, podActivityStreamUrl } from '../lib/api';
export function usePodActivity(podId: string | null, me: string) {
const [events, setEvents] = useState<PodActivityEvent[]>([]);
const [connected, setConnected] = useState(false);
const [error, setError] = useState<string | null>(null);
useEffect(() => {
if (!podId) return;
let alive = true;
const load = async () => {
try {
const snapshot = await getPodActivity(podId);
if (alive) {
setEvents(snapshot);
setError(null);
}
} catch (e) {
if (alive) setError((e as Error).message);
}
};
void load();
const source = new EventSource(podActivityStreamUrl(podId));
source.addEventListener('open', () => {
if (alive) setConnected(true);
});
source.addEventListener('snapshot', (event) => {
if (!alive) return;
setEvents(JSON.parse((event as MessageEvent<string>).data) as PodActivityEvent[]);
setConnected(true);
setError(null);
});
source.addEventListener('error', () => {
if (alive) {
setConnected(false);
setError('Realtime activity stream reconnecting');
}
});
return () => {
alive = false;
source.close();
};
}, [podId]);
return useMemo(() => {
const mine = events.filter((event) => belongsTo(event, me));
const team = events.filter((event) => !belongsTo(event, me));
return { events, mine, team, connected, error };
}, [connected, error, events, me]);
}
function belongsTo(event: PodActivityEvent, me: string): boolean {
const normalized = me.trim().toLowerCase();
if (!normalized) return false;
const names = [event.actor, ...(event.actors ?? [])]
.filter(Boolean)
.map((name) => name!.trim().toLowerCase());
return names.includes(normalized);
}
+168 -22
View File
@@ -1,4 +1,12 @@
import type { InterventionOutcome, Pod, PodInput } from '@podman/shared';
import type {
InterventionOutcome,
HermesJob,
HermesJobEvent,
MemberWorkHistory,
Pod,
PodActivityEvent,
PodInput,
} from '@podman/shared';
const BACKEND_URL =
import.meta.env.VITE_BACKEND_URL ||
@@ -6,11 +14,53 @@ const BACKEND_URL =
? 'http://localhost:8787'
: '');
type AuthTokenGetter = () => Promise<string | null>;
let authTokenGetter: AuthTokenGetter | null = null;
export function setAuthTokenGetter(getter: AuthTokenGetter | null): void {
authTokenGetter = getter;
}
async function requestHeaders(init?: HeadersInit): Promise<Headers> {
const next = new Headers(init);
const token = await authTokenGetter?.();
if (token) next.set('authorization', `Bearer ${token}`);
return next;
}
async function apiFetch(input: string, init: RequestInit = {}): Promise<Response> {
return fetch(input, {
...init,
headers: await requestHeaders(init.headers),
});
}
export interface MemoryStats {
observations: number;
collisions: number;
interventions: number;
outcomes: number;
userPodContext?: number;
}
export interface LiveConversationSession {
sessionId: string;
podId: string;
identity: string;
displayName: string;
room: string;
url: string;
token: string;
startedAt: string;
lastEventAt?: string;
endedAt?: string;
}
export interface UserProfilePayload {
displayName?: string;
email?: string;
imageUrl?: string;
}
async function json<T>(res: Response): Promise<T> {
@@ -21,16 +71,19 @@ async function json<T>(res: Response): Promise<T> {
return res.json() as Promise<T>;
}
const JSON_HEADERS = { 'content-type': 'application/json' } as const;
/** Mint a LiveKit token from the backend. */
export async function fetchToken(params: {
room: string;
identity: string;
name: string;
githubLogin?: string;
profile?: UserProfilePayload;
}): Promise<{ token: string; url: string }> {
const res = await fetch(`${BACKEND_URL}/api/token`, {
const res = await apiFetch(`${BACKEND_URL}/api/token`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
headers: JSON_HEADERS,
body: JSON.stringify(params),
});
return json(res);
@@ -38,9 +91,9 @@ export async function fetchToken(params: {
/** Record an intervention outcome for the policy learning loop. */
export async function postOutcome(outcome: InterventionOutcome): Promise<void> {
const res = await fetch(`${BACKEND_URL}/api/outcome`, {
const res = await apiFetch(`${BACKEND_URL}/api/outcome`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
headers: JSON_HEADERS,
body: JSON.stringify(outcome),
});
if (!res.ok) throw new Error(`outcome post failed: ${res.status}`);
@@ -52,9 +105,9 @@ export async function createSyncPr(input: {
summary?: string;
}): Promise<{ url: string; number: number }> {
return json(
await fetch(`${BACKEND_URL}/api/sync-pr`, {
await apiFetch(`${BACKEND_URL}/api/sync-pr`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
headers: JSON_HEADERS,
body: JSON.stringify(input),
}),
);
@@ -63,58 +116,151 @@ export async function createSyncPr(input: {
// --- Pods CRUD ---
export async function listPods(): Promise<Pod[]> {
return json(await fetch(`${BACKEND_URL}/api/pods`));
return json(await apiFetch(`${BACKEND_URL}/api/pods`));
}
/** Display names currently connected per pod id (= LiveKit room name). */
export async function getPresence(): Promise<Record<string, string[]>> {
return json(await fetch(`${BACKEND_URL}/api/presence`));
return json(await apiFetch(`${BACKEND_URL}/api/presence`));
}
export async function getMemoryStats(): Promise<MemoryStats> {
return json(await fetch(`${BACKEND_URL}/api/memory/stats`));
return json(await apiFetch(`${BACKEND_URL}/api/memory/stats`));
}
export async function createPod(input: PodInput): Promise<Pod> {
export async function getPodActivity(id: string, limit = 80): Promise<PodActivityEvent[]> {
return json(
await fetch(`${BACKEND_URL}/api/pods`, {
await apiFetch(`${BACKEND_URL}/api/pods/${encodeURIComponent(id)}/activity?limit=${limit}`),
);
}
export function podActivityStreamUrl(id: string): string {
return `${BACKEND_URL}/api/pods/${encodeURIComponent(id)}/activity/stream`;
}
/** URL of the pod's generated background-music MP3 (looped client-side). */
export function podMusicUrl(id: string): string {
return `${BACKEND_URL}/api/pods/${encodeURIComponent(id)}/music`;
}
export async function getMemberWorkHistory(
podId: string,
member: string,
): Promise<MemberWorkHistory> {
return json(
await apiFetch(
`${BACKEND_URL}/api/pods/${encodeURIComponent(podId)}/members/${encodeURIComponent(
member,
)}/history?hours=24&limit=80`,
),
);
}
export async function createPod(input: PodInput, profile?: UserProfilePayload): Promise<Pod> {
return json(
await apiFetch(`${BACKEND_URL}/api/pods`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify(input),
headers: JSON_HEADERS,
body: JSON.stringify({ ...input, profile }),
}),
);
}
export async function updatePod(id: string, patch: PodInput): Promise<Pod> {
return json(
await fetch(`${BACKEND_URL}/api/pods/${encodeURIComponent(id)}`, {
await apiFetch(`${BACKEND_URL}/api/pods/${encodeURIComponent(id)}`, {
method: 'PATCH',
headers: { 'content-type': 'application/json' },
headers: JSON_HEADERS,
body: JSON.stringify(patch),
}),
);
}
export async function deletePod(id: string): Promise<void> {
const res = await fetch(`${BACKEND_URL}/api/pods/${encodeURIComponent(id)}`, {
const res = await apiFetch(`${BACKEND_URL}/api/pods/${encodeURIComponent(id)}`, {
method: 'DELETE',
});
if (!res.ok) throw new Error(`delete pod failed: ${res.status}`);
}
export async function addMember(id: string, name: string): Promise<Pod> {
export async function addMember(
id: string,
name: string,
profile?: UserProfilePayload,
): Promise<Pod> {
return json(
await fetch(`${BACKEND_URL}/api/pods/${encodeURIComponent(id)}/members`, {
await apiFetch(`${BACKEND_URL}/api/pods/${encodeURIComponent(id)}/members`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ name }),
headers: JSON_HEADERS,
body: JSON.stringify({ name, profile }),
}),
);
}
export async function testPodVoice(id: string): Promise<void> {
const res = await apiFetch(`${BACKEND_URL}/api/pods/${encodeURIComponent(id)}/voice-test`, {
method: 'POST',
headers: JSON_HEADERS,
body: JSON.stringify({
message: 'PodMan voice test. Gemini TTS is playing through LiveKit.',
}),
});
if (!res.ok) throw new Error(`voice test failed: ${res.status}`);
}
export async function startLiveConversation(
podId: string,
input: { identity: string; displayName?: string },
): Promise<LiveConversationSession> {
return json(
await apiFetch(`${BACKEND_URL}/api/pods/${encodeURIComponent(podId)}/live-conversation/start`, {
method: 'POST',
headers: JSON_HEADERS,
body: JSON.stringify(input),
}),
);
}
export async function stopLiveConversation(podId: string, sessionId: string): Promise<void> {
const res = await apiFetch(
`${BACKEND_URL}/api/pods/${encodeURIComponent(
podId,
)}/live-conversation/${encodeURIComponent(sessionId)}/stop`,
{ method: 'POST' },
);
if (!res.ok) throw new Error(`live conversation stop failed: ${res.status}`);
}
export async function getLiveConversationHermesJob(
podId: string,
sessionId: string,
): Promise<{ job: HermesJob | null; events: HermesJobEvent[] }> {
return json(
await apiFetch(
`${BACKEND_URL}/api/pods/${encodeURIComponent(
podId,
)}/live-conversation/${encodeURIComponent(sessionId)}/hermes-job`,
),
);
}
export async function abortLiveConversationHermesJob(
podId: string,
sessionId: string,
): Promise<{ job: HermesJob | null }> {
return json(
await apiFetch(
`${BACKEND_URL}/api/pods/${encodeURIComponent(
podId,
)}/live-conversation/${encodeURIComponent(sessionId)}/hermes-job/abort`,
{ method: 'POST' },
),
);
}
export async function removeMember(id: string, name: string): Promise<Pod> {
return json(
await fetch(
await apiFetch(
`${BACKEND_URL}/api/pods/${encodeURIComponent(id)}/members/${encodeURIComponent(name)}`,
{ method: 'DELETE' },
),
+39
View File
@@ -75,3 +75,42 @@ export function startBeat(): BeatHandle {
},
};
}
/**
* Load an MP3 from `url` and loop it as an audio MediaStreamTrack to publish into
* a LiveKit room (and play on local speakers). Used for pod background music
* (Lyria-generated). Pure Web Audio no asset bundling.
*/
export async function startMusic(url: string): Promise<BeatHandle> {
const ctx = new AudioContext();
await ctx.resume();
const res = await fetch(url);
if (!res.ok) throw new Error(`music fetch failed: ${res.status}`);
const buffer = await ctx.decodeAudioData(await res.arrayBuffer());
const dest = ctx.createMediaStreamDestination();
const master = ctx.createGain();
master.gain.value = 0.6;
master.connect(dest); // -> published track (remote listeners)
master.connect(ctx.destination); // -> local speakers (publisher)
const src = ctx.createBufferSource();
src.buffer = buffer;
src.loop = true;
src.connect(master);
src.start();
const track = dest.stream.getAudioTracks()[0]!;
return {
track,
stop: () => {
try {
src.stop();
} catch {
/* already stopped */
}
track.stop();
void ctx.close();
},
};
}
+16 -4
View File
@@ -11,11 +11,17 @@ export async function fetchPodToken(
podId: string,
identity: string,
name: string,
getToken?: () => Promise<string | null>,
profile?: { displayName?: string; email?: string; imageUrl?: string },
): Promise<{ token: string; url: string }> {
const clerkToken = await getToken?.();
const res = await fetch(`${BACKEND_URL}/api/token`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ room: podId, identity, name }),
headers: {
'content-type': 'application/json',
...(clerkToken ? { authorization: `Bearer ${clerkToken}` } : {}),
},
body: JSON.stringify({ room: podId, identity, name, profile }),
});
if (!res.ok) throw new Error(`token request failed: ${res.status}`);
return res.json();
@@ -33,8 +39,14 @@ export type JoinResult = { mode: 'live'; room: Room } | { mode: 'dev'; room: nul
* connected screen sharing is a separate, deliberate action (see PodView) so
* a denied/slow screen prompt never blocks or fails the join.
*/
export async function joinPod(podId: string, identity: string, name: string): Promise<JoinResult> {
const { token, url } = await fetchPodToken(podId, identity, name);
export async function joinPod(
podId: string,
identity: string,
name: string,
getToken?: () => Promise<string | null>,
profile?: { displayName?: string; email?: string; imageUrl?: string },
): Promise<JoinResult> {
const { token, url } = await fetchPodToken(podId, identity, name, getToken, profile);
if (!isLiveKitConfigured(url)) {
console.warn('[podman] LiveKit not configured — dev mock join');
+150
View File
@@ -0,0 +1,150 @@
import { useCallback, useEffect, useRef, useState } from 'react';
import { RoomEvent, type Room } from 'livekit-client';
import { DATA_TOPIC, type DataMessage } from '@podman/shared';
import { startBeat, startMusic, type BeatHandle } from '../lib/beat.js';
/** Name of the test-audio track; its presence in the room IS the shared state. */
export const BEAT_TRACK = 'podman-beat';
export interface BeatState {
/** Is the test audio playing anywhere in the pod? */
on: boolean;
/** Display name of the participant who started it (the owner). */
by: string | null;
/** Do I own the beat track (so I can stop it directly)? */
mine: boolean;
}
const OFF: BeatState = { on: false, by: null, mine: false };
/**
* Shared, pod-wide test audio. One participant publishes the `podman-beat`
* track; everyone hears it and sees the same on/off state, derived directly
* from the track's presence (self-syncing across joins/leaves). Any participant
* can stop it: non-owners send BEAT_STOP and the owner unpublishes.
*/
export function useBeat(room: Room | null, musicUrl?: string) {
const [beat, setBeat] = useState<BeatState>(OFF);
const beatRef = useRef<BeatHandle | null>(null);
const stopLocal = useCallback(async () => {
const handle = beatRef.current;
if (!handle) return;
beatRef.current = null;
try {
await room?.localParticipant.unpublishTrack(handle.track);
} finally {
handle.stop();
}
}, [room]);
// Derive the shared state from the podman-beat track across all participants.
useEffect(() => {
if (!room) {
setBeat(OFF);
return;
}
const recompute = () => {
const lp = room.localParticipant;
const localPub = [...lp.trackPublications.values()].find((p) => p.trackName === BEAT_TRACK);
if (localPub) {
setBeat({ on: true, by: lp.name || lp.identity, mine: true });
return;
}
for (const p of room.remoteParticipants.values()) {
const pub = [...p.trackPublications.values()].find((tp) => tp.trackName === BEAT_TRACK);
if (pub) {
setBeat({ on: true, by: p.name || p.identity, mine: false });
return;
}
}
setBeat(OFF);
};
recompute();
const events = [
RoomEvent.LocalTrackPublished,
RoomEvent.LocalTrackUnpublished,
RoomEvent.TrackPublished,
RoomEvent.TrackUnpublished,
RoomEvent.TrackSubscribed,
RoomEvent.TrackUnsubscribed,
RoomEvent.ParticipantConnected,
RoomEvent.ParticipantDisconnected,
] as const;
events.forEach((e) => room.on(e, recompute));
return () => {
events.forEach((e) => room.off(e, recompute));
};
}, [room]);
// The owner honors stop requests from any participant.
useEffect(() => {
if (!room) return;
const onData = (payload: Uint8Array, _p: unknown, _k: unknown, topic?: string) => {
if (topic !== DATA_TOPIC) return;
let msg: DataMessage;
try {
msg = JSON.parse(new TextDecoder().decode(payload)) as DataMessage;
} catch {
return;
}
if (msg.type === 'BEAT_STOP') void stopLocal();
};
room.on(RoomEvent.DataReceived, onData);
return () => {
room.off(RoomEvent.DataReceived, onData);
};
}, [room, stopLocal]);
// Tear down my own track on unmount (e.g. leaving the pod).
const stopLocalRef = useRef(stopLocal);
stopLocalRef.current = stopLocal;
const startingRef = useRef(false);
const unmountedRef = useRef(false);
useEffect(() => {
return () => {
unmountedRef.current = true;
void stopLocalRef.current();
};
}, []);
const toggleBeat = useCallback(async () => {
if (!room) return;
// `beatRef` is the synchronous source of truth for "do I own it" — `beat.mine`
// lags behind the LiveKit track events that recompute it, so gate on the ref.
if (beatRef.current) {
await stopLocal();
return;
}
if (beat.on) {
// Someone else owns it — can't unpublish their track, so ask them to stop.
await room.startAudio().catch(() => {});
await room.localParticipant.publishData(
new TextEncoder().encode(JSON.stringify({ type: 'BEAT_STOP' } satisfies DataMessage)),
{ reliable: true, topic: DATA_TOPIC },
);
return;
}
// Start it. Guard against rapid double-clicks publishing two tracks before the
// LocalTrackPublished event has had a chance to update state.
if (startingRef.current) return;
startingRef.current = true;
try {
await room.startAudio().catch(() => {}); // unlock playback from this gesture
if (unmountedRef.current) return;
const handle = musicUrl ? await startMusic(musicUrl) : startBeat();
beatRef.current = handle;
await room.localParticipant.publishTrack(handle.track, { name: BEAT_TRACK });
if (unmountedRef.current) await stopLocal(); // left mid-publish — clean up
} catch (e) {
beatRef.current?.stop();
beatRef.current = null;
throw e;
} finally {
startingRef.current = false;
}
}, [room, beat, stopLocal, musicUrl]);
return { beat, toggleBeat };
}
+35 -2
View File
@@ -4,6 +4,27 @@ import type { DataMessage, HermesMessage, Intervention, InterventionStatus } fro
import { DATA_TOPIC } from '@podman/shared';
import { createSyncPr, postOutcome } from '../lib/api';
const browserTtsFallbackEnabled = import.meta.env.VITE_ENABLE_BROWSER_TTS_FALLBACK === 'true';
/** Speak a cue in the browser only when the explicit fallback flag is enabled. */
export function speakInBrowser(text: string): void {
if (!browserTtsFallbackEnabled) return;
if (typeof window === 'undefined' || !('speechSynthesis' in window) || !text) return;
const u = new SpeechSynthesisUtterance(text);
u.rate = 1.05;
window.speechSynthesis.cancel(); // drop any queued cue so the latest wins
window.speechSynthesis.speak(u);
}
/** Unlock speechSynthesis from a user gesture when the browser fallback is enabled. */
export function primeSpeech(): void {
if (!browserTtsFallbackEnabled) return;
if (typeof window === 'undefined' || !('speechSynthesis' in window)) return;
const u = new SpeechSynthesisUtterance(' ');
u.volume = 0;
window.speechSynthesis.speak(u);
}
export function useInterventions(room: Room | null) {
const [active, setActive] = useState<Intervention | null>(null);
const [hermes, setHermes] = useState<HermesMessage | null>(null);
@@ -18,9 +39,17 @@ export function useInterventions(room: Room | null) {
if (msg.type === 'COLLISION') {
setActive(msg.intervention);
setActionUrl(null);
// Clear the prior intervention's cue/message so a new card never shows a
// stale voice line. The fresh ones arrive right after on the same
// reliable channel (ordered: COLLISION -> HERMES_MESSAGE -> VOICE_CUE).
setVoiceCue(null);
setHermes(null);
}
if (msg.type === 'HERMES_MESSAGE') setHermes(msg.message);
if (msg.type === 'VOICE_CUE') setVoiceCue(msg.text);
if (msg.type === 'VOICE_CUE') {
setVoiceCue(msg.text);
speakInBrowser(msg.text);
}
};
room.on(RoomEvent.DataReceived, onData);
return () => {
@@ -42,7 +71,9 @@ export function useInterventions(room: Room | null) {
interventionId: active.id,
collisionId: active.collisionId,
podId: active.podId,
wasRealCollision: true,
// Placeholder only — the backend derives the authoritative value from
// git overlap at outcome time (the client cannot know). (RSI Step 3)
wasRealCollision: false,
accepted,
recordedAt: new Date().toISOString(),
});
@@ -53,6 +84,8 @@ export function useInterventions(room: Room | null) {
{ reliable: true, topic: DATA_TOPIC },
);
setActive(null);
setVoiceCue(null);
setHermes(null);
return status;
},
[active, room],
+11
View File
@@ -1,13 +1,24 @@
import { StrictMode } from 'react';
import { createRoot } from 'react-dom/client';
import { ClerkProvider } from '@clerk/react';
import { shadcn } from '@clerk/ui/themes';
import App from './App.js';
import './index.css';
import '@clerk/ui/themes/shadcn.css';
import { TooltipProvider } from '@/components/ui/tooltip';
const clerkPublishableKey = import.meta.env.VITE_CLERK_PUBLISHABLE_KEY;
if (!clerkPublishableKey) {
throw new Error('Missing VITE_CLERK_PUBLISHABLE_KEY');
}
createRoot(document.getElementById('root')!).render(
<StrictMode>
<ClerkProvider publishableKey={clerkPublishableKey} appearance={{ theme: shadcn }}>
<TooltipProvider>
<App />
</TooltipProvider>
</ClerkProvider>
</StrictMode>,
);
+5 -9
View File
@@ -1,18 +1,14 @@
import { defineConfig } from 'vite';
import react from '@vitejs/plugin-react';
import tailwindcss from '@tailwindcss/vite';
import { VitePWA } from 'vite-plugin-pwa';
import { fileURLToPath, URL } from 'node:url';
export default defineConfig({
plugins: [
react(),
tailwindcss(),
// Self-destroying during active development: unregisters any previously
// installed service worker and clears its caches so deploys are always
// fresh (no stale UI). Re-enable a precaching PWA before the demo.
VitePWA({ selfDestroying: true }),
],
// No service worker in production. We deploy continuously during the event and
// any SW (even vite-plugin-pwa's self-destroying one) forces open tabs to
// reload, which breaks the live demo. Existing SWs are torn down by the
// cleanup snippet in index.html. Re-add a precaching PWA only post-event.
plugins: [react(), tailwindcss()],
server: {
port: 5173,
},
+2
View File
@@ -49,6 +49,7 @@ services:
- { key: GEMINI_API_KEY, scope: RUN_TIME, type: SECRET }
- { key: GEMINI_VISION_MODEL, scope: RUN_TIME, value: gemini-2.0-flash }
- { key: GEMINI_LIVE_MODEL, scope: RUN_TIME, value: gemini-3.1-flash-tts-preview }
- { key: GEMINI_TTS_VOICE, scope: RUN_TIME, value: Charon }
- { key: GEMINI_EMBEDDING_MODEL, scope: RUN_TIME, value: gemini-embedding-001 }
- { key: GITHUB_TOKEN, scope: RUN_TIME, type: SECRET }
- { key: GITHUB_REPO, scope: RUN_TIME, value: karti-ai/podman }
@@ -75,6 +76,7 @@ workers:
- { key: GEMINI_API_KEY, scope: RUN_TIME, type: SECRET }
- { key: GEMINI_VISION_MODEL, scope: RUN_TIME, value: gemini-2.0-flash }
- { key: GEMINI_LIVE_MODEL, scope: RUN_TIME, value: gemini-3.1-flash-tts-preview }
- { key: GEMINI_TTS_VOICE, scope: RUN_TIME, value: Charon }
- { key: GEMINI_EMBEDDING_MODEL, scope: RUN_TIME, value: gemini-embedding-001 }
- { key: GITHUB_TOKEN, scope: RUN_TIME, type: SECRET }
- { key: GITHUB_REPO, scope: RUN_TIME, value: karti-ai/podman }
+4
View File
@@ -23,6 +23,10 @@ lk.165-22-129-249.sslip.io {
}
}
gemini.165-22-129-249.sslip.io {
reverse_proxy 127.0.0.1:3000
}
podman.live, www.podman.live {
route {
handle /api/* {
+2
View File
@@ -49,6 +49,7 @@ services:
- { key: GEMINI_API_KEY, scope: RUN_TIME, type: SECRET }
- { key: GEMINI_VISION_MODEL, scope: RUN_TIME, value: gemini-2.0-flash }
- { key: GEMINI_LIVE_MODEL, scope: RUN_TIME, value: gemini-3.1-flash-tts-preview }
- { key: GEMINI_TTS_VOICE, scope: RUN_TIME, value: Charon }
- { key: GEMINI_EMBEDDING_MODEL, scope: RUN_TIME, value: gemini-embedding-001 }
- { key: GITHUB_TOKEN, scope: RUN_TIME, type: SECRET }
- { key: GITHUB_REPO, scope: RUN_TIME, value: karti-ai/podman }
@@ -75,6 +76,7 @@ workers:
- { key: GEMINI_API_KEY, scope: RUN_TIME, type: SECRET }
- { key: GEMINI_VISION_MODEL, scope: RUN_TIME, value: gemini-2.0-flash }
- { key: GEMINI_LIVE_MODEL, scope: RUN_TIME, value: gemini-3.1-flash-tts-preview }
- { key: GEMINI_TTS_VOICE, scope: RUN_TIME, value: Charon }
- { key: GEMINI_EMBEDDING_MODEL, scope: RUN_TIME, value: gemini-embedding-001 }
- { key: GITHUB_TOKEN, scope: RUN_TIME, type: SECRET }
- { key: GITHUB_REPO, scope: RUN_TIME, value: karti-ai/podman }
@@ -0,0 +1,17 @@
[Unit]
Description=PodMan LiveKit Gemini starter agent
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
WorkingDirectory=/root/podman/examples/livekit-gemini-hacker-starter/agent
Environment=PATH=/root/.local/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
ExecStart=/root/.local/bin/uv run agent.py dev
Restart=always
RestartSec=3
KillSignal=SIGTERM
TimeoutStopSec=20
[Install]
WantedBy=multi-user.target
@@ -0,0 +1,18 @@
[Unit]
Description=PodMan LiveKit Gemini starter frontend
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
WorkingDirectory=/root/podman/examples/livekit-gemini-hacker-starter/frontend
Environment=NODE_ENV=production
Environment=NEXT_TELEMETRY_DISABLED=1
ExecStart=/usr/bin/pnpm start --hostname 127.0.0.1 --port 3000
Restart=always
RestartSec=3
KillSignal=SIGTERM
TimeoutStopSec=20
[Install]
WantedBy=multi-user.target
@@ -0,0 +1,20 @@
[Unit]
Description=PodMan private LiveKit/Gemini live conversation agent
After=network-online.target podman-platform-api.service
Wants=network-online.target
[Service]
Type=simple
WorkingDirectory=/root/podman/agents/podman-live-conversation
Environment=PATH=/root/.local/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
EnvironmentFile=/root/podman/backend/.env
ExecStart=/root/.local/bin/uv run agent.py dev
Restart=always
RestartSec=3
MemoryHigh=1024M
MemoryMax=1536M
KillSignal=SIGTERM
TimeoutStopSec=20
[Install]
WantedBy=multi-user.target
@@ -12,6 +12,9 @@ EnvironmentFile=/root/podman/backend/.env
ExecStart=/usr/bin/node dist/agent.js
Restart=always
RestartSec=3
MemoryHigh=1536M
MemoryMax=2G
OOMPolicy=stop
KillSignal=SIGTERM
TimeoutStopSec=20
+3
View File
@@ -22,9 +22,12 @@
"deploy:static:local": "node scripts/deploy-static-local.mjs",
"hermes:watchdog": "node scripts/hermes-watchdog.mjs",
"hermes:watchdog:strict": "node scripts/hermes-watchdog.mjs --strict",
"hermes:notify": "node scripts/hermes-notify.mjs",
"hermes:sync-deploy": "node scripts/hermes-sync-deploy.mjs",
"hermes:install": "node scripts/install-hermes-ops.mjs",
"healthcheck:public": "node scripts/healthcheck-public.mjs",
"livekit:conversation:agent": "cd agents/podman-live-conversation && uv run agent.py dev",
"livekit:conversation:test": "cd agents/podman-live-conversation && uv run pytest",
"verify": "pnpm lint && pnpm typecheck && pnpm build && pnpm verify:backend && pnpm verify:frontend",
"verify:full": "pnpm verify && pnpm verify:infra && pnpm build:container && pnpm verify:containers",
"verify:backend": "node scripts/verify-backend.mjs",
+3245 -20
View File
File diff suppressed because it is too large Load Diff
+17
View File
@@ -0,0 +1,17 @@
#!/bin/bash
REPO="/home/ramis/Programming/podman"
LOG="/home/ramis/Programming/podman/scripts/auto-pull.log"
cd "$REPO" || exit 1
# Stash any local changes, pull, pop
git fetch origin main 2>>"$LOG"
LOCAL=$(git rev-parse HEAD)
REMOTE=$(git rev-parse origin/main)
if [ "$LOCAL" != "$REMOTE" ]; then
echo "[$(date)] Pulling: $LOCAL -> $REMOTE" >> "$LOG"
git pull --ff-only origin main >> "$LOG" 2>&1
else
echo "[$(date)] Up to date" >> "$LOG"
fi
+12 -3
View File
@@ -240,6 +240,7 @@ async function checkGeminiVision() {
async function checkGeminiVoiceModel() {
const key = configuredGeminiKey();
const model = process.env.GEMINI_LIVE_MODEL ?? 'gemini-3.1-flash-tts-preview';
const voice = process.env.GEMINI_TTS_VOICE ?? 'Charon';
const res = await doFetch(
`https://generativelanguage.googleapis.com/v1beta/models?key=${encodeURIComponent(key.value)}`,
);
@@ -257,10 +258,18 @@ async function checkGeminiVoiceModel() {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({
contents: [{ parts: [{ text: 'Say clearly: PodMan voice check.' }] }],
contents: [
{
parts: [
{
text: 'Speak this as a calm engineering teammate. Say only: PodMan voice check.',
},
],
},
],
generationConfig: {
responseModalities: ['AUDIO'],
speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: 'Kore' } } },
speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: voice } } },
},
}),
},
@@ -269,7 +278,7 @@ async function checkGeminiVoiceModel() {
const ttsBody = await tts.json();
const audio = ttsBody.candidates?.[0]?.content?.parts?.[0]?.inlineData?.data;
if (!audio) throw new Error('Gemini voice response had no audio');
return `${model}, generated ${Buffer.from(audio, 'base64').byteLength} audio bytes`;
return `${model}/${voice}, generated ${Buffer.from(audio, 'base64').byteLength} audio bytes`;
}
async function checkGeminiEmbeddings() {
+54
View File
@@ -0,0 +1,54 @@
#!/usr/bin/env node
import { existsSync } from 'node:fs';
import { config as loadEnv } from 'dotenv';
const envPath = process.env.DOTENV_CONFIG_PATH ?? (existsSync('.env') ? '.env' : 'backend/.env');
loadEnv({ path: envPath, quiet: true });
function usage() {
console.error(
'Usage: node scripts/hermes-notify.mjs --pod <podId> --message <text> [--engineers alice,bob] [--file path] [--urgent]',
);
process.exit(1);
}
function arg(name) {
const index = process.argv.indexOf(name);
return index === -1 ? '' : (process.argv[index + 1] ?? '');
}
const podId = arg('--pod');
const message = arg('--message');
if (!podId || !message) usage();
const apiBase = (
process.env.PODMAN_API_URL ??
process.env.BACKEND_URL ??
`http://127.0.0.1:${process.env.PORT ?? '8787'}`
).replace(/\/$/, '');
const engineers = arg('--engineers')
.split(',')
.map((name) => name.trim())
.filter(Boolean);
const file = arg('--file');
const urgent = process.argv.includes('--urgent');
const res = await globalThis.fetch(`${apiBase}/api/pods/${encodeURIComponent(podId)}/hermes/notify`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({
message,
...(engineers.length ? { engineers } : {}),
...(file ? { file } : {}),
...(urgent ? { urgency: 'urgent' } : {}),
}),
});
const text = await res.text();
if (!res.ok) {
console.error(text);
process.exit(1);
}
console.log(text);
+103
View File
@@ -22,6 +22,7 @@ const env = {
GITHUB_TOKEN: process.env.GITHUB_TOKEN ?? 'verify-github',
GITHUB_REPO: process.env.GITHUB_REPO ?? 'karti-ai/podman',
MONGODB_URI: mongoUri,
INTERNAL_AGENT_TOKEN: process.env.INTERNAL_AGENT_TOKEN ?? 'verify-internal-agent-token',
};
function fail(message) {
@@ -95,6 +96,105 @@ async function verifyApi() {
);
if (!withMember.members.includes('Hermes')) fail('member add did not persist');
const hermesNotify = await json(
await doFetch(`${baseUrl}/api/pods/${encodeURIComponent(created.id)}/hermes/notify`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({
message: 'Hermes verification notification.',
engineers: ['Alice', 'Bob'],
file: 'src/verify-hermes.ts',
dryRun: true,
}),
}),
);
if (
hermesNotify.livekit !== 'dry-run' ||
hermesNotify.intervention?.message !== 'Hermes verification notification.'
) {
fail('Hermes notify endpoint returned unexpected payload');
}
const liveConversation = await json(
await doFetch(`${baseUrl}/api/pods/${encodeURIComponent(created.id)}/live-conversation/start`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ identity: 'Alice', displayName: 'Alice' }),
}),
);
if (
typeof liveConversation.token !== 'string' ||
liveConversation.token.split('.').length !== 3 ||
typeof liveConversation.room !== 'string' ||
!liveConversation.room.includes('podman-live:')
) {
fail('live conversation start did not return a private room JWT');
}
const liveStatus = await json(
await doFetch(
`${baseUrl}/api/pods/${encodeURIComponent(
created.id,
)}/live-conversation/status?identity=Alice`,
),
);
if (liveStatus.active?.sessionId !== liveConversation.sessionId) {
fail('live conversation status did not return the active session');
}
await json(
await doFetch(
`${baseUrl}/api/pods/${encodeURIComponent(
created.id,
)}/live-conversation/${encodeURIComponent(liveConversation.sessionId)}/stop`,
{ method: 'POST' },
),
);
const hermesJob = await json(
await doFetch(`${baseUrl}/api/internal/hermes/jobs`, {
method: 'POST',
headers: {
'content-type': 'application/json',
authorization: `Bearer ${env.INTERNAL_AGENT_TOKEN}`,
},
body: JSON.stringify({
prompt: 'Check repository state for backend verification.',
contextScope: 'current_repo',
riskLevel: 'read_only',
successCriteria: ['Git status is inspected.'],
podId: created.id,
identity: 'Alice',
sessionId: liveConversation.sessionId,
}),
}),
);
if (!hermesJob.id || hermesJob.status !== 'queued') {
fail('Hermes job create returned unexpected payload');
}
let finalJob = hermesJob;
for (let i = 0; i < 30; i++) {
finalJob = await json(
await doFetch(`${baseUrl}/api/internal/hermes/jobs/${encodeURIComponent(hermesJob.id)}`, {
headers: { authorization: `Bearer ${env.INTERNAL_AGENT_TOKEN}` },
}),
);
if (['completed', 'failed', 'aborted'].includes(finalJob.status)) break;
await delay(500);
}
if (finalJob.status !== 'completed') {
fail(`Hermes job did not complete: ${JSON.stringify(finalJob)}`);
}
const hermesEvents = await json(
await doFetch(
`${baseUrl}/api/internal/hermes/jobs/${encodeURIComponent(hermesJob.id)}/events`,
{
headers: { authorization: `Bearer ${env.INTERNAL_AGENT_TOKEN}` },
},
),
);
if (!Array.isArray(hermesEvents) || hermesEvents.length < 1) {
fail('Hermes job events were not persisted');
}
await json(
await doFetch(`${baseUrl}/api/pods/${encodeURIComponent(created.id)}`, { method: 'DELETE' }),
);
@@ -262,6 +362,9 @@ try {
'health',
'token',
'pod-crud',
'hermes-notify',
'live-conversation-session',
'hermes-job-lifecycle',
'collision',
'memory-recall',
'graph',
+1
View File
@@ -24,6 +24,7 @@ const containerEnv = {
GEMINI_API_KEY: process.env.GEMINI_API_KEY ?? 'verify-gemini',
GEMINI_VISION_MODEL: process.env.GEMINI_VISION_MODEL ?? 'gemini-2.0-flash',
GEMINI_LIVE_MODEL: process.env.GEMINI_LIVE_MODEL ?? 'gemini-3.1-flash-tts-preview',
GEMINI_TTS_VOICE: process.env.GEMINI_TTS_VOICE ?? 'Charon',
GEMINI_EMBEDDING_MODEL: process.env.GEMINI_EMBEDDING_MODEL ?? 'gemini-embedding-001',
GITHUB_TOKEN: process.env.GITHUB_TOKEN ?? 'verify-github',
GITHUB_REPO: process.env.GITHUB_REPO ?? 'karti-ai/podman',
+170 -17
View File
@@ -21,7 +21,17 @@ const { DATA_TOPIC } = await import('../shared/dist/messages.js').catch(() => ({
DATA_TOPIC: 'podman.intervention',
}));
const backendRequire = createRequire(new URL('../backend/package.json', import.meta.url));
const { Room } = backendRequire('@livekit/rtc-node');
const {
AudioFrame,
AudioSource,
LocalAudioTrack,
Room,
TrackPublishOptions,
TrackSource,
} = backendRequire('@livekit/rtc-node');
const pods = await fetchJson('/api/pods');
const verifyPod = pods.find((pod) => pod.id === 'frontend-pod') ?? pods[0];
if (!verifyPod) throw new Error('no pods available for frontend verification');
async function stopChild(child) {
if (!child || child.exitCode !== null || child.signalCode !== null) return;
@@ -89,7 +99,7 @@ async function publishIntervention(room, podId) {
podId,
kind: 'card',
message: 'Verification collision: two engineers are editing frontend/src/App.tsx.',
suggestedAction: { kind: 'sync_before_push' },
suggestedAction: { kind: 'open_sync_pr' },
status: 'pending',
createdAt: now,
};
@@ -120,6 +130,27 @@ async function publishDataMessage(room, message) {
});
}
async function publishAudioProbe(room) {
const source = new AudioSource(24_000, 1, 5_000);
const track = LocalAudioTrack.createAudioTrack(`verify-audio-${process.pid}`, source);
const options = new TrackPublishOptions();
options.source = TrackSource.SOURCE_MICROPHONE;
const publication = await room.localParticipant.publishTrack(track, options);
await source.captureFrame(new AudioFrame(new Int16Array(24_000), 24_000, 1, 24_000));
return { source, publication };
}
async function waitForAttachedAudio(page) {
const audioSink = page.getByTestId('livekit-audio-sink');
await audioSink.waitFor({ timeout: 15_000 });
for (let i = 0; i < 30; i++) {
const count = await audioSink.locator('audio').count();
if (count > 0) return;
await delay(250);
}
throw new Error('LiveKit audio track was not attached to the hidden audio sink');
}
async function waitForInterventionCard(page, room, podId) {
const cardText = 'Verification collision: two engineers are editing frontend/src/App.tsx.';
for (let attempt = 1; attempt <= 3; attempt++) {
@@ -168,6 +199,27 @@ async function waitForPublishedScreenShare(roomName) {
);
}
async function boxOf(locator, name) {
const box = await locator.boundingBox();
if (!box) throw new Error(`${name} did not have a visible bounding box`);
return box;
}
function assertNoOverlap(left, main, right, label) {
if (left.x + left.width > main.x + 2) {
throw new Error(`${label}: left sidebar overlaps main workspace`);
}
if (main.x + main.width > right.x + 2) {
throw new Error(`${label}: right sidebar overlaps main workspace`);
}
}
function assertContained(container, child, label) {
if (child.x < container.x - 2 || child.x + child.width > container.x + container.width + 2) {
throw new Error(`${label}: body summary is not contained inside the main workspace`);
}
}
let preview = null;
if (shouldStartPreview) {
preview = spawn(
@@ -227,7 +279,9 @@ page.on('console', (msg) => {
});
page.on('pageerror', (err) => pageErrors.push(err.message));
page.on('requestfailed', (req) => {
failedRequests.push(`${req.url()} ${req.failure()?.errorText ?? ''}`.trim());
const failure = req.failure()?.errorText ?? '';
if (req.url().includes('/activity/stream') && failure.includes('ERR_ABORTED')) return;
failedRequests.push(`${req.url()} ${failure}`.trim());
});
try {
@@ -235,29 +289,113 @@ try {
await page.waitForTimeout(500);
const bodyText = await page.locator('body').innerText();
const hasPodCards =
(await page.locator('text=/Frontend Pod|Backend Pod|graph pod/i').count()) > 0;
const hasPodCards = (await page.getByText(verifyPod.name, { exact: true }).count()) > 0;
const hasOverlay = (await page.locator('vite-error-overlay, .vite-error-overlay').count()) > 0;
if (bodyText.length < 100) throw new Error('frontend rendered too little text');
if (!hasPodCards) throw new Error('pod cards did not render');
if (hasOverlay) throw new Error('Vite error overlay is visible');
await page.getByRole('button', { name: 'Team memory' }).click();
await page.getByRole('button', { name: 'Pod actions' }).first().click();
await page.getByRole('menuitem', { name: 'Team memory' }).click();
await page.getByText('Workflow metrics').waitFor({ timeout: 15_000 });
await page.getByText('Learning edges').waitFor({ timeout: 15_000 });
await page.getByRole('img', { name: 'PodMan team-memory graph' }).waitFor({ timeout: 15_000 });
await page.getByRole('button', { name: 'engineer: Karti' }).click();
await page.getByText('Learned owner of auth; backend + DB wiring.').waitFor({ timeout: 15_000 });
await page
.getByRole('button', { name: /engineer:/ })
.first()
.click();
await page.getByText('Relationships').waitFor({ timeout: 15_000 });
await page.getByRole('button', { name: 'Whole graph' }).click();
await page.getByRole('button', { name: /Pods/i }).click();
const frontendPodCard = page
.getByText('Frontend Pod', { exact: true })
const podCard = page
.getByText(verifyPod.name, { exact: true })
.locator('xpath=ancestor::*[.//input[@placeholder="Your name"]][1]');
await frontendPodCard.getByPlaceholder('Your name').fill(verifyMember);
await frontendPodCard.getByRole('button', { name: 'Add and join' }).click();
await podCard.getByPlaceholder('Your name').first().fill(verifyMember);
await podCard.getByRole('button', { name: 'Join' }).first().click();
await page.getByRole('button', { name: 'Share screen' }).waitFor({ timeout: 15_000 });
await page.getByTestId('live-conversation-toggle').waitFor({ timeout: 15_000 });
await page.getByRole('button', { name: 'Start Live Conversation' }).waitFor({
timeout: 15_000,
});
if (new URL(page.url()).pathname !== `/${verifyPod.id}`) {
throw new Error(`join did not update URL to /${verifyPod.id}: ${page.url()}`);
}
await page.getByRole('heading', { name: 'My stream' }).waitFor({ timeout: 15_000 });
await page.getByRole('heading', { name: 'Team stream' }).waitFor({ timeout: 15_000 });
const bodySummary = page.getByTestId('pod-body-summary');
const mainWorkspace = page.getByTestId('pod-main-workspace');
const mySidebar = page.getByTestId('my-stream-sidebar');
const teamSidebar = page.getByTestId('team-stream-sidebar');
await bodySummary.waitFor({ timeout: 15_000 });
await mainWorkspace.waitFor({ timeout: 15_000 });
await mySidebar.waitFor({ timeout: 15_000 });
await teamSidebar.waitFor({ timeout: 15_000 });
if ((await page.getByTestId('pod-topbar').count()) > 0) {
throw new Error('pod detail view still renders a topbar test id');
}
if ((await bodySummary.getByRole('button', { name: /stream|team/i }).count()) > 0) {
throw new Error('pod body summary contains sidebar stream/team controls');
}
const expandedLayout = {
summary: await boxOf(bodySummary, 'body summary expanded'),
main: await boxOf(mainWorkspace, 'main workspace expanded'),
left: await boxOf(mySidebar, 'my stream sidebar expanded'),
right: await boxOf(teamSidebar, 'team stream sidebar expanded'),
};
assertNoOverlap(
expandedLayout.left,
expandedLayout.main,
expandedLayout.right,
'expanded layout',
);
assertContained(expandedLayout.main, expandedLayout.summary, 'expanded layout');
await page.locator('[data-testid="my-stream-toggle"]:visible').click();
await page.waitForTimeout(300);
const leftCollapsedLayout = {
main: await boxOf(mainWorkspace, 'main workspace after left collapse'),
left: await boxOf(mySidebar, 'my stream sidebar collapsed'),
right: await boxOf(teamSidebar, 'team stream sidebar with left collapsed'),
summary: await boxOf(bodySummary, 'body summary after left collapse'),
};
if (leftCollapsedLayout.left.width >= expandedLayout.left.width - 24) {
throw new Error('my stream sidebar did not collapse into a compact rail');
}
if (leftCollapsedLayout.main.width <= expandedLayout.main.width) {
throw new Error('main workspace did not expand after my stream collapsed');
}
assertNoOverlap(
leftCollapsedLayout.left,
leftCollapsedLayout.main,
leftCollapsedLayout.right,
'left collapsed layout',
);
assertContained(leftCollapsedLayout.main, leftCollapsedLayout.summary, 'left collapsed layout');
await page.locator('[data-testid="team-stream-toggle"]:visible').click();
await page.waitForTimeout(300);
const bothCollapsedLayout = {
main: await boxOf(mainWorkspace, 'main workspace after both collapse'),
left: await boxOf(mySidebar, 'my stream sidebar with both collapsed'),
right: await boxOf(teamSidebar, 'team stream sidebar collapsed'),
summary: await boxOf(bodySummary, 'body summary after both collapse'),
};
if (bothCollapsedLayout.right.width >= expandedLayout.right.width - 24) {
throw new Error('team stream sidebar did not collapse into a compact rail');
}
if (bothCollapsedLayout.main.width <= leftCollapsedLayout.main.width) {
throw new Error('main workspace did not expand after team stream collapsed');
}
assertNoOverlap(
bothCollapsedLayout.left,
bothCollapsedLayout.main,
bothCollapsedLayout.right,
'both collapsed layout',
);
assertContained(bothCollapsedLayout.main, bothCollapsedLayout.summary, 'both collapsed layout');
await page.locator('[data-testid="my-stream-toggle"]:visible').click();
await page.locator('[data-testid="team-stream-toggle"]:visible').click();
await page.waitForTimeout(300);
const joinedText = await page.locator('body').innerText();
const hasPodView =
@@ -272,18 +410,30 @@ try {
await page.getByRole('button', { name: 'Share screen' }).click();
await page.getByRole('button', { name: 'Stop sharing' }).waitFor({ timeout: 15_000 });
await page.getByText(/Screen\s*published/i).waitFor({ timeout: 15_000 });
await waitForPublishedScreenShare('frontend-pod');
await waitForPublishedScreenShare(verifyPod.id);
await page.getByRole('button', { name: 'Stop sharing' }).click();
await page.getByRole('button', { name: 'Share screen' }).waitFor({ timeout: 15_000 });
const publisher = await connectPublisher('frontend-pod');
const publisher = await connectPublisher(verifyPod.id);
try {
const intervention = await waitForInterventionCard(page, publisher, 'frontend-pod');
const audioProbe = await publishAudioProbe(publisher);
try {
await waitForAttachedAudio(page);
} finally {
if (audioProbe.publication.sid) {
await publisher.localParticipant
.unpublishTrack(audioProbe.publication.sid, true)
.catch(() => {});
}
await audioProbe.source.close().catch(() => {});
}
const intervention = await waitForInterventionCard(page, publisher, verifyPod.id);
await publishDataMessage(publisher, {
type: 'HERMES_MESSAGE',
message: {
id: `hermes-${process.pid}`,
podId: 'frontend-pod',
podId: verifyPod.id,
interventionId: intervention.id,
recipients: ['Verify'],
text: 'Hermes verification message routed to the team.',
@@ -310,6 +460,7 @@ try {
}
await page.getByRole('button', { name: 'Leave pod' }).click();
await page.waitForURL((url) => url.pathname === '/', { timeout: 15_000 });
if (consoleErrors.length) throw new Error(`console errors: ${consoleErrors.join(' | ')}`);
if (pageErrors.length) throw new Error(`page errors: ${pageErrors.join(' | ')}`);
@@ -325,7 +476,9 @@ try {
graph: true,
joined: true,
screenShare: 'livekit-published',
audioSink: 'livekit-attached',
intervention: 'collision-hermes-voice',
podId: verifyPod.id,
member: verifyMember,
},
null,
@@ -333,7 +486,7 @@ try {
),
);
} finally {
await doFetch(`${apiBase}/api/pods/frontend-pod/members/${encodeURIComponent(verifyMember)}`, {
await doFetch(`${apiBase}/api/pods/${verifyPod.id}/members/${encodeURIComponent(verifyMember)}`, {
method: 'DELETE',
}).catch(() => {});
await browser.close();

Some files were not shown because too many files have changed in this diff Show More