From 28f3236d76a4f3ad18cca066a8479f1d8c762824 Mon Sep 17 00:00:00 2001 From: Ramis Date: Sun, 28 Jun 2026 11:28:13 -0700 Subject: [PATCH] docs(demo): tighten 1-min script to a scannable half-page --- README.md | 89 +++++++++++++++ docs/demo.md | 196 ++++++++++++++------------------ frontend/public/favicon.svg | 69 ++++++++++- frontend/public/podman-logo.svg | 64 +++++++++++ frontend/src/App.tsx | 28 ++--- 5 files changed, 319 insertions(+), 127 deletions(-) create mode 100644 frontend/public/podman-logo.svg diff --git a/README.md b/README.md index 5d969a8..271bfcf 100644 --- a/README.md +++ b/README.md @@ -7,6 +7,37 @@ each engineer is doing, remembers which interventions helped, and nudges the team before duplicated work, merge collisions, or missed handoffs slow everyone down. +

+ + Try PodMan live in production + +

+ +

+ Open the live production app: https://podman.live/ +

+ +

+ PodMan is already deployed. Click the link above and try the production app. +

+ +

+ Watch the demo video +

+ +

+ + PodMan demo video thumbnail + +

+ [LiveKit](https://livekit.io/) · [MongoDB](https://www.mongodb.com/) · [Gemini](https://ai.google.dev/) · [Hermes](https://hermes-agent.nousresearch.com/) · [Modular MAX](https://www.modular.com/max) · @@ -38,6 +69,64 @@ Source image: `~/pic.jpg`, committed as | Start local development | API, vision agent, and frontend commands are in [Local Development](#local-development) | | Debug production | Public checks first, then systemd services in [Production Operations](#production-operations) | +## Hackathon Fit + +PodMan is built for the **2026 AI Engineer World's Fair Hackathon** theme: +**Continual Learning**. + +It is not a wrapper chatbot or static dashboard. It is a production-running +agentic coordination system that gets better from real use: every observation, +collision, intervention, accept, dismiss, suppression, and conversation note +feeds MongoDB-backed memory so future interventions become sharper and less +annoying. + +| Hackathon target | How PodMan addresses it | +| ---------------------- | ---------------------------------------------------------------------------------------- | +| Continual Learning | Learns from real team behavior and accepted/dismissed interventions. | +| Self-Improvement Stack | Uses memory stats, Hermes jobs, production watchdogs, and deployment checks. | +| Recursive Intelligence | Gives agents operational memory and feedback loops for improving future coordination. | +| Live Demo | Already deployed at [`https://podman.live/`](https://podman.live/). | +| Technicality | Combines realtime media, vision, vector recall, graph traversal, memory, and ops agents. | +| Creativity | Solves the coordination cost around AI-assisted engineering teams. | + +### Sponsor Tech Used + +| Sponsor / tech | Usage in PodMan | +| -------------- | -------------------------------------------------------------------------------- | +| DigitalOcean | Hosts the production app, API, workers, Caddy, and systemd-supervised services. | +| LiveKit | Realtime room layer for screen share, audio, data messages, and agent presence. | +| Gemini | Screen understanding, live room conversation, urgent TTS, embeddings, and music. | +| MongoDB Atlas | Durable memory, `$vectorSearch`, `$graphLookup`, user learning, and job logs. | +| Modular MAX | Serves `gemma-4-31B-it` as the long-context reasoning model. | + +### What To Try + +```mermaid +flowchart LR + A["Open podman.live"] --> B["Join a pod"] + B --> C["Share screen via LiveKit"] + C --> D["Gemini extracts work context"] + D --> E["MongoDB recalls prior outcomes"] + E --> F["PodMan nudges only when useful"] + F --> G["Accept or dismiss"] + G --> H["Memory improves next run"] + + classDef action fill:#e8f1ff,stroke:#3366cc,color:#0b1f44; + classDef ai fill:#f4edff,stroke:#805ad5,color:#2d1857; + classDef memory fill:#fff6df,stroke:#c47f00,color:#3d2b00; + class A,B,C,F,G action; + class D ai; + class E,H memory; +``` + +1. Open [`https://podman.live/`](https://podman.live/). +2. Create or join a pod. +3. Add teammates and share screens. +4. Watch PodMan build live context, detect overlap, and surface intervention + cards. +5. Accept or dismiss a card; that feedback becomes memory for the next similar + event. + --- ## Current Live System diff --git a/docs/demo.md b/docs/demo.md index aa35f4b..e331414 100644 --- a/docs/demo.md +++ b/docs/demo.md @@ -5,7 +5,7 @@ **4-minute script** for the complete story. Practice the 4-min to land at 3:45. **The thesis (say this first, in either script):** AI is collapsing the cost of -_writing_ code. More and more of every codebase is authored with AI assist — and +*writing* code. More and more of every codebase is authored with AI assist — and increasingly by **autonomous agents**. So the bottleneck of software engineering is shifting away from engineering itself toward **organization, management, and project coordination**. Pair a team with hundreds of agents all committing to the @@ -13,8 +13,8 @@ same repo at once and it becomes **physically impossible for a human to track progress or avoid stepping on someone else's work.** Code generation scaled; human coordination did not. That gap is the new bottleneck. -**The one-line story:** writing code isn't the bottleneck anymore — _coordinating -who (and what) is writing what_ is. PodMan is a pair programmer for the whole +**The one-line story:** writing code isn't the bottleneck anymore — *coordinating +who (and what) is writing what* is. PodMan is a pair programmer for the whole team — humans **and** agents: it watches every actor's work in real time, gives everyone live status without anyone having to interrupt anyone, catches collisions before they land, and learns your team's dynamics so it nudges less and helps more @@ -28,85 +28,52 @@ every actor, every day, is the value. --- + + ## 1-Minute Live Script (you + teammates — real-time interventions) -**Goal:** in 60 seconds, prove the unprecedented loop — real-time multimodal -perception → live collision intervention spoken by an **on-device Gemma model** -→ a **self-improving** coordination layer. Hits four judging tracks: Continual -Learning, the Self-Improvement Stack, Recursive Intelligence, and on-device -Gemma. **Hard limit:** 1:00. +**Hard limit 1:00.** Hits the tracks: real-time multimodal, on-device Gemma, +Continual Learning, Self-Improvement, Recursive Intelligence. -**Pre-flight (do NOT skip — these are the win conditions):** +**Pre-flight:** pod open on `podman.live`, all joined, screen share + **sound on**; +git watchers running; everyone has `README.md` open. On-device **Gemma** must be +up as Hermes's voice (else Gemini TTS auto-falls-back — just drop the "on-device" +word). -- **On-device Gemma must be committed and running** as Hermes's voice/reasoning - layer before showtime. Verify it speaks once in rehearsal. If it isn't up, - PodMan **automatically falls back to Gemini TTS** — the demo still works, you - just drop the "on-device" line. Never debug Gemma on stage. -- Pod open on the shared screen (`podman.live`), all teammates joined, screen - share on, **sound on** — voice is live for _every_ intervention now. -- Each teammate's git watcher running so local unpushed edits report in. -- Everyone has `README.md` open, ready to type. Pre-pick the first colliding - pair (you + a teammate). +**0:00–0:12 — Thesis** +> "Code is almost free to write now — humans with AI, soon swarms of agents, all +> on one repo. No human can track that. The bottleneck is coordination — and it +> has to fix itself, live. Watch." -### 0:00–0:12 — The thesis (one breath) +**0:12–0:35 — Collision caught live (money shot)** +- You + a teammate both edit `README.md`, save (unpushed). +- Card appears on every screen + voice fires: _" and are both + editing README.md. Sync before pushing."_ +- Say: "Gemini Vision reads our screens, fused with local git, in real time — + and the alert is **spoken by a Gemma model on-device**. Neither of us asked." -> "Code is almost free to write now — humans with AI today, swarms of autonomous -> agents tomorrow, all committing to one repo. No human can track that. The -> bottleneck stopped being engineering and became coordination — and it has to -> fix itself, on its own, in real time. Watch." +**0:35–0:52 — Self-improving (narrate)** +- Third teammate edits too → **new** card + voice for the new pair (never goes + silent). +- Say: "Every trace lands in **Atlas**; vector search recalls past collisions, + repeats come back tagged **'Seen before'** and escalate — no retraining. It + tunes its own coordination from its own outcomes." -_On screen:_ pod view, teammate tiles live, activity stream moving. +**0:52–1:00 — Close** +> "Real-time awareness, on-device voice, a coordination layer that improves +> itself — for humans and agents, zero interruptions." -### 0:12–0:32 — Real-time perception → live intervention (the money shot) - -- On cue, **you and one teammate both edit `README.md`** and save (unpushed). -- A **conflict card appears on everyone's screen** within a cycle: _"Conflict: - + both on README.md (unpushed)."_ -- The voice fires out loud over LiveKit: _" and are both editing - README.md. Please sync before pushing."_ -- Land both the perception and the on-device beat: "PodMan is reading our actual - screens with **Gemini Vision** and fusing local git truth in real time — then - that alert is reasoned and **spoken by a Gemma model running on-device**, no - round-trip to a frontier API. Real-time multimodal in, on-device voice out. - Neither of us asked. Neither of us was watching." - -### 0:32–0:48 — Self-improving coordination (narrate the stack) - -- Have a **third teammate** edit the same file (or a second one). A **new** card - + voice fires for the new pair — _it never goes silent on later collisions._ -- Narrate the self-improvement architecture (no trigger needed — it's running): - "Every intervention and every accept/dismiss is a trace in **MongoDB Atlas**. - PodMan recalls similar past collisions with **vector search**, an - outcome-conditioned policy adapts who to nudge and how, and repeats come back - tagged **'Seen before'** and escalate automatically. That's the - self-improvement stack — it gets sharper the more the team uses it, with no - retraining and almost no human input. The recursive part: it's tuning its own - coordination behavior from its own logged outcomes." - -### 0:48–1:00 — Close - -> "Real-time multimodal awareness, on-device voice, and a coordination layer that -> improves itself from every interaction — for a whole team of humans and agents, -> with zero interruptions. That's the layer code generation never had." - -**Track callouts (say at least the bolded word in each):** - -| Track | Spoken moment | -| ----------------------- | -------------------------------------------------------- | -| Real-time multimodal | "reading our screens with **Gemini Vision** … real time" | -| **On-device Gemma** | "spoken by a **Gemma** model running **on-device**" | -| Continual Learning | "recalls past collisions … **'Seen before'** … no retraining" | -| Self-Improvement Stack | "every trace in **Atlas** … outcome-conditioned policy" | -| Recursive Intelligence | "tuning its **own** coordination behavior from its outcomes" | - -_Fallbacks:_ Gemma down → drop the on-device line, Gemini TTS covers it. Voice -silent → read the line aloud, point at the card (cards are the default path). -Card won't trigger → switch the colliding pair to a fresh file and re-save. +_Fallback:_ Gemma down → drop the on-device line. Voice silent → read it, point +at the card. No card → switch the pair to a fresh file and re-save. --- + + ## The script (4:00) + + ### 0:00–0:30 — The problem + hook > "AI made writing code almost free — humans with AI today, swarms of autonomous @@ -119,48 +86,48 @@ Card won't trigger → switch the colliding pair to a fresh file and re-save. > interrupting them, catches collisions before they land, and learns your team as > it goes." -_On screen:_ the pod view, two teammates joined, screen-share tiles live. +*On screen:* the pod view, two teammates joined, screen-share tiles live. ### 0:30–1:05 — Real-time team awareness (LiveKit + Gemini Vision) - Point at the two live screen tiles. "These are real screen shares over - **LiveKit**. Our agent subscribes to the tracks and samples frames." +**LiveKit**. Our agent subscribes to the tracks and samples frames." - "Each frame goes to **Gemini Vision**, which returns structured context — file, - symbol, activity — not a chatbot, a perception layer." +symbol, activity — not a chatbot, a perception layer." - Show the live activity stream filling in (Signals vs Reasoning sections). - Land the value: "This is the part that replaces 'what are you working on?' — - every teammate's current work is just _visible_, in real time. Nobody had to - ask." +every teammate's current work is just *visible*, in real time. Nobody had to +ask." -_Built-by-us callout:_ `backend/src/vision/gemini.ts`, the LiveKit agent worker. +*Built-by-us callout:* `backend/src/vision/gemini.ts`, the LiveKit agent worker. ### 1:05–1:40 — The catch (detection + first intervention) - Have alice and bob both edit the **same file** with unpushed changes. - "Normally nobody notices until merge time. GitHub can't see this — nothing's - pushed. Our detector fuses live screen context with **local git truth** from a - watcher on each laptop." -- A collision card appears: _"alice + bob both on detector.ts (unpushed)."_ -- Let the **Gemini TTS** urgent voice fire once over LiveKit: _"alice and bob are - both editing detector.ts. Please sync before pushing."_ +pushed. Our detector fuses live screen context with **local git truth** from a +watcher on each laptop." +- A collision card appears: *"alice + bob both on detector.ts (unpushed)."* +- Let the **Gemini TTS** urgent voice fire once over LiveKit: *"alice and bob are +both editing detector.ts. Please sync before pushing."* - Land the value: "That's a merge conflict and a wasted afternoon caught before it - happened — and neither of them had to be tracking the other." +happened — and neither of them had to be tracking the other." -_Built-by-us callout:_ `collision/detector.ts`, `action/hermes.ts`, +*Built-by-us callout:* `collision/detector.ts`, `action/hermes.ts`, `voice/live.ts`. ### 1:40–2:10 — Cross-channel overlap (research + code) - Keep alice editing `livekit.py`. - Have bob share a browser tab on LiveKit docs/SDK pages. -- A collaboration nudge appears: _"🤝 bob is researching LiveKit agents - (docs.livekit.io) while alice edits livekit.py — sync up before duplicating - effort."_ +- A collaboration nudge appears: *"🤝 bob is researching LiveKit agents +(docs.livekit.io) while alice edits livekit.py — sync up before duplicating +effort."* - Land the value: "This is not a merge conflict. PodMan caught duplicated effort - across channels — code on one screen, research on another — and nudged the team - before two people solved the same problem twice." +across channels — code on one screen, research on another — and nudged the team +before two people solved the same problem twice." -_Built-by-us callout:_ `vision/gemini.ts`, `collision/research.ts`, +*Built-by-us callout:* `vision/gemini.ts`, `collision/research.ts`, `memory/vectors.ts`. ### 2:10–2:50 — Continual learning (the theme — the money shot) @@ -168,53 +135,58 @@ _Built-by-us callout:_ `vision/gemini.ts`, `collision/research.ts`, This is the differentiator. Two beats, both from pre-seeded memory: 1. **It learned to stay quiet.** Trigger a pattern that was dismissed as a false - alarm earlier. "Last session a teammate marked this kind of alert as not a + alarm earlier. "Last session a teammate marked this kind of alert as not a real conflict. Watch — PodMan stays silent. No nagging." (No card fires.) 2. **It learned to escalate.** Trigger the real-conflict pattern that was - accepted before. The card now says **"Seen before."** and goes straight to + accepted before. The card now says **"Seen before."** and goes straight to the spoken urgent cue. - "The only input was one accept/dismiss tap. No retraining, no labeling. This is - **MongoDB Atlas vector search** recalling similar past events plus a policy - that adapts on the recalled outcome." +**MongoDB Atlas vector search** recalling similar past events plus a policy +that adapts on the recalled outcome." - Optional: show `/api/memory/stats` counts climbing — accumulated experience. -_Built-by-us callout:_ `memory/vectors.ts` ($vectorSearch), `memory/policy.ts` +*Built-by-us callout:* `memory/vectors.ts` ($vectorSearch), `memory/policy.ts` (outcome-conditioned gate), `memory/store.ts`. ### 2:50–3:30 — The five-minute meeting, killed (Gemini Live API) - Frame it: "Instead of breaking a teammate's focus to ask what they're up to, - you ask PodMan." -- Open the live voice conversation. Ask out loud: _"PodMan, what is everyone - working on, and where is the collision detector implemented?"_ +you ask PodMan." +- Open the live voice conversation. Ask out loud: *"PodMan, what is everyone +working on, and where is the collision detector implemented?"* - It answers with **real tool calls** — `search_repo`, git history, current - collisions — not guesses. +collisions — not guesses. - "This is the **Gemini Live API**, streaming speech-to-speech over LiveKit, with - custom function tools we wrote so it grounds every answer in the actual repo - and live state. That's the status sync, answered in seconds, with zero recovery - tax on anyone else." +custom function tools we wrote so it grounds every answer in the actual repo +and live state. That's the status sync, answered in seconds, with zero recovery +tax on anyone else." -_Built-by-us callout:_ `agents/podman-live-conversation/agent.py`. +*Built-by-us callout:* `agents/podman-live-conversation/agent.py`. ### 3:30–3:50 — Stack + close - "All on **DigitalOcean** — static frontend, API, and agent workers, supervised - by systemd. The ambient score is **Gemini Lyria** generated per pod through the - Interactions API." +by systemd. The ambient score is **Gemini Lyria** generated per pod through the +Interactions API." - Close: "Engineering ability stopped being the bottleneck — coordination is, and - it only gets worse as agents start writing alongside us. PodMan gives a whole - team, humans and agents, real-time awareness without the interruptions, catches - collisions before they cost an afternoon, and learns each team's dynamics so it - helps more over time. Saved focus, multiplied across every actor. That's - continual learning, shipped." +it only gets worse as agents start writing alongside us. PodMan gives a whole +team, humans and agents, real-time awareness without the interruptions, catches +collisions before they cost an afternoon, and learns each team's dynamics so it +helps more over time. Saved focus, multiplied across every actor. That's +continual learning, shipped." + + ### 3:50–4:00 — Buffer / Q&A handoff --- + + ## Sponsor-prize coverage (say each at least once) + | Prize | Spoken moment | Segment | | ---------------- | -------------------------------------------------------------------------------- | ---------------------- | | **Gemini** | Vision perception, Live API agent w/ tools, TTS voice, Lyria score | 0:30, 1:05, 2:50, 3:30 | @@ -222,10 +194,14 @@ _Built-by-us callout:_ `agents/podman-live-conversation/agent.py`. | **MongoDB** | "Atlas vector search recalling past events" | 1:50 | | **DigitalOcean** | "all on DigitalOcean, systemd-supervised workers" | 3:30 | + --- + + ## If something breaks (live recovery) + | Failure | Recovery | | ----------------------- | ----------------------------------------------------------------------- | | Voice doesn't fire | Cut to the card; say the line aloud; cards are the default path anyway. | @@ -233,12 +209,16 @@ _Built-by-us callout:_ `agents/podman-live-conversation/agent.py`. | Collision won't trigger | Use the backup recording for that beat; keep narrating. | | Agent flapping | Pre-checked — but if so, `systemctl restart podman-platform-agent`. | + **Rule:** never debug on stage. Narrate, fall back to recording, keep moving. --- + + ## Tight timing summary + | Time | Beat | | ---- | ---------------------------------------------------------- | | 0:00 | Problem (coordination cost) + hook + original-work line | diff --git a/frontend/public/favicon.svg b/frontend/public/favicon.svg index c042fb2..bb59afb 100644 --- a/frontend/public/favicon.svg +++ b/frontend/public/favicon.svg @@ -1,7 +1,64 @@ - - - - P + + PodMan Logo Mark + A minimal rounded robot head with glowing eyes and a blue-violet antenna orb, optimized for white backgrounds. + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/frontend/public/podman-logo.svg b/frontend/public/podman-logo.svg new file mode 100644 index 0000000..bb59afb --- /dev/null +++ b/frontend/public/podman-logo.svg @@ -0,0 +1,64 @@ + + PodMan Logo Mark + A minimal rounded robot head with glowing eyes and a blue-violet antenna orb, optimized for white backgrounds. + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/frontend/src/App.tsx b/frontend/src/App.tsx index 02960d1..81cf07a 100644 --- a/frontend/src/App.tsx +++ b/frontend/src/App.tsx @@ -9,11 +9,7 @@ import { useUser, } from '@clerk/react'; import type { Room } from 'livekit-client'; -import { - AlertCircleIcon, - RefreshCwIcon, - SparklesIcon, -} from 'lucide-react'; +import { AlertCircleIcon, RefreshCwIcon, SparklesIcon } from 'lucide-react'; import type { Pod, PodInput } from '@podman/shared'; import { joinPod } from './lib/pod.js'; import * as api from './lib/api.js'; @@ -342,9 +338,12 @@ export default function App() {
-
- PM -
+

PodMan

@@ -468,9 +467,12 @@ function AuthGate() {
-
- PM -
+

PodMan

@@ -485,8 +487,8 @@ function AuthGate() { Create your account to enter PodMan

- PodMan saves your context across pods so agents can learn from your work in every - room you join. + PodMan saves your context across pods so agents can learn from your work in every room + you join.