docs(demo): tighten 1-min script to a scannable half-page

This commit is contained in:
Ramis
2026-06-28 11:28:13 -07:00
parent 70b41104ad
commit 28f3236d76
5 changed files with 319 additions and 127 deletions
+88 -108
View File
@@ -5,7 +5,7 @@
**4-minute script** for the complete story. Practice the 4-min to land at 3:45.
**The thesis (say this first, in either script):** AI is collapsing the cost of
_writing_ code. More and more of every codebase is authored with AI assist — and
*writing* code. More and more of every codebase is authored with AI assist — and
increasingly by **autonomous agents**. So the bottleneck of software engineering
is shifting away from engineering itself toward **organization, management, and
project coordination**. Pair a team with hundreds of agents all committing to the
@@ -13,8 +13,8 @@ same repo at once and it becomes **physically impossible for a human to track
progress or avoid stepping on someone else's work.** Code generation scaled;
human coordination did not. That gap is the new bottleneck.
**The one-line story:** writing code isn't the bottleneck anymore — _coordinating
who (and what) is writing what_ is. PodMan is a pair programmer for the whole
**The one-line story:** writing code isn't the bottleneck anymore — *coordinating
who (and what) is writing what* is. PodMan is a pair programmer for the whole
team — humans **and** agents: it watches every actor's work in real time, gives
everyone live status without anyone having to interrupt anyone, catches collisions
before they land, and learns your team's dynamics so it nudges less and helps more
@@ -28,85 +28,52 @@ every actor, every day, is the value.
---
## 1-Minute Live Script (you + teammates — real-time interventions)
**Goal:** in 60 seconds, prove the unprecedented loop — real-time multimodal
perception → live collision intervention spoken by an **on-device Gemma model**
→ a **self-improving** coordination layer. Hits four judging tracks: Continual
Learning, the Self-Improvement Stack, Recursive Intelligence, and on-device
Gemma. **Hard limit:** 1:00.
**Hard limit 1:00.** Hits the tracks: real-time multimodal, on-device Gemma,
Continual Learning, Self-Improvement, Recursive Intelligence.
**Pre-flight (do NOT skip — these are the win conditions):**
**Pre-flight:** pod open on `podman.live`, all joined, screen share + **sound on**;
git watchers running; everyone has `README.md` open. On-device **Gemma** must be
up as Hermes's voice (else Gemini TTS auto-falls-back — just drop the "on-device"
word).
- **On-device Gemma must be committed and running** as Hermes's voice/reasoning
layer before showtime. Verify it speaks once in rehearsal. If it isn't up,
PodMan **automatically falls back to Gemini TTS** — the demo still works, you
just drop the "on-device" line. Never debug Gemma on stage.
- Pod open on the shared screen (`podman.live`), all teammates joined, screen
share on, **sound on** — voice is live for _every_ intervention now.
- Each teammate's git watcher running so local unpushed edits report in.
- Everyone has `README.md` open, ready to type. Pre-pick the first colliding
pair (you + a teammate).
**0:000:12 — Thesis**
> "Code is almost free to write now — humans with AI, soon swarms of agents, all
> on one repo. No human can track that. The bottleneck is coordination — and it
> has to fix itself, live. Watch."
### 0:000:12The thesis (one breath)
**0:120:35Collision caught live (money shot)**
- You + a teammate both edit `README.md`, save (unpushed).
- Card appears on every screen + voice fires: _"<you> and <teammate> are both
editing README.md. Sync before pushing."_
- Say: "Gemini Vision reads our screens, fused with local git, in real time —
and the alert is **spoken by a Gemma model on-device**. Neither of us asked."
> "Code is almost free to write now — humans with AI today, swarms of autonomous
> agents tomorrow, all committing to one repo. No human can track that. The
> bottleneck stopped being engineering and became coordination — and it has to
> fix itself, on its own, in real time. Watch."
**0:350:52 — Self-improving (narrate)**
- Third teammate edits too → **new** card + voice for the new pair (never goes
silent).
- Say: "Every trace lands in **Atlas**; vector search recalls past collisions,
repeats come back tagged **'Seen before'** and escalate — no retraining. It
tunes its own coordination from its own outcomes."
_On screen:_ pod view, teammate tiles live, activity stream moving.
**0:521:00 — Close**
> "Real-time awareness, on-device voice, a coordination layer that improves
> itself — for humans and agents, zero interruptions."
### 0:120:32 — Real-time perception → live intervention (the money shot)
- On cue, **you and one teammate both edit `README.md`** and save (unpushed).
- A **conflict card appears on everyone's screen** within a cycle: _"Conflict:
<you> + <teammate> both on README.md (unpushed)."_
- The voice fires out loud over LiveKit: _"<you> and <teammate> are both editing
README.md. Please sync before pushing."_
- Land both the perception and the on-device beat: "PodMan is reading our actual
screens with **Gemini Vision** and fusing local git truth in real time — then
that alert is reasoned and **spoken by a Gemma model running on-device**, no
round-trip to a frontier API. Real-time multimodal in, on-device voice out.
Neither of us asked. Neither of us was watching."
### 0:320:48 — Self-improving coordination (narrate the stack)
- Have a **third teammate** edit the same file (or a second one). A **new** card
+ voice fires for the new pair — _it never goes silent on later collisions._
- Narrate the self-improvement architecture (no trigger needed — it's running):
"Every intervention and every accept/dismiss is a trace in **MongoDB Atlas**.
PodMan recalls similar past collisions with **vector search**, an
outcome-conditioned policy adapts who to nudge and how, and repeats come back
tagged **'Seen before'** and escalate automatically. That's the
self-improvement stack — it gets sharper the more the team uses it, with no
retraining and almost no human input. The recursive part: it's tuning its own
coordination behavior from its own logged outcomes."
### 0:481:00 — Close
> "Real-time multimodal awareness, on-device voice, and a coordination layer that
> improves itself from every interaction — for a whole team of humans and agents,
> with zero interruptions. That's the layer code generation never had."
**Track callouts (say at least the bolded word in each):**
| Track | Spoken moment |
| ----------------------- | -------------------------------------------------------- |
| Real-time multimodal | "reading our screens with **Gemini Vision** … real time" |
| **On-device Gemma** | "spoken by a **Gemma** model running **on-device**" |
| Continual Learning | "recalls past collisions … **'Seen before'** … no retraining" |
| Self-Improvement Stack | "every trace in **Atlas** … outcome-conditioned policy" |
| Recursive Intelligence | "tuning its **own** coordination behavior from its outcomes" |
_Fallbacks:_ Gemma down → drop the on-device line, Gemini TTS covers it. Voice
silent → read the line aloud, point at the card (cards are the default path).
Card won't trigger → switch the colliding pair to a fresh file and re-save.
_Fallback:_ Gemma down → drop the on-device line. Voice silent → read it, point
at the card. No card → switch the pair to a fresh file and re-save.
---
## The script (4:00)
### 0:000:30 — The problem + hook
> "AI made writing code almost free — humans with AI today, swarms of autonomous
@@ -119,48 +86,48 @@ Card won't trigger → switch the colliding pair to a fresh file and re-save.
> interrupting them, catches collisions before they land, and learns your team as
> it goes."
_On screen:_ the pod view, two teammates joined, screen-share tiles live.
*On screen:* the pod view, two teammates joined, screen-share tiles live.
### 0:301:05 — Real-time team awareness (LiveKit + Gemini Vision)
- Point at the two live screen tiles. "These are real screen shares over
**LiveKit**. Our agent subscribes to the tracks and samples frames."
**LiveKit**. Our agent subscribes to the tracks and samples frames."
- "Each frame goes to **Gemini Vision**, which returns structured context — file,
symbol, activity — not a chatbot, a perception layer."
symbol, activity — not a chatbot, a perception layer."
- Show the live activity stream filling in (Signals vs Reasoning sections).
- Land the value: "This is the part that replaces 'what are you working on?' —
every teammate's current work is just _visible_, in real time. Nobody had to
ask."
every teammate's current work is just *visible*, in real time. Nobody had to
ask."
_Built-by-us callout:_ `backend/src/vision/gemini.ts`, the LiveKit agent worker.
*Built-by-us callout:* `backend/src/vision/gemini.ts`, the LiveKit agent worker.
### 1:051:40 — The catch (detection + first intervention)
- Have alice and bob both edit the **same file** with unpushed changes.
- "Normally nobody notices until merge time. GitHub can't see this — nothing's
pushed. Our detector fuses live screen context with **local git truth** from a
watcher on each laptop."
- A collision card appears: _"alice + bob both on detector.ts (unpushed)."_
- Let the **Gemini TTS** urgent voice fire once over LiveKit: _"alice and bob are
both editing detector.ts. Please sync before pushing."_
pushed. Our detector fuses live screen context with **local git truth** from a
watcher on each laptop."
- A collision card appears: *"alice + bob both on detector.ts (unpushed)."*
- Let the **Gemini TTS** urgent voice fire once over LiveKit: *"alice and bob are
both editing detector.ts. Please sync before pushing."*
- Land the value: "That's a merge conflict and a wasted afternoon caught before it
happened — and neither of them had to be tracking the other."
happened — and neither of them had to be tracking the other."
_Built-by-us callout:_ `collision/detector.ts`, `action/hermes.ts`,
*Built-by-us callout:* `collision/detector.ts`, `action/hermes.ts`,
`voice/live.ts`.
### 1:402:10 — Cross-channel overlap (research + code)
- Keep alice editing `livekit.py`.
- Have bob share a browser tab on LiveKit docs/SDK pages.
- A collaboration nudge appears: _"🤝 bob is researching LiveKit agents
(docs.livekit.io) while alice edits livekit.py — sync up before duplicating
effort."_
- A collaboration nudge appears: *"🤝 bob is researching LiveKit agents
(docs.livekit.io) while alice edits livekit.py — sync up before duplicating
effort."*
- Land the value: "This is not a merge conflict. PodMan caught duplicated effort
across channels — code on one screen, research on another — and nudged the team
before two people solved the same problem twice."
across channels — code on one screen, research on another — and nudged the team
before two people solved the same problem twice."
_Built-by-us callout:_ `vision/gemini.ts`, `collision/research.ts`,
*Built-by-us callout:* `vision/gemini.ts`, `collision/research.ts`,
`memory/vectors.ts`.
### 2:102:50 — Continual learning (the theme — the money shot)
@@ -168,53 +135,58 @@ _Built-by-us callout:_ `vision/gemini.ts`, `collision/research.ts`,
This is the differentiator. Two beats, both from pre-seeded memory:
1. **It learned to stay quiet.** Trigger a pattern that was dismissed as a false
alarm earlier. "Last session a teammate marked this kind of alert as not a
alarm earlier. "Last session a teammate marked this kind of alert as not a
real conflict. Watch — PodMan stays silent. No nagging." (No card fires.)
2. **It learned to escalate.** Trigger the real-conflict pattern that was
accepted before. The card now says **"Seen before."** and goes straight to
accepted before. The card now says **"Seen before."** and goes straight to
the spoken urgent cue.
- "The only input was one accept/dismiss tap. No retraining, no labeling. This is
**MongoDB Atlas vector search** recalling similar past events plus a policy
that adapts on the recalled outcome."
**MongoDB Atlas vector search** recalling similar past events plus a policy
that adapts on the recalled outcome."
- Optional: show `/api/memory/stats` counts climbing — accumulated experience.
_Built-by-us callout:_ `memory/vectors.ts` ($vectorSearch), `memory/policy.ts`
*Built-by-us callout:* `memory/vectors.ts` ($vectorSearch), `memory/policy.ts`
(outcome-conditioned gate), `memory/store.ts`.
### 2:503:30 — The five-minute meeting, killed (Gemini Live API)
- Frame it: "Instead of breaking a teammate's focus to ask what they're up to,
you ask PodMan."
- Open the live voice conversation. Ask out loud: _"PodMan, what is everyone
working on, and where is the collision detector implemented?"_
you ask PodMan."
- Open the live voice conversation. Ask out loud: *"PodMan, what is everyone
working on, and where is the collision detector implemented?"*
- It answers with **real tool calls**`search_repo`, git history, current
collisions — not guesses.
collisions — not guesses.
- "This is the **Gemini Live API**, streaming speech-to-speech over LiveKit, with
custom function tools we wrote so it grounds every answer in the actual repo
and live state. That's the status sync, answered in seconds, with zero recovery
tax on anyone else."
custom function tools we wrote so it grounds every answer in the actual repo
and live state. That's the status sync, answered in seconds, with zero recovery
tax on anyone else."
_Built-by-us callout:_ `agents/podman-live-conversation/agent.py`.
*Built-by-us callout:* `agents/podman-live-conversation/agent.py`.
### 3:303:50 — Stack + close
- "All on **DigitalOcean** — static frontend, API, and agent workers, supervised
by systemd. The ambient score is **Gemini Lyria** generated per pod through the
Interactions API."
by systemd. The ambient score is **Gemini Lyria** generated per pod through the
Interactions API."
- Close: "Engineering ability stopped being the bottleneck — coordination is, and
it only gets worse as agents start writing alongside us. PodMan gives a whole
team, humans and agents, real-time awareness without the interruptions, catches
collisions before they cost an afternoon, and learns each team's dynamics so it
helps more over time. Saved focus, multiplied across every actor. That's
continual learning, shipped."
it only gets worse as agents start writing alongside us. PodMan gives a whole
team, humans and agents, real-time awareness without the interruptions, catches
collisions before they cost an afternoon, and learns each team's dynamics so it
helps more over time. Saved focus, multiplied across every actor. That's
continual learning, shipped."
### 3:504:00 — Buffer / Q&A handoff
---
## Sponsor-prize coverage (say each at least once)
| Prize | Spoken moment | Segment |
| ---------------- | -------------------------------------------------------------------------------- | ---------------------- |
| **Gemini** | Vision perception, Live API agent w/ tools, TTS voice, Lyria score | 0:30, 1:05, 2:50, 3:30 |
@@ -222,10 +194,14 @@ _Built-by-us callout:_ `agents/podman-live-conversation/agent.py`.
| **MongoDB** | "Atlas vector search recalling past events" | 1:50 |
| **DigitalOcean** | "all on DigitalOcean, systemd-supervised workers" | 3:30 |
---
## If something breaks (live recovery)
| Failure | Recovery |
| ----------------------- | ----------------------------------------------------------------------- |
| Voice doesn't fire | Cut to the card; say the line aloud; cards are the default path anyway. |
@@ -233,12 +209,16 @@ _Built-by-us callout:_ `agents/podman-live-conversation/agent.py`.
| Collision won't trigger | Use the backup recording for that beat; keep narrating. |
| Agent flapping | Pre-checked — but if so, `systemctl restart podman-platform-agent`. |
**Rule:** never debug on stage. Narrate, fall back to recording, keep moving.
---
## Tight timing summary
| Time | Beat |
| ---- | ---------------------------------------------------------- |
| 0:00 | Problem (coordination cost) + hook + original-work line |