Files
podman/docs/PLAN.md
T
Kartikeya 253c7438fb style: apply prettier formatting to docs
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 14:17:14 -07:00

189 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# PodMan — Master Plan (v1)
> Living doc. A deeper, API-accurate v2 (exact Gemini/LiveKit SDK calls, starter code,
> DO deploy steps) is being generated by the research workflow and will be merged in.
---
## 0. TL;DR
**PodMan** is an ambient AI teammate. Engineers join a **pod**, share screen + mic, and
PodMan watches everyone's screen in realtime, understands what each person is doing
(Gemini vision), fuses it with the team's GitHub state, and **interrupts like Jarvis to
prevent merge collisions before anyone pushes**, offering to open a sync PR.
- **Track:** Continual Learning
- **Prizes targeted (stacked):** Gemini 3.5 ($5k cash), LiveKit (keyboards), DigitalOcean (credits)
- **Hero moment:** two laptops editing the same file → PodMan _speaks up live_ and offers the fix.
---
## 1. Why this fits the track (and stays eligible)
### Track = Continual Learning
The official definition rewards systems that "continuously improve from real-world use…
becoming more useful the more they are used with as little user intervention as possible."
PodMan does exactly this:
- Builds and **continuously refines a live model of the team** — who owns which files/areas,
what's in-flight, recurring conflict patterns, each engineer's working style.
- **Self-improves its own intervention policy** from outcomes: did the collision it predicted
actually happen? Did the team accept the suggested PR? It tunes its thresholds/prompts so it
nags less and helps more over time.
- Grows a per-team **skill/memory store** (vector memory) that makes later sessions sharper.
### ⚠️ Disqualification traps — and how we dodge them
| Risk | Mitigation |
| --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **"Dashboard is the main feature" = auto-DQ** | The pod grid is _secondary_. The hero is PodMan's **proactive voice/card interventions**. In the demo we barely show the grid; we show PodMan _acting_. |
| Repo must be **public** | Make the GitHub repo public from the start. |
| **Only what you built** during the event | Everything in this monorepo is new, timestamped by commits. Demo narrates "built today." |
| New work only | No pre-existing project reuse. |
---
## 2. Prize-stacking map
| Prize | How PodMan earns it |
| --------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Best Gemini 3.5 ($5,000 cash)** | Realtime screen understanding via Gemini 3.5 Flash vision + PodMan's voice via Gemini Live API. Bonus: Live Translate so a multilingual pod hears PodMan in their language. |
| **Best LiveKit (keyboards)** | LiveKit is the realtime backbone: screen-share + mic + cam tracks in, PodMan voice + data-channel cards out. It's load-bearing, not bolted on. |
| **Best DigitalOcean (credits)** | Backend PodMan agent + LiveKit agent deployed on DigitalOcean; claim the $200 credits. |
---
## 3. Architecture (v1 — refined by workflow)
```
┌────────────── Engineer laptops (Chrome PWA) ──────────────┐
│ getDisplayMedia (screen) + mic + cam │
│ publish tracks ─────────────┐ ▲ PodMan voice │
└──────────────────────────────┼─────────┼──────────────────┘
│ │ data-channel cards
┌───────▼─────────┴────────┐
│ LiveKit room │ (one room per pod)
└───────┬─────────▲──────────┘
│ subscribe│ publish voice/data
┌───────▼─────────┴──────────────────────┐
│ BACKEND: PodMan agent (DigitalOcean) │
│ │
│ 1. grab frames from each screen track │
│ 2. Gemini 3.5 vision → structured │
│ "engineer context" (file, symbol, │
│ feature, action) │
│ 3. GitHub client → branches/PRs/commits │
│ 4. COLLISION DETECTOR (fuse 2+3) │
│ 5. continual-learning memory (Atlas + │
│ Voyage vectors): team model + policy │
│ 6. PodMan brain (Gemini) → intervention │
│ 7. speak (Gemini Live/TTS) + send card │
└──────────────────────────────────────────┘
```
**The critical reconciliation:** GitHub only knows _pushed_ state. The "X is editing this and
hasn't pushed" signal comes from **vision on the live screen** (filename in the editor tab,
visible diff/gutter), optionally cross-checked by an _optional_ lightweight local `git status`
reporter the engineer can run. Vision is the headline; the local reporter is a nice-to-have.
### Continual-learning loop
1. **Observe** — per-engineer context every few seconds (sampled frames, not every frame).
2. **Store** — append observations to the team model; embed file/feature notes into Voyage
vectors in Atlas for retrieval.
3. **Predict** — collision detector + PodMan brain decide if/when to intervene.
4. **Outcome** — record whether the warning was acted on / was a true positive.
5. **Adapt** — adjust intervention thresholds, ownership attribution, and prompt context from
outcomes → fewer false alarms, better targeting over the session. _(This is the "gets
better the more you use it" story judges want.)_
---
## 4. Stack per folder
| Folder | Stack |
| ----------- | ----------------------------------------------------------------------------------------------------------- |
| `frontend/` | React + Vite + TypeScript, `livekit-client`, PWA (vite-plugin-pwa), Tailwind |
| `backend/` | Node + TypeScript, LiveKit server SDK + agents, `@google/genai`, GitHub (Octokit or GitHub MCP), Express/ws |
| `database/` | MongoDB Atlas (team model, observations, outcomes) + Voyage embeddings for vector recall |
| `infra/` | DigitalOcean App Platform / Droplet, Dockerfile, app spec |
| `shared/` | TS types: `Pod`, `EngineerContext`, `Collision`, `Intervention` |
---
## 5. 20-hour build order (MVP-first)
1. **Plumbing** — monorepo installs, shared types, env wiring, LiveKit token endpoint. _(Karti)_
2. **Capture** — frontend: join pod + publish screen/mic; render PodMan card + play voice. _(Zander)_
3. **Eyes** — backend: subscribe to a screen track, grab a frame, Gemini vision → `EngineerContext`. _(Ramis)_
4. **Brain + collision** — fuse two engineers' contexts + GitHub state → detect same-file/feature; PodMan brain composes the intervention. _(Yahya)_
5. **Voice + action** — PodMan speaks (Live API/TTS) into the room + "Open sync PR" via GitHub. _(Yahya + Ramis)_
6. **Memory/continual learning** — store observations + outcomes; show the team model improving. _(Karti)_
7. **Deploy on DO + polish demo** — everyone. Rehearse the live demo 3×.
> If behind: cut webcam, cut multilingual, cut the local git reporter, **mock the QR join**,
> hardcode the demo repo. Never cut: realtime screen→vision→PodMan-speaks loop.
---
## 6. Team split (4 max — must drop to 4!)
> ⚠️ Roster has 5 (Karti, Ramis, Yahya, Zander, Shakthi). **Max team size is 4.** Decide who's
> the official 4 before submission, or one stays unofficial/support.
| Person | Owns |
| ---------- | ---------------------------------------------------------------------------------- |
| **Karti** | Repo/infra/plumbing, shared types, memory + continual-learning store, DO deploy |
| **Zander** | Frontend PWA: pod join, capture, PodMan card UI + voice playback |
| **Ramis** | Backend realtime: LiveKit room subscribe + frame grab + Gemini vision pipeline |
| **Yahya** | PodMan brain: collision detector, intervention policy, GitHub PR action, voice out |
---
## 7. Demo script (3 min — refined by workflow)
1. **(0:00)** Two laptops on screen. Both engineers "join the pod" (QR mock). PodMan greets them by voice.
2. **(0:30)** Engineer A opens `auth.ts` and starts editing. PodMan quietly notes it (show the team model tick).
3. **(1:00)** Engineer B opens the _same_ `auth.ts` and edits a related function — **neither has pushed.**
4. **(1:20) MONEY MOMENT** — PodMan _interrupts by voice_: "Heads up — Karti and Yahya are both in `auth.ts`, Yahya has unpushed changes. Here's the diff. Want me to open a sync PR?" Card appears with the diff.
5. **(1:50)** One click → PodMan opens a draft PR via GitHub (show it on github.com).
6. **(2:20)** Show it **learned**: PodMan now knows Karti owns auth; second scenario it's faster/quieter where appropriate → "more useful the more you use it."
7. **(2:45)** One-liner close: "PodMan — the teammate that sees what git can't."
---
## 8. Env vars (v1 — finalized by workflow)
```
# LiveKit
LIVEKIT_URL=
LIVEKIT_API_KEY=
LIVEKIT_API_SECRET=
# Gemini
GEMINI_API_KEY=
GEMINI_VISION_MODEL=gemini-3.5-flash # confirm exact id from research
GEMINI_LIVE_MODEL= # confirm from research
# GitHub
GITHUB_TOKEN=
GITHUB_REPO=owner/name
# MongoDB Atlas + Voyage
MONGODB_URI=
VOYAGE_API_KEY=
```
---
## 9. Open risks
| Risk | Mitigation |
| --------------------------------------------------- | -------------------------------------------------------------------------------- |
| Realtime vision latency/cost | Sample ~1 frame/sec or on-change; downscale frames; cache last context |
| LiveKit ↔ Gemini frame plumbing is the hardest part | Build & de-risk it **first** (step 3); have a screenshot-fallback path |
| "Unpushed" detection is fuzzy | Lead with vision; optional local `git status` reporter for accuracy |
| On-stage flakiness | Pre-stage the demo repo, rehearse 3×, have a recorded backup of the money moment |
| Dashboard-DQ optics | Keep UI minimal; demo PodMan _acting_, not a grid |