docs: update README with Modular MAX architecture
This commit is contained in:
@@ -9,7 +9,22 @@ down.
|
|||||||
|
|
||||||
[LiveKit](https://livekit.io/) · [MongoDB](https://www.mongodb.com/) ·
|
[LiveKit](https://livekit.io/) · [MongoDB](https://www.mongodb.com/) ·
|
||||||
[Gemini](https://ai.google.dev/) · [Hermes](https://hermes-agent.nousresearch.com/) ·
|
[Gemini](https://ai.google.dev/) · [Hermes](https://hermes-agent.nousresearch.com/) ·
|
||||||
[vLLM](https://vllm.ai/) · [DigitalOcean](https://www.digitalocean.com/)
|
[Modular MAX](https://www.modular.com/max) ·
|
||||||
|
[DigitalOcean](https://www.digitalocean.com/)
|
||||||
|
|
||||||
|
<p align="center">
|
||||||
|
<img
|
||||||
|
src="docs/assets/interruption-flow-recovery.jpg"
|
||||||
|
alt="A chart showing that a five minute interruption can create a much longer recovery period before a developer returns to flow state."
|
||||||
|
width="760"
|
||||||
|
/>
|
||||||
|
</p>
|
||||||
|
|
||||||
|
> PodMan exists to prevent the five-minute interruption from becoming a
|
||||||
|
> half-hour recovery tax.
|
||||||
|
|
||||||
|
Source image: `~/pic.jpg`, committed as
|
||||||
|
`docs/assets/interruption-flow-recovery.jpg` so it renders on GitHub.
|
||||||
|
|
||||||
## Read This First
|
## Read This First
|
||||||
|
|
||||||
@@ -19,7 +34,8 @@ down.
|
|||||||
| Check the app | `curl https://podman.live/health` |
|
| Check the app | `curl https://podman.live/health` |
|
||||||
| Check pods and memory | `curl https://podman.live/api/pods && curl https://podman.live/api/memory/stats` |
|
| Check pods and memory | `curl https://podman.live/api/pods && curl https://podman.live/api/memory/stats` |
|
||||||
| Use the LLM externally | Base URL `https://llm.alhinai.dev/v1`, model `gemma-4-31B-it` |
|
| Use the LLM externally | Base URL `https://llm.alhinai.dev/v1`, model `gemma-4-31B-it` |
|
||||||
| Test Hermes | `hermes -z 'Reply with exactly: working' --provider gemma4-31b-vllm --model gemma-4-31B-it` |
|
| Test Hermes | `hermes -z 'Reply with exactly: working' --provider gemma4-31b-max --model gemma-4-31B-it` |
|
||||||
|
| Gemma 4 bring-up | Modular/MAX lab at `/home/alhinai/modular-lab` |
|
||||||
| Start local development | API, vision agent, and frontend commands are in [Local Development](#local-development) |
|
| Start local development | API, vision agent, and frontend commands are in [Local Development](#local-development) |
|
||||||
| Debug production | Public checks first, then systemd services in [Production Operations](#production-operations) |
|
| Debug production | Public checks first, then systemd services in [Production Operations](#production-operations) |
|
||||||
|
|
||||||
@@ -33,10 +49,11 @@ down.
|
|||||||
| API | `https://podman.live/api/*` | Pods, tokens, outcomes, memory |
|
| API | `https://podman.live/api/*` | Pods, tokens, outcomes, memory |
|
||||||
| Health | `https://podman.live/health` | Backend readiness |
|
| Health | `https://podman.live/health` | Backend readiness |
|
||||||
| Local API | `127.0.0.1:8787` | Express service behind Caddy |
|
| Local API | `127.0.0.1:8787` | Express service behind Caddy |
|
||||||
| Reasoning LLM | `https://llm.alhinai.dev/v1` | OpenAI-compatible Gemma/vLLM endpoint |
|
| Reasoning LLM | `https://llm.alhinai.dev/v1` | OpenAI-compatible Modular/MAX endpoint |
|
||||||
|
| Gemma MAX lab | `/home/alhinai/modular-lab` | Modular/MAX bring-up path for Gemma 4 |
|
||||||
| API key | `not-needed` | Placeholder key for OpenAI clients |
|
| API key | `not-needed` | Placeholder key for OpenAI clients |
|
||||||
| Hermes model | `gemma-4-31B-it` | 262K-context tool-using agent |
|
| Hermes model | `gemma-4-31B-it` | 262K-context tool-using agent |
|
||||||
| Hermes config | `gemma4-31b-vllm` | Custom provider used by Hermes locally |
|
| Hermes config | `gemma4-31b-max` | Custom provider used by Hermes locally |
|
||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart TB
|
flowchart TB
|
||||||
@@ -66,8 +83,10 @@ flowchart TB
|
|||||||
|
|
||||||
subgraph Reasoning["External Reasoning Endpoint"]
|
subgraph Reasoning["External Reasoning Endpoint"]
|
||||||
Tunnel["Cloudflare Tunnel<br/>llm.alhinai.dev"]
|
Tunnel["Cloudflare Tunnel<br/>llm.alhinai.dev"]
|
||||||
VLLM["vLLM OpenAI Server<br/>gemma-4-31B-it<br/>262144 context"]
|
MaxLab["Modular MAX Lab<br/>/home/alhinai/modular-lab"]
|
||||||
Hermes["Hermes Agent<br/>provider: gemma4-31b-vllm"]
|
MaxServe["MAX OpenAI Server<br/>gemma-4-31B-it<br/>262144 context"]
|
||||||
|
MaxTest["MAX test port<br/>127.0.0.1:8001"]
|
||||||
|
Hermes["Hermes Agent<br/>provider: gemma4-31b-max"]
|
||||||
end
|
end
|
||||||
|
|
||||||
Browser --> Caddy --> API
|
Browser --> Caddy --> API
|
||||||
@@ -84,7 +103,10 @@ flowchart TB
|
|||||||
Voice --> GeminiVoice
|
Voice --> GeminiVoice
|
||||||
Vision --> GeminiEmbed
|
Vision --> GeminiEmbed
|
||||||
Voice --> GeminiEmbed
|
Voice --> GeminiEmbed
|
||||||
Ops --> Hermes --> Tunnel --> VLLM
|
Ops --> Hermes --> Tunnel --> MaxServe
|
||||||
|
MaxLab --> MaxTest
|
||||||
|
MaxTest -. "validated Gemma 4<br/>staging path" .-> MaxServe
|
||||||
|
Tunnel -. "public route<br/>llm.alhinai.dev/v1" .-> MaxServe
|
||||||
API --> Hermes
|
API --> Hermes
|
||||||
Lyria --> Voice
|
Lyria --> Voice
|
||||||
|
|
||||||
@@ -95,7 +117,7 @@ flowchart TB
|
|||||||
class Browser,Room user;
|
class Browser,Room user;
|
||||||
class Caddy,API,Vision,Voice,Ops app;
|
class Caddy,API,Vision,Voice,Ops app;
|
||||||
class Observations,State,Outcomes data;
|
class Observations,State,Outcomes data;
|
||||||
class GeminiVision,GeminiVoice,GeminiEmbed,Lyria,Tunnel,VLLM,Hermes ai;
|
class GeminiVision,GeminiVoice,GeminiEmbed,Lyria,Tunnel,MaxLab,MaxServe,MaxTest,Hermes ai;
|
||||||
```
|
```
|
||||||
|
|
||||||
## What It Does
|
## What It Does
|
||||||
@@ -138,7 +160,9 @@ flowchart LR
|
|||||||
VisionRoute --> Gemini20["gemini-2.0-flash<br/>screen understanding"]
|
VisionRoute --> Gemini20["gemini-2.0-flash<br/>screen understanding"]
|
||||||
ConversationRoute --> GeminiLive["gemini-3.1-flash-live-preview<br/>live Q&A"]
|
ConversationRoute --> GeminiLive["gemini-3.1-flash-live-preview<br/>live Q&A"]
|
||||||
ConversationRoute --> GeminiTTS["gemini-3.1-flash-tts-preview<br/>urgent speech"]
|
ConversationRoute --> GeminiTTS["gemini-3.1-flash-tts-preview<br/>urgent speech"]
|
||||||
HermesRoute --> Gemma["gemma-4-31B-it<br/>Hermes + vLLM + tools"]
|
HermesRoute --> Gemma["gemma-4-31B-it<br/>public serving via Modular MAX"]
|
||||||
|
HermesRoute --> Max["Modular MAX lab<br/>GB10 Gemma 4 validation"]
|
||||||
|
Max -. "test on :8001<br/>promote to public endpoint" .-> Gemma
|
||||||
EmbedRoute --> Voyage["voyage-4-lite<br/>primary embeddings"]
|
EmbedRoute --> Voyage["voyage-4-lite<br/>primary embeddings"]
|
||||||
EmbedRoute --> GeminiEmbed["gemini-embedding-001<br/>fallback embeddings"]
|
EmbedRoute --> GeminiEmbed["gemini-embedding-001<br/>fallback embeddings"]
|
||||||
MusicRoute --> Lyria["lyria-3-clip-preview<br/>background music"]
|
MusicRoute --> Lyria["lyria-3-clip-preview<br/>background music"]
|
||||||
@@ -148,10 +172,10 @@ flowchart LR
|
|||||||
classDef model fill:#f4edff,stroke:#805ad5,color:#2d1857;
|
classDef model fill:#f4edff,stroke:#805ad5,color:#2d1857;
|
||||||
class Work,VoiceInput,OpsNeed,MemoryNeed,Ambient signal;
|
class Work,VoiceInput,OpsNeed,MemoryNeed,Ambient signal;
|
||||||
class VisionRoute,ConversationRoute,HermesRoute,EmbedRoute,MusicRoute route;
|
class VisionRoute,ConversationRoute,HermesRoute,EmbedRoute,MusicRoute route;
|
||||||
class Gemini20,GeminiLive,GeminiTTS,Gemma,Voyage,GeminiEmbed,Lyria model;
|
class Gemini20,GeminiLive,GeminiTTS,Gemma,Max,Voyage,GeminiEmbed,Lyria model;
|
||||||
```
|
```
|
||||||
|
|
||||||
### Gemma 4 31B via vLLM
|
### Gemma 4 31B via Modular MAX
|
||||||
|
|
||||||
Hermes uses the external OpenAI-compatible endpoint:
|
Hermes uses the external OpenAI-compatible endpoint:
|
||||||
|
|
||||||
@@ -162,17 +186,16 @@ Model: gemma-4-31B-it
|
|||||||
Context: 262144 tokens
|
Context: 262144 tokens
|
||||||
```
|
```
|
||||||
|
|
||||||
The vLLM server is configured for Hermes-style tool use:
|
The Modular/MAX server is configured for long-context Gemma 4 serving:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
--max-model-len 262144
|
--max-length 262144
|
||||||
--max-num-seqs 1
|
--device-memory-utilization 0.85
|
||||||
--max-num-batched-tokens 16384
|
--kv-cache-format float8_e4m3fn
|
||||||
--gpu-memory-utilization 0.85
|
--enable-prefix-caching
|
||||||
--kv-cache-dtype fp8
|
|
||||||
--enable-auto-tool-choice
|
|
||||||
--tool-call-parser gemma4
|
|
||||||
--enable-chunked-prefill
|
--enable-chunked-prefill
|
||||||
|
--max-batch-size 1
|
||||||
|
--max-batch-input-tokens 16384
|
||||||
```
|
```
|
||||||
|
|
||||||
Hermes should point at that endpoint with this provider shape:
|
Hermes should point at that endpoint with this provider shape:
|
||||||
@@ -180,11 +203,11 @@ Hermes should point at that endpoint with this provider shape:
|
|||||||
```yaml
|
```yaml
|
||||||
model:
|
model:
|
||||||
default: gemma-4-31B-it
|
default: gemma-4-31B-it
|
||||||
provider: gemma4-31b-vllm
|
provider: gemma4-31b-max
|
||||||
|
|
||||||
providers:
|
providers:
|
||||||
gemma4-31b-vllm:
|
gemma4-31b-max:
|
||||||
name: Gemma 4 31B vLLM (256K)
|
name: Gemma 4 31B Modular MAX (256K)
|
||||||
api: https://llm.alhinai.dev/v1
|
api: https://llm.alhinai.dev/v1
|
||||||
api_key: not-needed
|
api_key: not-needed
|
||||||
transport: chat_completions
|
transport: chat_completions
|
||||||
@@ -199,8 +222,102 @@ agent:
|
|||||||
```
|
```
|
||||||
|
|
||||||
Why this matters: Hermes sends OpenAI tool schemas and `tool_choice: "auto"`.
|
Why this matters: Hermes sends OpenAI tool schemas and `tool_choice: "auto"`.
|
||||||
Without `--enable-auto-tool-choice` and `--tool-call-parser gemma4`, vLLM returns
|
The model endpoint must support automatic tool choice so Hermes can initialize
|
||||||
HTTP 400 before Hermes can initialize an agent.
|
the agent without a client-side workaround.
|
||||||
|
|
||||||
|
### Modular MAX for Gemma 4
|
||||||
|
|
||||||
|
The GB10 box uses a working Modular/MAX lab as the Gemma 4 bring-up,
|
||||||
|
validation, and serving path. Run MAX on `8001` for testing, then promote it to
|
||||||
|
the public endpoint when ready.
|
||||||
|
|
||||||
|
Installed local Codex skills:
|
||||||
|
|
||||||
|
```text
|
||||||
|
/home/alhinai/.codex/skills/new-modular-project
|
||||||
|
/home/alhinai/.codex/skills/mojo-syntax
|
||||||
|
/home/alhinai/.codex/skills/mojo-gpu-fundamentals
|
||||||
|
/home/alhinai/.codex/skills/mojo-python-interop
|
||||||
|
/home/alhinai/.codex/skills/import-model
|
||||||
|
/home/alhinai/.codex/skills/debug-model
|
||||||
|
/home/alhinai/.codex/skills/profile-model
|
||||||
|
```
|
||||||
|
|
||||||
|
Restart Codex after installing or changing those skills so they are picked up.
|
||||||
|
|
||||||
|
Working lab:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd /home/alhinai/modular-lab
|
||||||
|
uv run mojo --version
|
||||||
|
uv run max --version
|
||||||
|
uv run mojo hello.mojo
|
||||||
|
```
|
||||||
|
|
||||||
|
Verified output includes:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Mojo 1.0.0b3.dev2026062806
|
||||||
|
MAX 26.5.0.dev2026062806
|
||||||
|
hello from Mojo on spark
|
||||||
|
```
|
||||||
|
|
||||||
|
Important finding from `max list`:
|
||||||
|
|
||||||
|
| MAX support | Value |
|
||||||
|
| ------------ | ---------------------------------------- |
|
||||||
|
| Architecture | `Gemma4ForConditionalGeneration` |
|
||||||
|
| Example | `google/gemma-4-31B-it` |
|
||||||
|
| Encodings | `float4_e2m1fnx2`, `float16`, `bfloat16` |
|
||||||
|
|
||||||
|
That means Modular/MAX recognizes the Gemma 4 31B architecture on this GB10
|
||||||
|
machine.
|
||||||
|
|
||||||
|
Do not run two GPU model servers at once. Stop the existing Gemma server before
|
||||||
|
starting another MAX serve process:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker stop <current-gemma-container>
|
||||||
|
```
|
||||||
|
|
||||||
|
Then test MAX on a separate port:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd /home/alhinai/modular-lab
|
||||||
|
|
||||||
|
uv run max serve \
|
||||||
|
--model /home/alhinai/models/gemma-4-31B-it \
|
||||||
|
--served-model-name gemma-4-31B-it \
|
||||||
|
--devices gpu:0 \
|
||||||
|
--port 8001 \
|
||||||
|
--trust-remote-code \
|
||||||
|
--quantization-encoding bfloat16 \
|
||||||
|
--max-length 262144 \
|
||||||
|
--kv-cache-format float8_e4m3fn \
|
||||||
|
--device-memory-utilization 0.85 \
|
||||||
|
--enable-prefix-caching \
|
||||||
|
--enable-chunked-prefill \
|
||||||
|
--max-batch-size 1 \
|
||||||
|
--max-batch-input-tokens 16384
|
||||||
|
```
|
||||||
|
|
||||||
|
Verify MAX:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl http://127.0.0.1:8001/v1/models
|
||||||
|
```
|
||||||
|
|
||||||
|
#### Replace the current Gemma server with Modular/MAX
|
||||||
|
|
||||||
|
If MAX works on `8001`, keep it there for testing until the public serving path
|
||||||
|
is intentionally changed.
|
||||||
|
|
||||||
|
Official resources:
|
||||||
|
|
||||||
|
- [MAX LLM Book](https://llm.modular.com/)
|
||||||
|
- [Model bring-up workflow](https://docs.modular.com/max/develop/model-bringup-workflow/)
|
||||||
|
- [Mojo](https://mojolang.org/)
|
||||||
|
- [Modular skills](https://github.com/modular/skills)
|
||||||
|
|
||||||
### Gemini Surfaces
|
### Gemini Surfaces
|
||||||
|
|
||||||
@@ -251,7 +368,7 @@ observe -> detect -> recall prior outcomes -> policy gate -> act -> record outco
|
|||||||
| Vision agent | Node + `@livekit/rtc-node` | Subscribes to screen tracks, samples frames, publishes interventions |
|
| Vision agent | Node + `@livekit/rtc-node` | Subscribes to screen tracks, samples frames, publishes interventions |
|
||||||
| Live voice agent | Python LiveKit Agents + Gemini Live | Real-time voice Q&A in a pod room |
|
| Live voice agent | Python LiveKit Agents + Gemini Live | Real-time voice Q&A in a pod room |
|
||||||
| Memory | MongoDB Atlas | Observations, engineer state, collisions, interventions, outcomes |
|
| Memory | MongoDB Atlas | Observations, engineer state, collisions, interventions, outcomes |
|
||||||
| Reasoning agent | Hermes + Gemma vLLM | Tool-using autonomous assistant and ops layer |
|
| Reasoning agent | Hermes + Gemma Modular/MAX | Tool-using autonomous assistant and ops layer |
|
||||||
| Realtime transport | LiveKit Cloud | Screen tracks, audio tracks, data messages |
|
| Realtime transport | LiveKit Cloud | Screen tracks, audio tracks, data messages |
|
||||||
| Hosting | DigitalOcean + Caddy + systemd | Static frontend, API, workers, watchdog timers |
|
| Hosting | DigitalOcean + Caddy + systemd | Static frontend, API, workers, watchdog timers |
|
||||||
|
|
||||||
@@ -433,7 +550,7 @@ pnpm hermes:sync-deploy
|
|||||||
pnpm deploy:doctor:strict
|
pnpm deploy:doctor:strict
|
||||||
```
|
```
|
||||||
|
|
||||||
Gemma/vLLM endpoint checks:
|
Gemma Modular/MAX endpoint checks:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
curl https://llm.alhinai.dev/v1/models \
|
curl https://llm.alhinai.dev/v1/models \
|
||||||
@@ -481,7 +598,7 @@ POD_ROOM=demo-pod
|
|||||||
Gemma/Hermes provider values live in Hermes config, not PodMan `.env`:
|
Gemma/Hermes provider values live in Hermes config, not PodMan `.env`:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
provider: gemma4-31b-vllm
|
provider: gemma4-31b-max
|
||||||
model: gemma-4-31B-it
|
model: gemma-4-31B-it
|
||||||
base_url: https://llm.alhinai.dev/v1
|
base_url: https://llm.alhinai.dev/v1
|
||||||
api_key: not-needed
|
api_key: not-needed
|
||||||
@@ -514,7 +631,7 @@ For Hermes/Gemma:
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
hermes -z 'Reply with exactly: working' \
|
hermes -z 'Reply with exactly: working' \
|
||||||
--provider gemma4-31b-vllm \
|
--provider gemma4-31b-max \
|
||||||
--model gemma-4-31B-it
|
--model gemma-4-31B-it
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -528,15 +645,15 @@ working
|
|||||||
|
|
||||||
## Troubleshooting
|
## Troubleshooting
|
||||||
|
|
||||||
| Symptom | Likely cause | Check |
|
| Symptom | Likely cause | Check |
|
||||||
| ---------------------------------------- | ---------------------------------------------- | ------------------------------------------------------------ |
|
| ----------------------------- | ------------------------------------------------ | ---------------------------------------------------- |
|
||||||
| Frontend loads but API fails | Backend or Caddy proxy issue | `systemctl status podman-platform-api caddy` |
|
| Frontend loads but API fails | Backend or Caddy proxy issue | `systemctl status podman-platform-api caddy` |
|
||||||
| `/api/*` returns 502 | API not listening on `8787` | `ss -ltnp`, `curl http://127.0.0.1:8787/health` |
|
| `/api/*` returns 502 | API not listening on `8787` | `ss -ltnp`, `curl http://127.0.0.1:8787/health` |
|
||||||
| No screen observations | Vision agent not in LiveKit room | `journalctl -u podman-platform-agent -n 80` |
|
| No screen observations | Vision agent not in LiveKit room | `journalctl -u podman-platform-agent -n 80` |
|
||||||
| Live voice missing | Python conversation agent down | `journalctl -u podman-live-conversation-agent -n 80` |
|
| Live voice missing | Python conversation agent down | `journalctl -u podman-live-conversation-agent -n 80` |
|
||||||
| Memory empty | MongoDB unavailable or env missing | `curl /api/memory/stats`, backend logs |
|
| Memory empty | MongoDB unavailable or env missing | `curl /api/memory/stats`, backend logs |
|
||||||
| Hermes says tool auto-choice is disabled | vLLM missing tool flags | verify `--enable-auto-tool-choice --tool-call-parser gemma4` |
|
| Hermes tool calls fail | Modular/MAX endpoint does not accept tool schema | verify Hermes provider and `/v1/chat/completions` |
|
||||||
| `llm.alhinai.dev` returns 502 | Gemma vLLM still loading or tunnel target down | `curl /v1/models`, vLLM logs on Gemma host |
|
| `llm.alhinai.dev` returns 502 | Gemma server still loading or tunnel target down | `curl /v1/models`, MAX logs on Gemma host |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
Binary file not shown.
|
After Width: | Height: | Size: 40 KiB |
Reference in New Issue
Block a user