3.0 KiB
3.0 KiB
Positioning
Where Lumbridge Compute sits relative to the tools people already know. The short version: the neighbors solve different problems, and the closest one has exactly the gap Lumbridge Compute fills.
The landscape
| Tool | Axis | What it does | Gap for a unified-memory box |
|---|---|---|---|
| exo | scale out | Shards one big model across many boxes (Thunderbolt/RDMA, MLX, Apple-first) | Doesn't run many models on one box; not GB10/GB300-native |
| Ollama | single model | Dead-simple pull/run one LLM, huge library | No multi-service orchestration, no memory governor |
| llama-swap | swap in/out | On-demand load/unload, TTL idle-unload, co-run matrix, any OpenAI server | "No memory budgeting or OOM prevention" (their words), LLM-swapping not capability-stacks, no scheduling/sharing |
| LiteLLM | routing | One OpenAI endpoint in front of many backends | A proxy — doesn't manage lifecycle or memory |
The one-liner
exo runs one model across many boxes. Lumbridge Compute runs many models on one box — safely.
Orthogonal, not competing — you can run Lumbridge Compute on each node of an exo cluster.
Where Lumbridge Compute is unique
- The Governor. No one else does unified-memory admission control + an OOM-wedge watchdog. It's the pain specific to Grace-Blackwell shared memory, where over-commit wedges the whole box instead of failing a single allocation. This is the moat.
- Scenes are capability stacks, not model swaps. llama-swap picks which LLM answers. Scenes compose heterogeneous services together — ASR + LLM + TTS + image + music — as one activatable unit.
- Scheduled scenes — time-multiplex a box by hour (live assistant by day, image farm overnight). Nobody does this.
- Shareable scene registry — publish and pull scenes. The community flywheel.
- GB10 / GB300 / RTX-native — NVFP4,
flashinferMoE backend,sm_121a. exo is Apple-first; Lumbridge Compute is Grace-Blackwell-first.
Ideas worth borrowing (don't reinvent)
- From llama-swap: TTL idle-unload; an OpenAI-compatible gateway in front so apps hit one stable endpoint regardless of which scene is live; a web UI; Prometheus metrics; install via brew/binary/docker.
- From Ollama: one-command onboarding and a friendly
model pull/ls/rmUX + a library. - From LiteLLM: the router-in-front pattern — pairs with scenes (the
served_namealiases already point this way). - From exo: OpenAI + Anthropic + Ollama-compatible APIs; later, auto-discovery so Lumbridge Compute can coordinate scenes across two Sparks.
The model manager angle
Lumbridge Compute's model commands aren't just another Ollama. They're governor-aware and
scene-aware: Lumbridge Compute knows each model's footprint, warns before you pull something no
scene could ever hold, and shows which scenes a model belongs to. Ollama manages files;
Lumbridge Compute manages what can actually run together.