2584ea908e
ci / rust (push) Successful in 2m27s
Governed compute for unified-memory AI hardware — machines where CPU and GPU share one pool and there is no separate VRAM allocation to bounce off. Over-commit that pool and the box thrashes and wedges, SSH and ping included, before the OOM killer gets a turn. Compute does not run inference. It supervises the servers that do: - Admission control against two ceilings — a declared budget, and what the machine actually has free. The refusal is the feature. - A 1 Hz watchdog on MemAvailable that stops the newest model before thrash, and defers to a Scene transition rather than racing it. - Scenes: named sets of models activated as one transactional unit, with pre-flight validation and rollback to the previously active Scene on failure. Scenes reference model ids, never weight paths or commands, so a Scene obtained from elsewhere cannot introduce code. - Process ownership bound to (boot_id, pid, start_time_ticks, pgid == pid), so a reused PID can never be group-killed. - A protocol-transparent TCP gateway, so clients keep one address while model runtimes move behind it. - An MCP server, so agents drive the node as tools. - A read-only HTTP API for dashboards: loopback by default, CORS off unless an origin is named, and it never mutates. One binary, six direct dependencies, no async runtime outside the MCP and API surfaces. Apache-2.0. Generated from the internal monorepo by scripts/publish-compute.sh, which refuses to publish a tree it cannot prove clean.
38 lines
1.8 KiB
Markdown
38 lines
1.8 KiB
Markdown
# Lumbridge Compute evaluations (`lumbridge/v1`)
|
|
|
|
Lumbridge Compute evaluates the model configuration that is actually serving: weights,
|
|
quantization, context, runtime, parsers, and speculative decoder. A model name
|
|
without its serving configuration is not a reproducible benchmark target.
|
|
|
|
```bash
|
|
lumbridge-compute eval ls
|
|
lumbridge-compute eval run smoke
|
|
lumbridge-compute eval run performance --model brain --repeat 5
|
|
lumbridge-compute eval run finance-core --base-url http://your-node:8001/v1
|
|
```
|
|
|
|
Suites live in `evals/*.eval.yaml`. Each case has a stable id, prompt, category,
|
|
generation limit, and deterministic assertions. Runs produce append-only JSON in
|
|
`eval-results/` with raw outputs and per-sample metrics.
|
|
|
|
## Metrics
|
|
|
|
- **TTFT**: wall time until the first streamed content or reasoning token.
|
|
- **Prefill tok/s (approximate)**: API-reported prompt tokens divided by TTFT.
|
|
This is client-observed and includes queueing/scheduling; server-native prefill
|
|
metrics should be added as a separate source rather than conflated with it.
|
|
- **Decode tok/s**: completion tokens divided by time after the first token.
|
|
- **Score**: share of samples satisfying every declared assertion.
|
|
|
|
Performance runs should include warmups in automation and record hardware, Lumbridge Compute
|
|
scene, runtime version, model revision, and cold/warm cache state. The v1 artifact
|
|
is deliberately local and portable; a future registry can ingest the same JSON.
|
|
|
|
## Eval-to-training bridge
|
|
|
|
Capability suites should graduate into environment packages containing a dataset,
|
|
harness, and reward function. That common contract can be adapted to Prime
|
|
Intellect `verifiers` for evaluation, synthetic-data generation, SFT, or RL with
|
|
`prime-rl`. Lumbridge Compute owns scene scheduling, memory admission, checkpoints, and process
|
|
lifecycle; the training framework owns optimization and distributed execution.
|