Files
compute/evals/performance.eval.yaml
T
Karti Tripathi 75aa63d737
ci / rust (push) Successful in 2m26s
Lumbridge Compute
Governed compute for unified-memory AI hardware — the machines where CPU and
GPU share one pool and there is no separate VRAM allocation to bounce off.
Over-commit that pool and the box thrashes and wedges, SSH and ping included,
before the OOM killer gets a turn.

Compute does not run inference. It supervises the servers that do:

- Admission control against two ceilings: a declared budget, and what the
  machine actually has free. The refusal is the feature.
- A 1 Hz watchdog on MemAvailable that stops the newest model before thrash,
  and defers to a Scene transition rather than racing it.
- Scenes: named sets of models activated as one transactional unit, with
  pre-flight validation and rollback to the previously active Scene on
  failure. Scenes reference model ids, never weight paths or commands, so a
  Scene obtained from elsewhere cannot introduce code.
- Process ownership bound to (boot_id, pid, start_time_ticks, pgid == pid),
  so a reused PID can never be group-killed.
- A protocol-transparent TCP gateway, so clients keep one address while model
  runtimes move behind it.
- An MCP server, so agents drive the node as tools rather than as a CLI.

One binary, no async runtime outside the MCP surface. Apache-2.0.

Generated from the internal monorepo by scripts/publish-compute.sh, which
refuses to publish a tree it cannot prove clean.
2026-08-03 16:47:06 -07:00

22 lines
686 B
YAML

apiVersion: lumbridge/v1
kind: EvalSuite
metadata:
name: performance
version: 1
description: "Repeated streamed requests measuring TTFT, client-observed prefill, and decode throughput."
tags: [performance, latency, throughput]
defaults:
max_tokens: 256
temperature: 0.0
repeat: 3
system: "Answer directly in plain text."
cases:
- id: short-prefill
category: latency
prompt: "Explain why unified-memory admission control prevents system thrashing. Give a detailed answer."
max_tokens: 256
- id: structured-decode
category: throughput
prompt: "Write twenty numbered, one-sentence operational checks for an AI inference server."
max_tokens: 384