a4490ec80e
ci / rust (push) Successful in 2m26s
Governed compute for unified-memory AI hardware — machines where CPU and GPU share one pool and there is no separate VRAM allocation to bounce off. Over-commit that pool and the box thrashes and wedges, SSH and ping included, before the OOM killer gets a turn. Compute does not run inference. It supervises the servers that do: admission control against both a declared budget and what the machine actually has free, a 1 Hz watchdog that stops the newest model before thrash, Scenes activated as one transactional unit with rollback, process ownership bound to (boot_id, pid, start_time_ticks, pgid) so a reused PID can never be group-killed, a protocol-transparent gateway, an MCP server, and a read-only HTTP API for dashboards. Registry footprints in this release are measured on a live node rather than estimated. One binary, six direct dependencies. Apache-2.0. Generated by scripts/publish-compute.sh, which refuses to publish a tree it cannot prove clean.
22 lines
686 B
YAML
22 lines
686 B
YAML
apiVersion: lumbridge/v1
|
|
kind: EvalSuite
|
|
metadata:
|
|
name: performance
|
|
version: 1
|
|
description: "Repeated streamed requests measuring TTFT, client-observed prefill, and decode throughput."
|
|
tags: [performance, latency, throughput]
|
|
defaults:
|
|
max_tokens: 256
|
|
temperature: 0.0
|
|
repeat: 3
|
|
system: "Answer directly in plain text."
|
|
cases:
|
|
- id: short-prefill
|
|
category: latency
|
|
prompt: "Explain why unified-memory admission control prevents system thrashing. Give a detailed answer."
|
|
max_tokens: 256
|
|
- id: structured-decode
|
|
category: throughput
|
|
prompt: "Write twenty numbered, one-sentence operational checks for an AI inference server."
|
|
max_tokens: 384
|