This commit is contained in:
@@ -0,0 +1,36 @@
|
||||
apiVersion: lumbridge/v1
|
||||
kind: Scene
|
||||
metadata:
|
||||
name: nemotron
|
||||
version: 1
|
||||
description: "The assistant stack with NVIDIA Nemotron 3.5 Lightning 30B-A3B in place of the Qwen MoE brain — the A/B candidate for replacing `brain`."
|
||||
author: karti
|
||||
tags: [assistant, voice, "downstream", candidate, nemotron]
|
||||
|
||||
# Deliberately the assistant stack with ONE substitution. `brain-nemotron` takes :8001
|
||||
# in place of `brain`, so the two are mutually exclusive by construction — and
|
||||
# that is also what the hardware wants: 68GB is already spoken for with brain
|
||||
# up, and two MoE engines would each spike during KV profiling.
|
||||
#
|
||||
# Keeping ears/voice/embed/ocr identical is the point. The only variable
|
||||
# between the primary scene and this scene is which brain answers :8001, so an eval run
|
||||
# against one is comparable to an eval run against the other.
|
||||
#
|
||||
# ORDER MATTERS, same rule as the primary scene: whatever holds :8001 comes up
|
||||
# first and alone. brain's KV-profiling spike is not bounded by
|
||||
# gpu-memory-utilization and this engine is unproven on the box, so it gets
|
||||
# the empty machine.
|
||||
models:
|
||||
- brain-nemotron
|
||||
- ears
|
||||
- voice
|
||||
- embed
|
||||
- ocr
|
||||
# The eye, added 2026-08-12. Cosmos3-Edge is BF16-only and cannot be
|
||||
# quantized, so it costs a real 12GB that the other slots do not.
|
||||
- eye
|
||||
|
||||
budget_gb: 100
|
||||
activation:
|
||||
order: listed
|
||||
wait_healthy: true
|
||||
Reference in New Issue
Block a user