# Agents - Local Studio DLTL

Company: Sybil Solutions (https://www.sybilsolutions.ai/).

Compact instruction sheet for coding agents covering controllers, providers, runtimes, and Pi.

## Scope

- Controllers stay saved; switching is non-destructive.
- Provider keys live in controller config, not prompts.
- `provider/model` routes to that provider.
- Default model names hit the active backend.
- Pi sessions load selected skills and local tools.

## Hard rules

- Never use max_tokens.
- For vLLM/SGLang, never add --disable-cuda-graphs or --enforce-eager.
- Do not bypass SSH host-key verification.
- Keep keys in env, secure local files, or app settings.

## Controller

1. Verify GET /status, /gpus, /config, /v1/models
2. Local default: http://localhost:8080
3. Remote GPU boxes expose controller API, not raw inference ports
4. Settings → Connection; keep all saved controllers
5. Switch active target; confirm Settings → System

## Providers

OpenAI-compatible /v1 upstreams via POST /studio/providers. Route as `provider-id/model-name`.

## Runtimes

vLLM (CUDA), SGLang (structured), llama.cpp (GGUF), MLX (Apple Silicon). Launch through recipes/UI.

## Acceptance

Settings switches controllers. System shows runtime. /v1/chat/completions works locally and through one provider. /agent completes a turn. No secrets in artifacts.
