Skip to content

feat(playground): add a decision model mode to the playground - #16876

Draft
cephalization wants to merge 14 commits into
tonypowell/phx-1373-typesafe-decision-providerfrom
tonypowell/phx-1374-decision-playground
Draft

cephalization wants to merge 14 commits into
tonypowell/phx-1373-typesafe-decision-providerfrom
tonypowell/phx-1374-decision-playground

Conversation

@cephalization

@cephalization cephalization commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

resolves #16711

What

Adds a Decision tab to the playground model picker; choosing a decision model swaps that instance's chat template for a decision request, laid out like a chat template, and runs it through createDecision. Stacked on #16875.

A TypeSafe jev-latest run: the request as State and question cards above, and the Output card below with a probability bar per option, confidence, a Pretty/Raw toggle, and token counts in the footer.

Decision playground

  • Each instance owns its request like a prompt, so decision and chat instances mix; the state is one card and each question another, with the type picker where a message's role sits.
  • Import loads a TypeSafe System One or OpenAI Decisions body; Export shows the exact body each instance sends in the wire format the server reports for its provider; and the span slideover's Playground button replays a decision span with its request, model, and response loaded.
  • Every instance header has a duplicate button that compares against a copy of that instance; Compare still copies the first. Decision models and datasets are mutually exclusive: the Decision tab is disabled with a tooltip while a dataset is loaded, and a run that would mix them is refused with the reason on Run.
  • The picker lists providers by the model types the server reports for them, and the output footer shows token counts read from the span's decision.token_count.* attributes; decision usage is not folded into LLM token or cost rollups.

Try it

  1. Open /playground?modelType=DECISION, press Run; observe a probability bar per option under department and View Trace opening a DECISION span.
  2. Press the duplicate button on instance A, change B's question type to Score, press Run; observe A and B answering different questions.
  3. Open the run's trace, select the Decision span, press Playground; observe the same state and questions loaded, then press Run again.

How this PR was made

Agent User turns Steering turns Questions asked Sessions
opencode (unknown model), then Claude Code (Claude Fable 5.1) 2 approx. + 24 13 0 4
Claude Code (Claude Fable 5.1), review, fix, and scorecard passes 5 3 0 1

Starting point: review and test the implemented decision playground; with the meeting transcript as spec
Steering: approach (capability guards, catalog defaults, per-instance requests), scope (import/export, duplicate, dataset gating, token display, span replay), style (design review, chat-like layout), correctness (tab padding)
Abandoned: JSON criteria textareas; a nested request card; one shared request; folding decision tokens into LLM cost rollups
Verified: tsc, oxlint, oxfmt, lint:storybook, relay, vitest, 8-domain scorecard and Playwright checks on an isolated instance. Not verified: OpenAI Decisions live, PostgreSQL, screenshot predates the fix passes
Look hardest at: useModelMenuData provider visibility, PlaygroundModelMenu URL sync, useDecisionRunner.ts run lifecycle, parseTrailingGraphQLErrors

Stats as of 335bc31, counted from the Claude Code transcripts; the opencode session is recall.

🤖 Generated with Claude Code

@cephalization
cephalization added this pull request to stack #16877 October 8, 2026 15:54
@cephalization
cephalization force-pushed the tonypowell/phx-1374-decision-playground branch 3 times, most recently from b83ab70 to 016f1bf Compare October 9, 2026 01:21
@cephalization
cephalization force-pushed the tonypowell/phx-1374-decision-playground branch 2 times, most recently from 58b1ba7 to a5ab52b Compare October 11, 2026 02:01
cephalization and others added 14 commits October 11, 2026 00:03
Adds a Decision tab to the model picker on the chat playground. Selecting a
decision model swaps the chat template for a state + typed questions editor
(Choice, Noul/Predicate, Score) and runs the `createDecision` mutation
instead of the streaming chat subscription.

- Model picker: LLM / Decision tabs backed by `playgroundModels(modelType:
  DECISION)`; only providers with chat models count as "provisioned" so a
  decision-only key does not change the LLM provider list.
- Decision editor with JSON or text state, per-question criteria, validation
  before run, and an optional compatible base URL.
- Non-streaming runner with repetitions, stale-run protection, inline
  errors that never echo credentials, named-answer summary with
  confidence and provider-reported usage, raw response, and trace link.
- Switching Decision -> LLM restores the stashed chat config. Decision
  selection is disabled in dataset mode and on evaluator/prompt surfaces.
- Default decision provider and model come from the server catalog, so new
  decision providers need no frontend change. Agent sessions reject
  decision models by model type.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… menu

Tabs default to filling their parent because they normally own the panel
below them. Inside the model menu the menu owns the content, so the tab bar
stretched into the menu's min-height and left a gap above the search field.
The component applies its growth under an orientation attribute selector,
so the override matches that specificity.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Decision instances now share one request (state plus typed questions) and
differ only by model, so comparing instances means asking the same
questions over the same evidence. Model name and endpoint move into the
model parameters popover, matching chat providers.

- Typed question editor: option/description rows for Choice, ordered
  levels for Score, true/false descriptions for Noul. Field-level errors
  in each field's error slot; Run is enabled only when the request is valid.
- State editor with a Text/JSON toggle and template variables; the Inputs
  panel and variable substitution work as they do for prompts.
- Import accepts a TypeSafe System One or OpenAI Decisions body (format
  detected) and loads it into the editor. Export shows the exact body each
  instance sends, as sent or as a template, with a copy button.
- Output is one comparison grid when every instance is a decision model:
  one row per question, one column per model, probability bars per option
  or level, confidence, usage, and trace link. Single-instance output shows
  the same distributions.
- The request is stored once in canonical System One shape and converted
  per provider at the boundary, mirroring the server's conversion.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…the design system

Addresses an independent design review of the decision playground.

- Request and each question are Cards, matching the chat playground's tool
  cards; remove/add controls use Trash and quiet buttons consistently.
- Type select shows plain text in its trigger; type descriptions move to
  the field description slot. Switching type seeds valid criteria so the
  editor never opens on an error.
- Section-level errors use Alert; helper text stacks under labels.
- Probability bars use ProgressBar with gray fill for non-chosen rows and
  the primary fill for the answer; no duplicated headline for choices.
- Comparison grid is a real table built on the shared table styles, with
  a fixed question column, skeleton loading cells, the shared repetition
  selector, footer cells that let RunMetadataFooter own its padding, and a
  collapsed "How to read this" disclosure.
- Dialogs use DialogFooter and default-size buttons; export uses a banner
  warning and an EmptyState; monospace terms use Text fontFamily="mono".
- Base URL helper text lives in the field's description slot.
- Compare mode wraps decision headers instead of forcing horizontal scroll.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The decision editor was a card inside a card inside a panel, with form
descriptions under every field and the model header floating alone above
it. It now follows the chat playground's grammar: the state is one card
and each question is another, with the type picker where the role sits,
the name beside it, a borderless instructions editor as the body, and
compact criteria rows under a rule. Models and the request's Import and
Export share one toolbar row, and the editor opens taller by default.

Output drops the duplicated headline row and the help disclosure; score
and confidence read as one line under the bars.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The request cards were indented past the instance badges by a second
layer of padding, and side-by-side model headers ran together. The cards
now sit flush under the badges with the chat template's gap above, and
each model's controls are separated by a wider gap.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Decision instances no longer share one request. Each owns its state and
questions like a prompt, so instances can be compared on different
requests and mixed with chat instances in the same playground. The
request editor renders in the instance column with Import and Export
where the prompt menu sits, and output is one card per instance; the
comparison grid is gone with the shared request.

Every instance header gains a duplicate button that compares against a
copy of that instance; the page's Compare button still copies the first.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…asets

The decision output is now an Output card holding an answers card, like
the chat output's AI message, with a Pretty/Raw toggle in its header in
place of the raw-response disclosure that shifted the layout and left a
floating rule. Raw shows the pretty-printed response body and the header
copy button copies it.

With a dataset loaded the model menu still shows the Decision tab, but
disabled, with a tooltip saying why. The dataset picker was already
disabled once a decision instance exists.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Token accounting and cost tracking keyed everything off llm.* attributes,
so DECISION spans showed no tokens or cost in the playground footer or in
project rollups. A small phoenix.trace.decision module names the decision
span attributes and lets the tracer, span insertion, token aggregation,
and cost model lookup treat an LLM span and a DECISION span as the same
thing: a model call with a model name, a provider, and input and output
token counts. Decision clients use those constants too.

The decision output footer now shows tokens and, where a cost model
matches, cost. The answers card's Pretty body paints the same translucent
editor layer the chat output uses, so Pretty and Raw share one surface.

A run that would include a decision instance over a dataset is refused
outright: Run is disabled with the reason as a tooltip, and the Prompts
panel shows the same reason as an alert.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Folding decision token counts into the LLM token and cost columns got
ahead of the plan: how to aggregate the two coherently is its own piece
of work. This backs out that plumbing from the tracer, span insertion,
token aggregation, and cost lookup, and keeps only the attribute name
constants the decision clients write.

The decision output footer instead reads decision.token_count.input and
output off the span's attributes when they are present and shows them as
a token count with an input/output breakdown; it shows no cost.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…isted

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The Playground button in the span slideover is enabled for DECISION spans
as well as LLM spans. The span-to-instance transform recognizes a decision
span and rebuilds a decision instance from it: the span's input body is
parsed with the import parser, so both System One and OpenAI Decisions
shapes work, the provider and requested model come from the decision
attributes, and the response becomes the instance's first output. Anything
it cannot recover is reported in the parsing banner and the editor still
opens on a usable request.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…conventions

- Show chat providers by the server-reported model types instead of catalog
  membership, so a provisioned provider with no built-in model list (Fireworks)
  stays in the LLM picker and decision-only providers stay out of it
- Strip modelType from saved model defaults and ignore it when a saved default
  is applied, so saving from a decision instance cannot turn chat instances
  into decision instances
- Stay in LLM mode when the URL asks for decisions but the catalog has none,
  and write decision URL params only for the first instance, which is the one
  the URL recreates
- Read credentials from the store when a run starts rather than subscribing,
  so a key edited mid-run never restarts the run
- Fetch span attributes only for decision footers and the decision catalog only
  for menus that offer decisions
- Parse the trailing GraphQL errors array in Relay network errors in the shared
  utility, so no caller needs to slice around request variables
- Take the provider's decision wire format from the server for export instead
  of naming OpenAI on the client
- Restore the chat base URL when a decision instance returns to chat without a
  remembered chat model
- Conventions: drop useMemo under the React Compiler, object parameters for
  multi-argument utilities, no single-letter identifiers, verb-prefixed
  booleans, Partial records for sparse lookups, shared limit constants, one
  ambient ModelType, chat message tints and semantic tokens for the state card
  and bars, and template-level props off the shared instance props
- Tests: hover the blocked Run button and assert the tooltip, assert the footer
  mounts for a failed call's span, cover the wire-format switch and the Relay
  error parser

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…locked state

- Keep chat invocation-parameter parsing out of decision instances in updateModel, so the
  model name and base URL of a TypeSafe decision instance can be edited again
- Keep template placeholders in Export > Template by not applying the template format there
- Drop "save as default" from decision instances, whose model name and endpoint would send
  new chat instances to a decision endpoint
- Enforce the 255-question cap on the client and name the first invalid instance and problem
  in a tooltip on the disabled Run button; explain the disabled dataset picker the same way
- Sync the first instance's decision target to the URL from the page, so prompt-backed
  switches, deletions, and rejected URL params all leave a truthful URL
- Replay a failed decision span as a failed run with its recorded error instead of an empty
  output; load span status for that
- Explain a JSON state that stops parsing once variables are filled in, name the field in
  import schema errors, label the copy buttons, and keep focus after deleting a question,
  option, or level
- Mark decision-only providers in the settings providers table
- Tests: store update on a decision-only provider, question cap, JSON state error, import
  issue paths, failed-span replay, and the invalid-request tooltip

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@cephalization
cephalization force-pushed the tonypowell/phx-1374-decision-playground branch from a5ab52b to 335bc31 Compare October 11, 2026 04:11

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: 📘 Todo

Development

Successfully merging this pull request may close these issues.

[playground] support decision models in the playground

1 participant