Skip to main content

User Guide

A comprehensive guide to using Delstorm for multi-agent experimentation.

1. Overview

Delstorm is a multi-agent experimentation platform. You design experiments where AI agents — each with their own persona, backed by an LLM of your choice — engage in structured discourse across configurable rounds.

Each round has a condition prompt that sets the social scenario: a debate topic, a collaborative task, a negotiation, a simulated classroom, or any situation you can describe. Agents respond based on their personas and the conversation history. You control the interaction pattern, number of turns, which agents participate, and how long responses should be.

Use cases include:

  • Simulating academic debates grounded in real papers
  • Modeling stakeholder negotiations with distinct interests
  • Exploring how different ideological perspectives interact on a topic
  • Testing how agent personas, temperature, or round structure affect discourse outcomes
  • Running batch experiments to observe variance across identical setups
  • Generating structured multi-perspective analysis of documents or topics

2. Getting Started

To run your first experiment, follow these steps:

  1. 1.

    Create an account

    Click "Sign In" in the top right. If you don't have an account, click "Create one" to sign up with your email and password. If this address is new, you'll receive a confirmation email: click the link to activate your account, then sign in. If you already have an account, no email is sent, so sign in instead (and use "Forgot password?" on the sign-in page if you need a new password).

  2. 2.

    Check your providers

    Go to Settings. You should see at least one provider already configured (e.g., "deepseek" or "claude"). If not, you'll need to add one — see Setting Up Providers below.

  3. 3.

    Create an experiment

    Go to New Experiment. Give it a name, select your default provider (e.g., "deepseek"), enter the default model name (e.g., "deepseek-chat"), and add at least one agent and one round.

  4. 4.

    Launch

    Click "Launch Experiment" and watch the agents respond in real-time. Results appear as they're generated.

Forgot your password? Use the "Forgot password?" link on the sign-in page: enter your account email and open the link in the message that arrives. It takes you to a page where you confirm your account email and set a new password, and you land back on Home already signed in. Links expire after a short while and each works once; if yours has expired, the page says so and offers to send a fresh one. For privacy, the request page shows the same confirmation whether or not an account exists for the address you typed, so if no email arrives, check your spam folder and double-check the address. This is also how you change your password: there is no separate password field in Settings.

Quick test: Try 2 agents, 1 round, "Short" response length. That is a small, fast test, so it lets you verify everything works before building a larger experiment.

Your Home Page

Home is where signing in lands you: a greeting, a focus card of the work to continue, your classrooms, and then your lists.

  • The focus card offers the design you saved most recently, with an Open in builder button. Before you have saved any design it offers your most recent run instead. Students always see the run card, since their starting point is Classrooms rather than the builder.
  • Your designs is a grid of your saved designs, most recently saved first, each showing when you last saved it and an Open in builder link that loads it into the Experiment Builder ready to edit and run. It shows the first 12 with a "Show all N designs" button underneath for the rest.
  • Recent Experiments lists your runs, newest first, each card carrying the run's name, its state, and the date it was launched. It shows the first 12 and puts a "Show all N runs" button underneath for the rest, so a long history doesn't bury everything below it.
  • Recent Batches lists batch runs and how many of their runs have finished.

Finding an older run

The search box above Recent Experiments filters the list as you type, matching a run's name, its state (completed, running, error), or its date. Dates match either way you type them: as the card prints it (Mar 7) or as a plain calendar date (2019-03-07). Search reads every run you have, not only the 12 on screen, so an old run turns up without expanding the list first. If nothing matches, the list tells you so; clear the box to go back to the normal view.

Appearance: System, Light & Dark

Delstorm’s appearance control changes presentation only. It does not alter your account, permissions, experiments, assignments, or saved work.

  • System follows your device’s current light or dark appearance and updates if the device setting changes.
  • Light uses the warm paper interface.
  • Dark uses the neutral charcoal interface.

When signed in, open the appearance menu in the top navigation or use the Appearance section in Settings. On the sign-in and create-account pages, use the compact appearance button in the top bar.

Your selection is stored only in this browser and remains after you sign out. It is not synchronized to your account or shared with other users.

3. Setting Up Providers

A provider is an LLM backend that powers your agents. Providers are configured in Settings.

Shared vs. Personal Providers

Your administrator may have pre-configured shared providers (like Claude) that are available to all users automatically. These appear in your provider list without any setup. You can also add your own personal providers — these are visible only to you.

Provider Types

Anthropic (Claude)

For Claude models by Anthropic.

  • Backend: Select "Anthropic"
  • Base URL: Leave blank (the SDK handles this automatically)
  • API Key: Your Anthropic API key (starts with sk-ant-). Get one at console.anthropic.com.
  • Default Model: claude-sonnet-4-6 (recommended) or any Claude model ID

DeepSeek

A shared deepseek provider is available to all users — no setup needed. To add your own DeepSeek provider:

  • Backend: Select "OpenAI Compatible"
  • Base URL: https://api.deepseek.com/v1
  • API Key: Your DeepSeek API key. Get one at platform.deepseek.com.
  • Default Model: deepseek-chat (or deepseek-reasoner for reasoning tasks)

OpenAI Compatible

Works with any server that exposes the OpenAI-style /v1/chat/completions endpoint. This includes Ollama, vLLM, LMStudio, RunPod, OpenAI itself, and many others.

  • Backend: Select "OpenAI Compatible"
  • Base URL: The server's API URL ending in /v1 (e.g., http://localhost:11434/v1 for Ollama, or a RunPod proxy URL)
  • API Key: Whatever the server requires. For Ollama, any string works (e.g., "ollama"). For OpenAI, your OpenAI API key.
  • Default Model: The model name as the server knows it (e.g., gemma2:latest for Ollama, gpt-4o for OpenAI)

Per-Agent Provider Override

By default, all agents in an experiment use the experiment's default provider. However, you can override this on individual agents — for example, to have one agent powered by Claude and another by GPT-4. Set this in the "Provider Override" dropdown on each agent card.

4. Building an Experiment

The Experiment Builder is a three-pane workspace: your rounds (plus an "⚙ Experiment setup" entry) list in the left rail, your agents list in the right rail, and whichever item you pick opens as a full editor in the center. On smaller screens the rails hide, every section simply stacks down the page, and the "+ Add Agent" / "+ Add Round" buttons appear inline beneath their lists.

  1. 1.

    ⚙ Experiment setup (the default view)

    Name your experiment, select the default provider (e.g., "deepseek"), and enter the default model ID (e.g., "deepseek-chat"). The model ID must match what the provider expects. This view also holds Document Context (optional): upload a document to extract themes and generate grounded agent personas — see Document-Grounded Personas.

  2. 2.

    Agents (right rail)

    Define the participants. Click "+ Add Agent" at the bottom of the right rail; the new agent's editor opens in the center. Click any agent in the rail to edit it. See Agents & Personas.

  3. 3.

    Rounds (left rail)

    Define the sequence of interactions. Click "+ Add Round" at the bottom of the left rail; reorder rounds with the ↑/↓ buttons on each rail entry. See Rounds, Conditions & Settings.

  4. 4.

    Bottom bar: Launch / Save / Batch + drawers

    "Launch Experiment" runs it once. "Save Design" stores the config for reuse, updating the design you loaded rather than duplicating it; "Save as new" beside it always creates a separate copy. "Batch Run" runs it multiple times. Two panels slide in from the right: Designs (load a saved design, start from a built-in template, import/export/delete — templates that bundle documents are first imported into your saved designs so their documents land in your library) and Data extraction (see Data extraction; a green dot on the button means extraction is configured).

If Launch or Save flags a problem, the builder jumps to the item that needs fixing and highlights the field in red.

5. Agents & Personas

Agents are the participants in your experiment. Each agent is defined by these fields:

Name / ID

A unique identifier shown in the results (e.g., "Left_Coalition_Theorist", "Budget_Hawk"). Use descriptive names — they help you read the results and they're included in the agent's prompt context.

Specialty

The agent's area of expertise (e.g., "Progressive policy coherence", "Fiscal conservatism"). This is included in the agent's system prompt to focus its responses.

Persona

The most important field. This is the system prompt that defines who the agent is. It should describe: what the agent believes, what evidence they draw on, how they communicate, and what perspective they bring. The richer and more specific this is, the more distinctive and grounded the agent's responses will be.

Example of a weak persona: "A political analyst."

Example of a strong persona: "You are a political analyst specializing in left-wing party coherence. You argue that egalitarianism functions as a structural unifying principle across wealth redistribution, social morality, and immigration. You draw on Cochrane's analysis of Benoit and Laver's 22-country expert survey data, which shows 100% of far-left economic parties also hold left-wing positions on social and immigration dimensions. You are confrontational toward claims of left-right symmetry."

Temperature

Controls response variability. Range: 0.0 to 1.5.

  • Low (0.1-0.3): More deterministic, focused, predictable. Good for analytical or conservative agents.
  • Medium (0.4-0.6): Balanced. The default (0.5) works well for most cases.
  • High (0.7-1.0): More creative, varied, surprising. Good for innovative or provocative agents.

Tip: Pairing agents with different temperatures (e.g., 0.3 and 0.8) produces more interesting discourse than uniform temperatures.

Memory Window

How many past rounds the agent remembers. Default is 0 (all rounds). Set to a number (e.g., 2) to limit memory to the last N rounds. This is useful for simulating bounded memory or preventing context from growing too large. The round's Context Window still governs what's available in the first place — an agent's memory limits within the rounds the round offers, so it can never see more than the round does.

Memory Overflow

What happens to memories beyond the window. "Summary" compresses older rounds into a summary. "Forget" drops them entirely. This is an experimental variable — agents with different memory settings can behave very differently over many rounds.

Agent Type, Model Override, Provider Override

"Agent Type" is a legacy label (default "Custom") kept so older designs keep working. It inserts a labeled line into the agent's system prompt, the same way Specialty does, so it is not inert; it is largely redundant with Specialty, which is the field to write in instead, and it is never offered to students as an editable field. "Model Override" lets you use a different model than the experiment default. "Provider Override" lets this agent use a different LLM provider entirely — e.g., one agent on Claude, another on GPT-4.

Personal Document

A document from your Library that this agent carries with them across every round they participate in. Unlike a round-level reference document (shared by all agents in that round), a personal document is private to this agent and persists for the whole experiment. Use it when one agent needs expertise the others don't have — e.g., a "Methodologist" agent with a statistics textbook, or a "Historian" agent with a primary source.

Document Depth (agent)

Same three modes as round-level documents: Summary (extraction metadata only), Relevant Excerpts (default — paragraphs matching the current round's topic), or Full Text. Excerpts re-query against each round's condition prompt, so the same personal document surfaces different passages in different rounds.

How strongly the persona is held

Agents hold their persona on their own. You do not need to repeat "stay in character" or restate persona-specific instructions inside the round's condition prompt to be heard: every agent turn is told to answer in its persona's own voice, and is reminded of that persona again right before it speaks. Write the condition prompt for the situation and let each agent's Persona field carry who they are.

The moderator is deliberately different. It narrates and synthesizes about the participants rather than speaking as one of them, so it keeps a neutral voice. The same is true of the behind-the-scenes calls that decide whether a round is finished and which round runs next.

If you are running a research sequence, note that this behaviour changed on 9 August 2026. Transcripts produced before that date and transcripts produced after it are not directly comparable: the same design will hold its personas more strongly now. Re-run your baseline if a comparison has to hold across that date.

Prompt Library (instructors)

The Prompt Library is a built-in collection of course prompts you can reuse instead of retyping. It ships with the Classical Social Theory corpus — 91 prompts from SOCB42 and SOCB43 covering Smith, Marx, Wollstonecraft, Tocqueville, and Weber. It is instructor-only: students never see it.

Filter by theory, by assignment, or by prompt kind, and search across both titles and prompt text. There are four kinds:

  • Agent — a persona, inserted into an agent's Persona field.
  • Round — scene-setting instructions, inserted into a round's Condition Prompt.
  • Moderator — end-of-round instructions, inserted into a round's Completion criteria.
  • Transition — describes moving from one round to another based on what happened, inserted into a round's Transitions section as its Decision guidance.

In the Experiment Builder, open Prompt library in the right-hand agents rail. Pick a prompt, choose the agent or round to receive it, and click Insert. Insert appends to the chosen field rather than replacing it, so anything you have already written is kept.

Every round remains available as a target for a Moderator prompt. If the chosen round currently ends after its set turns, the action says Enable moderator ending + insert: it changes the existing Round ends control to When the moderator decides, reveals Completion criteria, and inserts there. Turns remains the hard cap, and you still choose the Moderator before launch. Nothing is changed until you click that explicit action.

In the assignment Studio the same library sits under Prompt library. While the design editor drawer is closed it stays in Copy mode, so you copy a prompt for reference. Open the design editor to edit an agent or round and the same library switches to Insert mode: it inserts into the drawer card you have focused and appends to that field, exactly as it does in the Builder. On a wide screen the library stays usable beside the open drawer; on a narrow screen it rides the drawer's own Prompt library tab. Closing the drawer returns it to copy mode. Transition prompts insert into the round's Transitions section in either mode: because guidance alone does not make a design branch, the insert also reveals that section and starts one empty condition row for you to fill in.

Some prompts are labelled Used by: Alpha, Beta, Gamma. That means one prompt is shared across several agents or rounds in the original course material. A visible source note preserves and explains a known issue in the source material before you reuse that record; it does not silently rewrite or remove the source prompt. The shipped corpora themselves stay read-only: you cannot add to or edit the course prompts. You can, however, save your own agents alongside them. See Agent library.

Use the Source filter to separate the two. It lists My agent library first, then each course corpus. The count line below the filters reads "12 of 94 prompts and saved agents" once you have saved anything, so the second number always means everything this browser can reach.

Agent library (instructors)

Your agent library saves a whole agent so you can use it again in another design. It is instructor-only, and it is private to your account: nothing you save is shared with students or with other instructors. Sharing a library, to a classroom or to a colleague, is not part of this version.

Saving. In the Experiment Builder, each agent card has a Save to my library button in its header row. The saved entry keeps everything on that card: the name and ID, specialty, persona, agent type, model, provider, temperature, memory window and overflow, and any attached document reference. Saving never interrupts what you are building. It changes nothing on the card, and if it fails you get a notice you can dismiss and carry on.

Names. An entry is named after the agent's ID by default. If you already have an entry with that name, Delstorm asks whether to replace the saved one or give this one a different name. It never overwrites silently, and the comparison ignores capitalisation and surrounding spaces, so "The Preacher" and "the preacher" count as the same name.

Browsing and inserting. Saved agents appear in the same Prompt Library browser as the course corpora, under the source My agent library. Inserting one adds a new agent card to your design with every field filled in, rather than pasting text into a card you already have. In the assignment Studio, saved agents are visible and you can copy their persona text, but the whole-card insert lives in the Builder, which is where agents are authored.

Routing is re-checked when you insert. If the saved provider is not set up in your account, Delstorm clears the provider and the model together, tells you which provider was dropped, and the new card uses your session's default provider and model. It never does that silently. A saved document reference that is no longer in your document library is dropped the same visible way. Model names are free text everywhere in Delstorm, so a model that carries over is exactly as checked as one you type in yourself.

Size limits: a name is up to 100 characters, and a whole saved agent has to fit in 64 KB, which is roughly thirty times the longest persona in the shipped course material. If either is exceeded the save is refused and says so; nothing is silently shortened.

Importing an Assignment Builder draft (instructors)

If you built an assignment with the Assignment Builder GPT, you can bring its draft into Delstorm at Import assignment draft. Paste the AssignmentDraftV1 JSON the GPT produced, or upload it as a file.

The quickest way in is from the classroom itself: open a classroom you instruct, and under New assignment choose Import a draft from the Assignment Builder. Arriving that way carries the classroom with you, so you do not have to pick it again later.

Delstorm checks the draft on its own server against the exact schema the GPT was configured with, then shows you what it would create and anything blocking it. Nothing is created by that check. Saving stores the draft in your private intake inbox — it does not create an assignment.

The intake inbox (Saved drafts, also linked from each classroom) lists every draft you saved. Opening one re-checks it fresh and shows the full record: instructions, agents, rounds, what students will do, the builder's sources and assumptions, and any blockers. Submitting a revised draft with a higher revision number updates the same intake and keeps the older revisions as history; submitting the same draft again is recognized and stored only once. When a draft is not ready, Copy feedback for the builder gives you a structured note to paste back into the GPT.

When a draft is ready, choose one of your classrooms on the review page and create the assignment. It is always created unpublished. You land in Studio to review and edit it, and publishing stays a separate deliberate step — importing a draft never exposes anything to students. Each intake can create one assignment; deleting an intake never deletes an assignment already created from it.

Two things worth knowing. Delstorm re-checks the stored draft when you create the assignment, not only when you saved it, so nothing stored is trusted stale. And uploaded source files do not travel: the draft names the documents it used, but the DOCX, PDF, or notes themselves stay in ChatGPT. If a design points at a document you have not uploaded to your Delstorm Document Library, creation is blocked rather than quietly dropping it.

A draft the GPT marked as blocked or awaiting your decision can still be saved and reviewed, but offers no create button until you resolve it with the GPT and submit a new revision. A draft still being interpreted cannot be saved at all. The GPT cannot write into Delstorm by itself — this manual import is the supported path.

Building an assignment with the assistant (instructors)

The assignment assistant builds an assignment with you by conversation, at Build an assignment with the assistant. The way in is from a classroom you instruct: under New assignment, choose Build an assignment with the assistant, which carries that classroom along for you. It is a tool for instructors. Students do not use it and cannot open it.

Describe the assignment you want in your own words and add your course material: a PDF or a text file, and slides work best exported to PDF. The assistant reads what you upload, asks what it still needs, and drafts the assignment with you. Ordinary questions get short answers; when the plan looks right, ask it for the draft.

From there the draft rides exactly the same path as an imported one. Check this draft validates it on the Delstorm server, the same check the import page runs, and shows you what it would create along with anything blocking it. Nothing is created by that check. Save to my drafts inbox stores it in your private intake inbox, the very same inbox manual imports land in. You create the assignment later from the review page, and it always arrives unpublished.

What it does not do. The assistant creates new assignment drafts only: it does not edit assignments that already exist, and it cannot publish anything or create the assignment for you. If the reply gets cut off before the draft finishes, it says so and offers nothing to save, rather than handing you an older draft that looks current.

Two names, two different things. This assistant lives inside Delstorm. The Assignment Builder GPT is a separate tool that runs in ChatGPT, and its drafts come over by hand on the import page. They both end at the same inbox, so use whichever suits you.

Your conversation stays in your own browser so you can come back to it, and it is cleared when you sign out. Uploaded material is read for the conversation and is not added to your Document Library; if a design needs a document at run time, upload it there as usual.

6. Document-Grounded Personas

You can upload a document (PDF, TXT, or MD) to automatically extract themes, positions, and stakeholders, then generate agent personas grounded in the document's actual content.

Step-by-step:

  1. 1.

    In the Experiment Builder, go to the "Document Context" section. Upload your file and select which provider to use for analysis (Claude is recommended for best extraction quality).

  2. 2.

    Click "Analyze Document". The LLM reads the document and extracts: title, summary, key themes, distinct positions/stances, stakeholders, key arguments, and domain terminology. This takes 10-30 seconds.

  3. 3.

    Review the extraction results. The themes, positions, and stakeholders should give you ideas for what agents to create.

  4. 4.

    Add agents and give them relevant names and specialties (e.g., if the document is about left-right politics, create "Left_Coalition_Theorist" with specialty "Progressive policy coherence").

  5. 5.

    On each agent card, click "Generate Persona from Document". The LLM creates a detailed persona for that specific agent, using the document's themes, arguments, data points, and terminology. The persona is tailored to the agent's name and specialty — not a generic summary.

  6. 6.

    Review and edit the generated personas. The generation is a starting point — you can refine it, add specific instructions, or combine it with your own text.

Important notes:

  • Document analysis is ephemeral — it's not saved with the experiment. The generated personas ARE saved as part of the agent configuration.
  • The more specific the agent name and specialty, the more distinctive the generated persona.
  • Supported file types: PDF, TXT, MD (max 10MB).
  • For academic papers, the extraction identifies the thesis, counterarguments, methodology, and key findings.
  • For narratives or stories, characters become stakeholders and their motivations become positions.

7. Document Library

The Document Library provides persistent storage for documents you want to reuse across experiments. Unlike the ephemeral document upload in the experiment builder (which is lost after analysis), library documents are permanently stored in your account.

Uploading Documents

Go to Library and upload a PDF, TXT, or MD file. The document is automatically analyzed by the LLM — themes, positions, stakeholders, and key arguments are extracted and stored alongside the full text. Supported files up to 10MB.

Using Documents in Experiments

Library documents can be attached at two levels:

Round-level — "Reference Document"

Each round card has a "Reference Document" dropdown. The selected document is shared by all agents in that round. Use this when the whole group should discuss the same text (e.g., a paper everyone is critiquing together). Different rounds can reference different documents — good for walking through a paper section-by-section.

Agent-level — "Personal Document"

Each agent card has a "Personal Document" dropdown. The selected document stays with that agent across every round, and it's private — other agents don't see it. Use this to give different agents different expertise (a Methodologist with a stats textbook, a Historian with a primary source, a Critic with a rival paper).

Both at once

An agent can have a personal document and participate in a round with a reference document — the two are injected as separate sections in the prompt, so there's no conflict.

Document Depth

For both round-level and agent-level documents, the "Document Depth" setting controls how much of the document is injected into the prompt:

Summary

Agents receive only the extraction metadata — themes, positions, and key arguments. Fast and cheap. Good for general grounding when agents just need to know what the paper argues.

Relevant Excerpts (recommended)

The system uses keyword matching against the round's condition prompt to pull the most relevant paragraphs from the actual document text. Agents get real text — specific data points, quotes, tables — but only the parts relevant to this round's topic. The default and best balance of cost and depth.

Full Text

The entire document is injected into agents' prompts. Comprehensive but uses significantly more tokens. Use when agents genuinely need access to every part of the document.

Example: Section-by-Section Analysis

Upload an academic paper to your library, then design rounds that walk through it:

Round 1: "Discuss the theoretical framework." (Document: paper, Depth: relevant excerpts)

Round 2: "Critique the methodology." (Document: paper, Depth: relevant excerpts)

Round 3: "Analyze the findings and data." (Document: paper, Depth: relevant excerpts)

Round 4: "Synthesize your conclusions." (No document — agents discuss from memory)

Each round's condition prompt guides the excerpt extraction, so agents see the parts of the paper most relevant to that round's focus.

8. Rounds, Conditions & Settings

Each round represents a phase of the experiment. Rounds are independent by default — you create continuity through how you write your condition prompts.

The Condition Prompt

This is the most important part of a round. It sets the social scenario that all participating agents operate within. Think of it as setting the stage for a scene.

Example — Academic conference:

"You are at an academic conference panel. Present your interpretation of the asymmetry finding and what explains it."

Example — Escalating challenge:

"A critic challenges your methodology. Defend your position and address the critique directly."

Example — Resource scarcity:

"Resources have become scarce. A drought has hit the village. Renegotiate your trade agreements."

Round Settings

On the round card these controls are grouped into three sections: Turn-Taking (speaking order, first speaker, turns, moderator, when the round ends), Memory & Context (context depth, context window, overflow, private self-reflections), and Output & Participants (response length, participating agents, reference document and depth) — so it's clear at a glance which settings shape how the group takes turns versus what each agent remembers. Each section starts collapsed — click a section heading to expand it (and again to hide it) — so you can focus on one group at a time and see at a glance that the settings are organised into expandable groups.

Round name (optional)

A short human-readable label for the round — e.g. "Opening positions", "Challenge round", "Synthesis". In a flat design it is a label only: it never changes how agents behave, but it makes multi-round experiments easier to scan in the builder, while a run is streaming, and in the saved transcript. In a branching design the name is load-bearing, because transitions and "After this round" reference rounds by name; the builder keeps those references in sync when you rename a round, and a round that is referenced needs a name. Leave it blank and rounds are shown as "Round 1", "Round 2", … exactly as before.

Interaction Pattern

How agents interact within the round. See Interaction Patterns for details on all five options. Note that two of them — Chain and Blind Parallel — are really about what each agent sees (a memory effect), while Free Discussion, Debate, and Custom set how the group takes turns.

Turns per Round

How many back-and-forth exchanges within this round. Default is 1 (each agent speaks once). Set to 3 and agents will go back and forth three times, building on what was said before. How much each agent sees of the others depends on the interaction pattern. More turns = deeper conversation within a single round.

Response Length

Controls how long each agent's responses should be:

  • Dynamic: agents match their response length to the conversational moment. A brief interjection, a question, or a one-line reply counts as a full response when that is what the exchange calls for. Dynamic adds no length cap of its own, though your provider's and model's own limits still apply.
  • Short: 3-5 sentences. Good for quick iteration, testing, and batch runs.
  • Medium: 1-2 short paragraphs. Room to develop a point beyond a few sentences.
  • Thorough: Detailed responses with evidence and elaboration.

This is set per-round, so you can have a short opening round followed by a thorough deep-dive round.

Participating Agents

Which agents are active in this round. Leave empty for all agents. Enter comma-separated agent IDs to restrict participation (e.g., "Agent_1, Agent_3"). Click "Fill from agents" to auto-populate from your current agent list. This lets you design rounds where only certain agents interact.

Context Window

How many past rounds of shared discourse are visible to agents. Default is 0 (all rounds visible). Set to 2 to only show the last 2 rounds. This is an experimental variable — limiting context simulates bounded attention.

Context Overflow

What happens to round history beyond the context window. "Summary" compresses it. "Truncate" drops older rounds. "Forget" removes them entirely.

Context Depth

How much of each earlier message agents actually see. Full (untruncated) — the default — passes the complete text, so agents can react to specific details, quotes, or numbers. Summary (300-char cap) trims every earlier response to its first 300 characters, keeping prompts small and cheap. This is separate from Context Window: the window decides which rounds are in view, the depth decides how much of each one.

Moderator

An optional dedicated agent that guides this round instead of participating in it. See Moderators.

Round ends

Whether the round stops after the set turns or when the moderator decides it's complete. See Round Endings.

Private self-reflections

When ticked, each agent writes a private note after the round that only it can see later. See Private Self-Reflections.

9. Interaction Patterns

Each round uses one of five interaction patterns that define how agents communicate:

Free Discussion

All agents respond to the condition prompt in your agent-list order, and each agent sees what the earlier speakers said — both within the same turn (a later speaker reacts to whoever already spoke this turn) and across previous turns in a multi-turn round. This is the default and most common pattern — it simulates an open conversation.

Best for: general discussion, brainstorming, collaborative analysis

Debate

Agents are split into teams that argue opposing positions. You define team assignments as JSON — for example: {"for_regulation": ["Agent_1", "Agent_2"], "against_regulation": ["Agent_3"]}. Each team sees the other team's arguments.

Best for: adversarial discourse, policy debates, exploring opposing viewpoints

Chain

Sequential — each agent sees only the immediately previous agent's response, not the full discussion. You specify the order as comma-separated IDs (e.g., "Agent_1, Agent_2, Agent_3"). Ideas build incrementally, like a game of telephone.

Best for: iterative refinement, building on ideas, examining how information transforms

Blind Parallel

All agents respond independently with no visibility of each other's responses — even from previous turns. Each agent only sees the condition prompt.

Best for: comparing uninfluenced perspectives, measuring baseline positions before discussion

Custom

Define your own interaction rules in natural language using the "Custom Instructions" field. These instructions are injected into each agent's prompt alongside the condition. Use this for scenarios that don't fit the presets — e.g., "Only respond to the agent who spoke before you" or "You can only communicate in questions."

Best for: novel interaction designs, creative constraints, specialized scenarios

Tip: You can use different patterns across rounds in the same experiment. Start with Blind Parallel (to capture uninfluenced positions), then Free Discussion (to let agents interact), then Debate (to force a conclusion).

10. Speaking Order

On a Free Discussion (or Custom) round, the "Speaking order" control on the round card decides who speaks next. The other patterns — Debate, Chain, Blind Parallel — have their own built-in orders and ignore this control, so it only appears for Free Discussion and Custom.

Sequential (default — one speaker at a time)

Agents speak one at a time in your agent-list order, and each speaker sees the entire conversation so far — so a later speaker genuinely reacts to what was just said instead of everyone answering the prompt at once. This is the default for Free Discussion and Custom rounds. Use the optional First speaker control to pick which agent opens; leave it blank to lead with the first agent in your list.

Round-robin (all at once)

The older behaviour where every agent answers once per turn. Pick this to opt back into all-at-once rounds — it reproduces a pre-2026 round exactly (its stored value is default).

Random

A random pick from whoever has spoken least so far, so the round keeps participation balanced without a fixed rotation. The optional First speaker control still applies — set it to name which agent opens before the random picks begin.

Handoff (the current speaker decides)

The agent who just spoke names the next speaker. If it names no one or an unknown agent, the round falls back to the least-spoken agent so it never stalls. The hand-off marker itself is stripped from the transcript — you only see the conversation.

Moderator (the moderator decides)

The round's moderator chooses who speaks next and says so out loud: before each turn it posts a short 🎙 message naming the next speaker and why, then that agent speaks. Tell it how to choose in the "How the moderator picks" box that appears (e.g. "call on whoever was challenged directly; favour quieter voices") — leave it blank to use sensible defaults (honour direct requests, balance participation, favour quieter voices). This mode requires the round to have a moderator assigned; on an empty or unrecognised pick it falls back to the design order.

Turns become a per-speaker budget

For every order except Round-robin — so including Sequential — the "Turns per round" field relabels to "Turns per speaker". The number is now a per-agent budget, and the round's hard cap is that number times the number of participants. Because the moderator or the hand-off can re-call or skip an agent, a speaker may take several turns while another takes none — the cap is the ceiling, not a promise that everyone speaks the same amount. The round still ends early if its moderator decides it's complete.

Each utterance streams into its own bubble in the live and completed views, so a re-called speaker reads as a fresh turn rather than appended text.

11. Moderators

A moderator is a dedicated agent that guides a round without taking part in the debate. It frames the scene, comments between turns, and wraps the round up — like a panel chair or a narrator.

Assigning one

On the round card, use the "Moderator (optional)" dropdown. It lists every agent you've defined and reads — none — until you pick one, so add your agents first. The same agent can moderate one round and participate normally in another.

What the moderator does

  • It doesn't participate. The moderator is removed from the round's participant list, so it never argues a position or takes a turn.
  • It speaks to guide the round: an opening message and a closing message always, plus comments between turns on some speaking orders, and, when it drives the speaking order itself, a short message before each turn naming who speaks next and why.
  • It sees everything. The moderator reads the full, untruncated round so far (it ignores the round's Context Depth setting).
  • Its comments steer the room. Moderator messages are passed to the participants on their next turn under a "Moderator guidance" heading — so a nudge or a question shapes what comes next.
  • Blind Parallel is the exception. In a Blind Parallel round agents see only the condition prompt, so moderator guidance doesn't reach them (the moderator still speaks and its messages are still recorded).

Moderator messages render as their own 🎙 moderator bubbles in the live and completed views, and they're included in the exported transcript.

The default narration, and the end-of-round summary

Assigning a moderator to a round assigns a fixed script, not a personality that may or may not speak. Every moderated round gets an opening message that sets the scene and poses the central question, and a closing message that synthesizes the discussion and concludes. In between, the moderator may interject, depending on the round's speaking order: some orders give it a short comment between turns, and when the moderator drives the speaking order itself it speaks before each turn to name who goes next and why. Other orders produce no between-turn narration at all. The close always happens, and it is an extra utterance: it does not count as one of the round's turns, so the closing synthesis is never paid for out of the participants' turn budget.

That closing synthesis is the platform's instruction to the moderator, not something the agent decided to do, so the end-of-round summary many designs show is expected behaviour rather than drift. When the round's ending is set to "When the moderator decides" and you have written completion criteria, the close is also told to say what it observed that satisfied those completion criteria, or which criteria were still unmet if the round hit its turn cap instead. The criteria box is optional: leave it empty and the close is a plain synthesis with no met-or-unmet verdict. If the round branches, the transition announcement ("Moving to ...") is recorded as one more moderator message after the close, so the reason the run took that path sits in the transcript beside the summary.

This is not configurable today: there is no setting that turns the closing summary off or rewrites it. If you want a round with no narration and no summary, do not assign a moderator to that round; the round then runs on its turn count with participant turns only. To shape the summary rather than remove it, write the moderator's persona to say how it should close.

Good to know

  • Keep the moderator's persona neutral and procedural — a facilitator or narrator, distinct from the substantive characters. Reusing a character as the moderator muddies its role.
  • A moderator can also decide when the round ends — see Round Endings.
  • A moderator can also decide who speaks next — set the round's Speaking order to "Moderator (moderator decides)".

12. Round Endings

Every round card has a "Round ends" control that decides when the round stops. There are two options.

After the set turns

The default. The round runs exactly the number of turns you set, then stops — a classic fixed-length round.

When the moderator decides

The round's moderator ends the round once a goal is reached. After each turn it privately checks whether the round is complete against the "Completion criteria" you write in the box that appears, and stops as soon as they're met. This option requires the round to have a moderator assigned.

Turns is always the cap

Even with "When the moderator decides", the turn count you set stays the hard cap — the round can end early on the moderator's call, but never runs longer than the set turns. So set Turns to the most you'd ever want, and let the moderator stop sooner.

Completion criteria are private while the round runs

The completion criteria are moderator-private during the round: the behind-the-scenes end-check is the only place they appear while agents are still speaking, so you can describe the outcome you're waiting for — even in the theory's own terms — without steering the discussion. The end-check itself is a separate, quiet call that never appears in the transcript. When the round closes, and only if you filled in the completion criteria, the moderator's closing message announces what it detected: which criteria it observed being met, or, if the turn cap ended the round first, which were still unmet. By then every participant turn is already spoken (and private self-reflections never see moderator messages), so the announcement explains the ending without influencing it. The box is optional: leave it empty and the round still ends on the moderator's own judgement, but the close is a plain synthesis with no met-or-unmet verdict.

Example completion criteria:

"End the round once every worker has either accepted the new schedule or openly refused it."

13. Private Self-Reflections

Tick "Private self-reflections" on a round card and, after the round's turns finish, each participating agent writes a short private note on what just happened — from its own point of view.

What "private" means

  • A reflection never enters the conversation — it isn't part of any response or turn, and no other agent can read it.
  • An agent's own earlier reflections are fed back into that same agent's prompts in later rounds.

So a private inner state can build up across a simulation while the outward conversation runs separately. That's the point for any theory where what an agent feels should diverge from what it says — resentment or alienation accumulating, private doubts, an inner self at odds with a public role.

In the live, completed, and submission views, a round's reflections start collapsed at the end of the round behind a 💭 N private reflections toggle — click it to read them as dashed 💭 bubbles (see Navigating a Transcript). They're included in the exported transcript, so you can analyse them afterwards. Reflections always see the full round, regardless of the round's Context Depth setting.

14. Running Experiments

Click "Launch Experiment" to start. You'll be taken to the live view.

The live view shows:

  • A progress indicator while the experiment is running
  • Live streaming of each agent's response as tokens are generated (you see the text appearing in real-time)
  • Color-coded agent names for easy identification
  • Round-by-round results with full markdown rendering (headers, bold, lists, etc.)
  • Agent cards in a right-hand rail beside the transcript, so who's who stays in view while you read
  • An action bar along the bottom with "Export JSON", "Data CSV", and "Evaluate"

Duration estimates:

With Claude Sonnet (3 agents, 3 rounds):

  • Short responses: a few sentences per agent
  • Medium responses: a paragraph or two per agent
  • Dynamic responses: varies with the discussion, since each turn is sized to the moment rather than run to a fixed length

More agents, more rounds, and more turns per round increase duration proportionally.

Going back to the builder

The same action bar carries the way back to the design the run came from:

  • "Edit design & re-run" appears when the run was launched from a saved design. It opens the Experiment Builder with that design loaded, so you can adjust a persona, a condition prompt, or a round setting and launch again. Save Design updates that design in place; use "Save as new" to keep the original and file the edited version separately.
  • "Back to builder" appears when the run came from a design you never saved.

A launch never discards your work. The builder keeps a local draft of whatever you are building, and launching saves that draft instead of clearing it, either way:

  • An unsaved build is offered back with a Restore banner the next time you open the plain builder.
  • While the Restore banner is showing, everything that acts on the design waits: saving, launching, batch running, and deleting a design. Each one would decide for you which copy to keep, so Delstorm asks you to Restore or Dismiss it first and then does what you asked.
  • A saved design is protected the same way, because launching runs what is on screen rather than what the design row holds: edit a persona and launch without saving, and the run used your edit while the saved design still has the old one. Coming back through "Edit design & re-run" loads the saved design and then offers those edits back with the same banner, so you can take them or dismiss them. Launch without changing anything and there is nothing to offer, so no banner appears.

Launching produces a new run every time, so re-running never overwrites the transcript you were just reading. Both runs stay in your experiment history. Runs launched before this link existed show "Back to builder", because they recorded no design.

A run that failed keeps the way back and nothing else: there is no transcript to export, chart, or evaluate, so those buttons are not offered on it.

Workspace assignment runs work differently on purpose: a run launched from a workspace assignment carries no builder link of either kind. That loop goes back to the assignment workspace, where the run was launched and where the next step of the assignment lives. Runs from a freeform build or a classroom's builder space keep both controls, leading back to the builder that launched them (the assignment's own, or the classroom's), never the generic personal builder.

15. Navigating a Transcript

Completed transcripts — the live results view, the guided workspace’s Analyze step, and the submission review page — share the same reading aids, so long runs don’t mean endless scrolling.

The Contents sidebar

  • A docked Contents sidebar lists every round with its messages. Click a round heading to jump to that round; click an agent’s name to jump straight to that message — it briefly pulses so you can spot it.
  • Analysis questions appear in the list at their place in the transcript, with a state icon: ✓ answered, ▸ current, ❓ open. Clicking one jumps to its ❓ marker in the transcript — the answers themselves live in the questions panel beside it.
  • As you scroll, the sidebar highlights the message you’re reading.
  • In a gated assignment, sections past your current question show as gray 🔒 locked rows — you can see the shape of what’s coming, but the content (and the question text) unlocks only as you answer.
  • On narrow screens the sidebar tucks away behind a floating ☰ button.

Folding

  • Rounds are collapsible — click a round’s divider to fold it away; click again to reopen. Rounds start open.
  • Private reflections start collapsed at the end of each round behind a 💭 N private reflections toggle — click it to read them.

Message details (ⓘ)

Every message has a small ⓘ button next to the agent’s name. Click it for that message’s context: the agent’s persona, specialty, and model; the round, turn, and condition prompt; and — for newer runs — the time the message was generated. Press Escape or click anywhere else to close it.

Citing a message

Every message carries a short citation reference beside the speaker’s name, in a plain monospace chip. It names exactly one message, so you can point at it in an answer instead of describing it (“the bit where Elena pushes back”).

An agent’s turn reads R2.T3.Elena:

  • R2 is the round, counted in the order the rounds actually ran. If a round repeats (a design can loop back), each pass gets its own number, so a reference is never ambiguous.
  • T3 is the turn within that round, counted from 1.
  • Elena is the agent, written exactly as the agent is named in the design.

Two other kinds of message use the same shape:

  • A moderator message reads R2.M.open, where the last part is the moment it spoke: open, turn_2, close, transition, or a speaker pick.
  • A private reflection reads R2.X.Elena.

Click the chip to copy the reference, then paste it where pasting is allowed, such as the note's quote field. The full reference is always shown, never shortened, so where an answer box requires typing you read it off the screen and type it in; it is only a few characters.

References appear while the run is streaming as well as on the finished transcript. An agent turn, a moderator message, and a private reflection each carry the same reference in both, so one you copy mid-run still points at the right message afterwards. There is one exception: the moderator’s speaker-pick narration shows its reference only once the run has finished, because a pick is not always kept in the record. A reference is tied to one experiment record: the same design run again produces a new record, and R2.T3.Elena in that run means that run’s message, so say which run you mean when you compare two.

16. Batch Running

Batch running lets you execute the same experiment design multiple times to observe how agent behavior varies across identical setups.

How to use:

  1. Build your experiment as usual in the experiment builder.
  2. Instead of "Launch Experiment", click the purple "Batch Run" button.
  3. Enter the number of runs (1-20).
  4. A batch progress page shows completion status for each run.
  5. Click into individual runs to see their full results.

Each run uses the exact same configuration — same agents, personas, rounds, and conditions — but produces different responses due to LLM sampling variability. Runs execute sequentially (not in parallel) to avoid API rate limits.

Tip: Use "Short" response length for batch runs to save time and API costs. Run 3-5 short batch runs to identify interesting patterns, then do a single thorough run to explore them in depth.

17. Saving & Loading Designs

Designs let you save and reuse experiment configurations:

  • Save Design: Click the green "Save Design" button in the experiment builder's bottom bar. Your agents (including personas), rounds, and all settings are saved to your account. If you loaded a saved design first, this updates that design in place (one design, not a second copy under the same name) and the confirmation reads "Design updated".
  • Save as new: Beside Save Design. This always creates a separate design, so it is how you branch a variant off a design you loaded while keeping the original untouched. The builder then follows the new copy, so the next Save Design updates the variant. A builder with nothing loaded creates a new design either way: a fresh page, and a built-in template with no documents (nothing is stored until you save it). A template that bundles documents is imported into your designs first, so the builder is then holding that imported design and Save Design updates it. Restoring unsaved work from a previous session keeps whichever design you were editing, so Save Design updates that one, but only once Delstorm has confirmed that design is still in your list. If you deleted it in the meantime, or if that check cannot run (your saved designs failed to load just then), the restored work comes back unattached and saving creates a fresh design instead. Either way the work itself is never lost. Saving into a design that no longer exists is refused with a message rather than reported as a success, so use "Save as new" to keep the work.
  • Load Design: Open the Designs drawer from the builder's bottom bar, select a saved design from the dropdown and click "Load". All fields are populated with the saved values, and a note in the drawer reminds you that Save Design will now update this design. Every dropdown that lists your designs (here, the Studio's attached design, and a classroom's attach and share pickers) labels each one with its name and the date you last saved it, most recently saved first.
  • Delete Design: In the same Designs drawer, select a design from the dropdown and click "Delete" to remove it.

What gets saved:

  • Experiment name, default provider, default model
  • All agents: names, specialties, personas, temperatures, memory settings, provider overrides, personal document references
  • All rounds: condition prompts, interaction patterns, turns, response length, participating agents, context settings, reference document

Saved designs store document references (library IDs), not the documents themselves. The documents remain in your Library.

Sharing designs with other users

You can share an entire experiment setup — agents, rounds, and the documents they reference — with another Delstorm user (or with yourself on a different account) via JSON export/import:

  • Export: In the Designs drawer, select a saved design and click "Export". A <name>.delstorm.json file is downloaded. By default, every referenced document (from rounds and from agent-level personal documents) is bundled inside the file so nothing is lost on the other end.
  • Import: In the Designs drawer, click "Import" and choose a .delstorm.json file. The design is added to your list (suffixed with "(imported)") and auto-loaded into the builder. Any bundled documents that don't already exist in your library are added to it.

How references are handled on import

  • Documents: Deduped by name + filename. If your library already has a document with the same name and filename, that existing document is reused — no duplicate is created. Otherwise, the bundled document is added with a fresh ID, and the design's references are rewritten to point to it.
  • Dangling references: If the bundle has no embedded content for a referenced document, that reference is stripped (the design still imports, but with a warning).
  • Providers: If the design uses a provider name you don't have configured (e.g., the sender's "my_openai"), the default provider is cleared and you're prompted to pick one before launching.

Note: export files can be large if the design references big documents. There's a 25 MB cap on import uploads.

Your work is protected

The Experiment Builder keeps a local draft of whatever you're building and offers it back with a Restore banner if you reload or come back later, including after you launch a build you never saved, so neither an interrupted session nor a run loses your in-progress design. You also stay signed in across a working session (your sign-in renews automatically), so you shouldn't be bumped to the login screen mid-build. If you ever do see "Session expired — please sign in again", just sign back in and your draft will be waiting.

18. Exporting Results

Click "Export JSON" on any completed experiment to download the full results. The JSON file includes:

  • Experiment metadata (name, duration, session ID)
  • Agent specifications (names, personas, temperatures, providers)
  • All rounds with condition prompts, interaction patterns, and full agent responses
  • Turn-by-turn data for multi-turn rounds
  • Timing data per round

The exported JSON can be used for further analysis in Python, R, or any data processing tool.

19. Evaluation

The Evaluate page lets you analyze completed experiments by defining structured evaluation fields and running a hybrid parser + LLM extraction system.

When to Use Evaluation

Evaluation is post-hoc and on-demand — you run it after an experiment completes, not as part of the experiment design. This means you can evaluate the same experiment multiple ways, or decide what to measure after seeing the results.

Access the evaluate page from:

  • The "Evaluate" link in the navigation bar
  • The "Evaluate" button on any completed experiment's results page
  • The "Evaluate All Runs" button on a completed batch page

How It Works

The evaluation system is hybrid — it combines two approaches:

1. Deterministic Parser (always runs)

Scans the transcript using regex patterns for explicit values. If an agent wrote "Score: 7/10" or "Consensus: yes", the parser extracts it directly. This is instant, free, and exact.

2. LLM Evaluator (optional)

Sends the transcript + field definitions + parser findings to an LLM for nuanced analysis. The LLM can understand context ("the agents generally agreed" → Consensus: true), fill in text fields ("summarize the key agreement"), and validate parser findings. Costs one LLM call per evaluation.

The parser runs first. The LLM runs second, informed by what the parser found. Results are merged with confidence tracking so you know whether each value came from the parser (deterministic), the LLM (inferred), or both (confirmed).

Field Types

Number — A numeric value, optionally with min/max range. Parser looks for patterns like "Score: 7" or "8/10". Example: "Consensus Score (1-10)".

Scale — Like number but intended for Likert-type ratings. Same extraction logic with range validation.

Text — Free-form text answer. Always requires the LLM (the parser can't infer text). Example: "Key Agreement Points".

Boolean — Yes/no, true/false. Parser looks for "Resolved: yes" style patterns. Example: "Reached Consensus".

Choice — Pick from predefined options. Parser looks for explicit mentions near the field name, or falls back to most-mentioned option. Example: "Winner" with choices ["Agent_1", "Agent_2", "Tie"].

Source Filtering

Each field can optionally specify a source round and/or source agent. When set, the parser and LLM only look at that portion of the transcript. This is useful when you have a dedicated evaluation round — you want scores from the evaluator agent's output, not from the full debate.

Evaluation Templates

Save your field definitions as reusable templates. If you always evaluate political debates with the same criteria (Consensus Score, Winner, Key Arguments), save those fields as a template and load it for each new evaluation. Templates are saved to your account.

Batch Comparison

When evaluating multiple experiments (from a batch run), results display as a comparison table — fields as rows, experiments as columns. Numeric fields show an average across runs. This is the core tool for measuring variance: "Did the agents reach consensus more often in runs with lower temperature?"

Example Workflow: Evaluator Agent

A powerful pattern is to include a dedicated evaluator agent in your experiment:

  1. Create an agent named "Evaluator" with a persona like "You are an impartial judge who evaluates the quality of discourse."
  2. Add a final round with only the Evaluator participating, with a condition like: "Score the preceding discussion on: consensus (1-10), argument quality (1-10), and name the strongest contributor."
  3. After the experiment, go to Evaluate and define fields: "Consensus (number, 1-10, source agent: Evaluator)", "Argument Quality (number, 1-10, source agent: Evaluator)", "Strongest Contributor (choice, source agent: Evaluator)".
  4. The parser will extract the scores directly from the evaluator's structured output.

Confidence Badges

parser — Value was extracted deterministically from the transcript text.

llm — Value was inferred by the LLM evaluator.

both — Both parser and LLM found the value (highest confidence).

none — Neither parser nor LLM could determine a value.

not_executed — On a branching run, the round this field points at did not run, so there was nothing to read.

recorded — Value was taken from the routing decision Delstorm recorded on a branching run, not read from the transcript. Only fields that report a routing outcome get this badge.

20. Data Extraction & CSV Export

The Evaluation page is something you run by hand after an experiment. Data extraction is the same idea built into the design, so the metrics you care about are pulled out of every run automatically.

Setting it up

In the experiment builder, open the "Data extraction" drawer from the bottom bar (a green dot on the button means extraction is already configured). Then:

  1. Click "+ Add field" for each thing you want to measure. A field has a Name, a Type (Number, Text, Scale, Boolean, or Choice), and a Description of what to extract. A Choice field also takes its allowed values, one per line.
  2. Optionally add "Extraction instructions" to guide the evaluator.
  3. Tick "Auto-extract after each run".

Now, whenever a run (or each run in a batch) finishes, the hybrid evaluator scores those fields and stores the result on the experiment — in the same place the Evaluate page reads from, so the Evaluate view shows it with no extra step. If extraction ever fails it's logged but never fails the run; your experiment still completes normally.

Builder drawer vs. the Evaluate page

The builder's drawer covers every field type (Number, Text, Scale, Boolean, or Choice, with a one-per-line choices box) and it keeps every setting a loaded design already carries, so loading a template or saved design and launching it never strips value ranges or source scoping. To author value ranges or to limit extraction to one round or one agent, use the full Evaluate page on demand. One setting has no authoring control anywhere yet: the flag that makes a field report a branching run's routing outcome arrives with a GPT draft, a template, or design JSON. The builder's drawer preserves it; the Evaluate page rebuilds a field from the controls it shows, so a field re-saved there comes back without the flag. Re-typing such a field to Number, Scale, or Boolean in the drawer also drops it, with a note on the row, because the outcome it reports is a round name.

Exporting to CSV

Once an experiment has extracted data, download it as a spreadsheet: on the finished experiment's page, click "Data CSV" in the action bar along the bottom, beside Export JSON. The same file is served from the experiment's data endpoint, which is also how a batch exports (the batch page has no button yet) and how you choose columns:

One experiment (a single row): /api/experiment/<session_id>/data.csv

A batch (one row per run): /api/batch/<batch_id>/data.csv

Add ?fields=Name1,Name2 to choose which columns to include, and (batch only) ?confidence=1 to add a confidence column per field. Cells are left blank where a run is missing a value.

The CSV opens cleanly in Excel, Google Sheets, Python, or R for further analysis.

21. AI Assistant

Delstorm includes a built-in AI assistant. Look for Pal, the small assistant character, in the bottom-right corner of most pages; in the Assignment Studio he sits in the bottom-left corner instead, so he stays clear of the design editor drawer. You need to be signed in to use it. The panel header has a context selector that always offers Platform help, plus one entry for each of your classrooms whose instructor turned on a Course TA. Each context keeps its own conversation (saved in your browser for 7 days; the trash icon clears only the active context's chat).

Platform Help

A how-to guide for Delstorm itself. Ask anything about features — setting up providers, building experiments, configuring agents, documents, evaluation, classrooms. It answers from this guide, names the actual pages and buttons, and is happy to write full worked examples — personas, condition prompts, settings — for you.

Example questions:

  • "How do I set up a Claude provider?"
  • "What's the difference between free discussion and debate patterns?"
  • "How do I use document grounding to create agent personas?"
  • "What evaluation field types are available?"

Course TA

If one of your classrooms has a course TA enabled, it appears in the selector too. A course TA answers from the materials its instructor loaded and follows the assistance level the instructor set. On a classroom, assignment, or workspace page the selector switches to that classroom's TA automatically. On an assignment's own pages (the assignment page, its workspace, or a freeform build's builder) the TA reads only that assignment's instructions and hints, never another assignment's; on the classroom page it sees every published assignment. Pasting is off for the course TA when you are a student in that classroom unless your instructor turned it on: type your message in your own words, the same rule as the answer boxes, and a short cue says so if you try. See Course TA below for the full picture.

Heads up: conversations with a course TA are logged and visible to that classroom's instructors — the note under the context selector reminds you whenever a classroom context is active. Platform help stays private to your browser.

Pal and page walkthroughs

Pal is the visual and text-status layer for the existing assistant, not a second chatbot. The same Platform Help and Course TA contexts, permissions, histories, and logging rules still apply. Pal's expression reinforces a short status such as Ready, Listening, Thinking, Done, or Something went wrong; that status is also announced as text for assistive technology.

On Home, the experiment builder, and the student workspace, type "Show me around" to start the walkthrough registered for that page and your role. This command is resolved locally, without an AI model request or a new chat log entry. Tours are available only through an explicit page-and-role registry; chat cannot supply arbitrary selectors, URLs, or scripts.

The workspace also has a "Walk me through" button beside its Guide button. That walkthrough is built from what the assignment actually has: it walks the three stages, reads out the guide sections your instructor wrote, points out the pre-run questions and the "Check my work" button where they exist, and skips whatever the assignment does not use. If your instructor wrote no guide, the workspace has no Guide button and therefore no walkthrough button, so use "Show me around" in the assistant instead.

The Builder walkthrough waits on real Builder fields as you work. On an empty Builder, completing it leaves you with two interacting agents and one round. Pal guides and waits; Pal does not write field content, choose a provider, save, or launch the experiment. If you begin with already populated work, it is reused and preserved rather than cleared or replaced. The final step means ready to launch, not launched; only your separate click on Launch Experiment starts the real run.

Assistant replies do not currently offer or start tours on their own. "Show me around" is the local way to begin one.

Use Back, Next, and Skip in the walkthrough panel. Some steps wait for you to perform the real page action. The page remains usable while the walkthrough is open; if the highlighted control is outside the panel, follow the prompt "Go to the highlighted control". Press Escape to exit at any time. When the walkthrough ends, focus returns to the assistant launcher.

Whichever context is selected, the assistant still sees the page you are on, and on the experiment builder page it can see your work in progress — a "Design shared with assistant" chip appears when your current design is attached, so its answers can target what you are actually building. Every context answers only from its curated knowledge base and will say so plainly when the material doesn't cover something.

22. Classrooms

Classrooms connect instructors and students for coursework: instructors post assignments built on experiment designs, students run them and submit their results, and anyone in the class can share designs. The Classrooms page lists every classroom you teach or have joined.

There are two roles. Instructors create classrooms, post assignments, and grade submissions — instructor accounts are granted by the platform administrator (an email allowlist), so contact the admin if you need one. Everyone else is a student.

Joining a classroom

Your instructor gives you a join code. On the Classrooms page, enter it in the "Join code" box and click Join. The classroom appears in your list; open it to find three tabs: Assignments, Shared Designs, and Members. You can leave a classroom from the Members tab.

Some classrooms set join requirements. A classroom can restrict joining to certain email domains (for example, only university addresses). If your account's email doesn't match, the join is refused with a message naming the accepted domains, and you'll need an account at one of them. A classroom can also ask for your official name as it appears in your school's course system: the join box reveals first and last name fields, and the name you enter shows on your instructor's roster and grade export so they can match your work to the course roster. If your classroom starts asking for names after you joined, a card on the classroom page asks for yours on your next visit. Neither requirement ever locks out someone already in the classroom.

Working on an assignment

These four steps are for an assignment that opens in the builder. If your instructor built a workspace for it, the assignment page shows "Open workspace" instead, with "Start in builder" beside it if you also want the design in your own builder to explore. On that kind of assignment you work through the workspace's own steps and submit there, so step 2's button and step 4's picker are not what you will see.

"Start in builder" copies the design, and only the design. You get the baseline agents, rounds, and any referenced documents in your own saved designs. The declared runs, the fields your instructor left editable, the analysis questions, the required transcript notes, and the instructor's guide are all part of the workspace, not part of the design, so none of them come with the copy. It is there to take the design apart and experiment with it. The attempt your instructor marks still happens in the workspace, and you still submit there.

  1. 1.

    Open the assignment from the classroom's Assignments tab. It shows the instructions, due date, and any resource links. A due date never blocks a submission: you can still submit after it passes, and Delstorm accepts the work exactly as it would before, and your instructor simply sees it marked as late.

    Whether you see your own late status is up to your instructor, and it is off unless they turn it on. When it is on, a Late marker sits beside your submission on the assignment page and on your submission page, and the workspace header shows a Past due note once the due date has passed. It is a marker, never a block: the workspace still lets you submit, and the work is accepted the same way.

  2. 2.

    Click "Start assignment". The attached design template — agents, rounds, and any referenced documents — is imported into your own saved designs, ready to load in the experiment builder. Clicking it again does not make a second copy: you are handed the copy you already have, edits and all, and renaming it makes no difference (Delstorm records which assignment the copy came from). To start over from the instructor's version, delete your copy first. What is imported is the design template on its own: anything the instructor set up as workspace content (declared runs, editable slots, analysis questions, required notes, the guide) stays on the assignment and does not travel into your saved design.

  3. 3.

    Run the experiment as usual, adjusting the design first if the instructions ask you to.

  4. 4.

    Back on the assignment page, pick the completed experiment, optionally add notes for your instructor, and click "Submit". A workspace assignment has no completed-experiment picker on this page: you submit inside the workspace, and a run you launched from the builder is not accepted there.

Practice runs

The experiment builder is yours to use whenever you like. Build a design of your own, run it as often as you want, change it and run it again: nothing there is submitted to anyone, so that is the place to practice. When an assignment opens in the builder (step 2 above), the copy imported into your saved designs runs there too, so you can run it, read the transcript, adjust it, and run it again before you submit anything.

An assignment that opens in a guided workspace works differently: the runs you launch there are your real attempt. You can still go back to Step 1, adjust whatever the assignment leaves editable, and launch again as often as you need before you submit. Each launch runs a new experiment, and your earlier attempts stay on record. There is no separate practice mode inside the workspace, and your instructor sees your work when you submit it.

Submissions are snapshots

Submitting freezes a copy of the experiment — editing or deleting the experiment afterwards doesn't change what your instructor reviews. To hand in a newer version, submit again: Resubmit replaces your previous submission (an existing posted grade is kept until the instructor re-reviews). On a resubmit the experiment picker opens on the run you submitted last time, so adding a note does not mean finding it again. Pick a different run if you mean to hand in a different one. Grades and feedback appear when your instructor posts them: instructors grade in private drafts, often finishing the whole class first, so there may be a gap between handing in and seeing feedback. Status badges track the lifecycle: submitted (awaiting review, or reviewed but not yet posted), reviewed (your instructor posted the grade and feedback; they are on your submission page), resubmitted (changed since the last posted review).

For instructors

Create a classroom & share the join code

On the Classrooms page, fill in a name and description and click Create. The 6-character join code appears on the classroom page (with a "Copy code" button) — share it with your class. Archiving a classroom blocks new joins and asks first (students keep their work and access; unarchive to open joins again).

Co-instructors: from the Members tab, the classroom owner can add co-instructors by their account email. Co-instructors can create and edit assignments, see drafts, review submissions, view the roster and join code, and remove members. Only the owner can delete the classroom or manage co-instructors (a co-instructor can remove themselves). Adding someone who already joined as a student promotes them out of the student list.

Course TA

Each classroom can have its own AI course TA — off by default. Open "Course TA settings" from the classroom page (any instructor of the classroom can manage it) and configure:

  • Enable — until you turn this on, students see only "Platform help" in the assistant; nothing changes for them.
  • TA name — how the TA introduces itself and appears in the selector.
  • Persona — private instructions that shape the TA's voice and rules (students never see the text itself).
  • Assist level — Hints (questions and pointers only, never drafts text a student could submit), Guided (explains and gives worked examples but won't answer your assignment questions), or Open (ordinary helpfulness, including drafting).
  • Course materials — paste text or upload PDF/TXT/MD files; the TA answers from them. Give each a descriptive title (the TA finds material by title). Up to 40 materials. Enrolled students can read the materials and persona — the TA draws its answers from them — so don't include answer keys or private rubrics.
  • Include published assignments — on by default, so the TA knows the instructions students are working from. When a student chats from an assignment's own pages the TA reads only that assignment; on the classroom page it sees all of them.
  • Allow students to paste into the TA chat — off by default: students type their messages in their own words, the same rule as the answer boxes, and a short cue says so if they try. Turn it on when you want them to paste, for example a passage from the readings. Instructors can always paste, whatever the setting.

The "Start from SOCB43 template" button fills all of this in with a ready-made Socratic design TA (persona, hints level, and the course materials) and enables it in one click — you can then edit anything. Students pick the course TA from the assistant's context selector; on classroom, assignment, and workspace pages it is selected automatically. Page and design awareness work in the course TA just as they do in Platform help.

TA activity (monitoring)

Every conversation members have with your course TA is logged. Open "TA activity" (from the classroom page or Course TA settings) to read them grouped by student, expand any conversation in place, download a single conversation as markdown, or export the whole classroom's chats as CSV. Students see a standing disclosure whenever they talk to a course TA. The log is append-only and is deleted with the classroom. Logging is best-effort: a storage failure never blocks the TA from answering, so a message can occasionally be missing — treat the log as a record, not a proof.

Create assignments

On the Assignments tab, give the assignment a title, markdown instructions, an optional due date, and optionally attach one of your saved designs as the template. The design is attached as a snapshot (documents bundled) — editing your original design later doesn't change the assignment. Untick "Published" to keep it a draft only you can see. To build an assignment on one of the built-in templates, load it in the experiment builder via "Start from template" first — it lands in your saved designs, ready to attach.

The "Student guide" builder in the Assignment Studio's Student guide & support section lets you write step-by-step guidance directly into the platform — ordered markdown sections, each optionally tagged to a workspace step, that students open from a "📖 Guide" button. Unlike the analysis questions, which freeze onto each submission once answered, guide edits reach students immediately, so you can fix or expand it mid-assignment. Loading a built-in starter can pre-fill draft guide sections for you to adjust.

What the step tags mean: General shows across the whole workspace rather than tied to one step; Set up is what students see before configuring and starting a run; Run shows while a simulation runs; Analyze & submit is the write-up and submission step.

Copy an assignment

The assignment page's Copy button (beside Edit in Studio) duplicates an assignment so you can iterate on a new version without touching the one your class is using. A dialog opens prefilled with "Copy of" plus the title, editable before the copy is created, and you land on the copy's own page. The dialog's "Copy to" chooser picks where the copy lands: "Make a copy in this classroom" is the default, and any other classroom you instruct (as owner or co-instructor) can be chosen instead, which is how a finished assignment moves from a testing classroom into the real one. The copy carries the attached design snapshot, instructions, resources, due date, workspace scaffolding, analysis questions, and the student guide. That snapshot includes the provider and model the assignment runs on, so nothing has to be set again on the copy, and how the name resolves depends on who runs it: a shared platform provider (the usual classroom case, such as a course-wide deepseek) needs nothing set anywhere, a guided or open workspace run falls back to the runner's own default provider and model when the named one is not one of theirs, and a freeform or "Start in builder" import blanks a provider the runner does not have, so the builder asks them to choose one before launching. It always starts unpublished, whatever the original's state, so students never see a "Copy of" draft appear mid-course, wherever it lands; publish it when it is ready. The copy and the original are independent snapshots: editing one never changes the other, and submissions, drafts, and student work stay with the original. One setting is deliberately not carried: "Show students their own late status" starts off on the copy, so turn it back on there if you use it.

Review & grade

A due date never blocks a submission; work that arrives after it is marked Late on your submissions roster and on the review page header. A submission landing exactly on the due timestamp counts as on time, and since resubmitting replaces the submission, a resubmit after the due date reads Late even when the first attempt was not: the badge describes the snapshot you are reviewing.

Students are never blocked, and whether they see the marker themselves is yours to set per assignment. The Assignment Studio's Purpose & brief section has a "Show students their own late status" checkbox, off by default. Turn it on and a student sees the Late marker on their own submission line and submission page, plus a "Past due" note in the workspace header once the date has passed. Your roster, this review page, and the grades CSV show late work either way: the CSV carries a late column beside the submitted timestamp, reading true or false for anyone who submitted and blank for anyone who has not.

The assignment page lists every student's submission with its status. Open one to read the full experiment transcript — the student's answers and the grading form sit in a panel beside it, so you can grade against the evidence without scrolling back and forth (anchored answers link back to their spot in the run). When the reading is done and the grading starts, press "Hide transcript" at the top of the transcript: the transcript and its rail tuck out of the way and the answers panel takes the width at a comfortable reading measure, with a docked "Show transcript" tab to bring it back. Anything that points into the transcript reopens it for you first. It is a per-visit choice, not a setting.

The overall grade (free text — "A-", "85/100") and overall feedback sit at the top of the review panel. Per-answer mark and feedback boxes are there when you want them, behind an "Add per-question marks" reveal, so grading overall-only never means scrolling past unused boxes, while a review that already has per-question entries opens with them showing; one Save review button stores all of it, and marks typed on one run tab ride the save from any other. When every mark you've entered is a plain number, a line under the grade box offers their total ("Marks total: 85", with a "Use as overall" button); it fills the grade box only when you press the button, so a letter grade or your own total always wins. If a student resubmits after you review, feedback tied to questions their new submission no longer has is kept in a clearly labeled card rather than lost. For a multi-run assignment, the submission page adds run tabs — one frozen transcript per run, with that run's notes and answers beside it — and a Comparison card with the cross-run answers; there is still one review per submission.

A new review saves as a private draft: the student sees nothing (no grade, no feedback, and their status badge stays submitted) until you post it. Saving never changes posted-or-draft state, so an unposted review stays private through every edit, and an already-posted one stays visible. Draft as you like, revise across sittings, grade the whole class before anyone sees anything. Post one student's review with "Post to student" on their review page, or the whole class at once with "Post all grades" on the assignment's submissions roster (it names how many drafts it will post and never touches already-posted reviews). Posting is reversible: "Unpost" returns a review to a private draft. Once posted, the student sees the overall grade, the overall feedback, and each answer's mark and feedback beside the answer itself, and edits to a posted review stay visible as you save them; unpost first to rework in private. The roster marks each review Draft or Posted so you always know who can see what. Reviews saved before this feature existed were already visible to their students and stay that way.

The grades CSV now leads with identity columns: each student's official first and last name (collected at join when you turn that on in the classroom settings) and their account email, ahead of status, grade, and the timestamps, so grades can be matched back into your course system's gradebook. In the classroom's Edit panel you can restrict joining to certain email domains and ask new members for their official name; both apply to new joins only.

Download a submission as JSON. Every hand-in can be downloaded whole, for grading by script or by eye. On a student's review page, "Download JSON" (in the row with Save review) saves one file holding everything on that submission: the assignment's brief and every editable setting with the design's current baseline and your expected answer where you set one; per run, a settings table with the student's value, whether they changed it from the baseline (judged against the baseline the run was launched under for a typed setting, the current one otherwise) and whether it matches the expected answer, what a text setting actually ran with where the run record holds it (the experiment name, personas, specialties, round names and condition prompts), and their "Why this choice?" text; their transcript notes with the transcript position each one points at (a round and turn with its speaker, a moderator message, or a reflection); their answers paired with the question as it was asked; the full transcript in reading order (the moderator's messages and the agents' reflections included) beside the stored run record; their notes to you; the names of any attached files; and your review as it stands, draft or posted. On the assignment page, "Download all submissions (JSON)" beside the grades CSV saves one file with the same block for every hand-in, in order of submission, plus the list of members who have not handed in. Both are instructor-only: a student cannot fetch either, and nothing about the submission changes when you download it. The files are versioned ("version": 1) so a checker written against them keeps working.

Reopen answers (a redo). In a staged-reveal workspace an answer locks once it is submitted, so a student you allow to redo an assignment cannot revise the answers they already locked in, even after running again. On that student's submission review page, "Reopen answers" (below the review buttons; press it once, then confirm; the button says how many answers it will reopen) clears from that student's workspace draft the answers the student has locked in by submitting them (work saved before the fix of 2026-09-24 reached the student's page is judged by where they had got to: an earlier answer counts as locked in when a later one holds text, and their last started answer comes back open on its own): on every run, each in-transcript question they pressed "Submit your answer" on, with the evidence notes attached to it (the notes themselves stay in their notepad). An answer not locked in on the student's draft is left alone, whatever its state (short of a minimum, past a maximum, missing evidence); so are unanchored answers, comparison questions and before-the-run commitments; a submitted answer stays locked in even if you change a limit or an evidence minimum afterwards, and this control is how it reopens; their notes, runs and setup stay as they are. On a multi-run assignment a draft parked at Compare is taken to the first run with a reopened question. It also reopens answers the student locked in on this submission that are missing from their workspace draft (a draft that lost its answers after submitting, or a student with no draft at all): those questions get the same reopen note on the student’s next load, any older copy of the answer still in their browser is set aside rather than restored, and a page they left open cannot submit the old copy. The line under the button says which questions are locked in on their draft, which are locked in on this submission but not locked in on their draft (a partial they have started there stays theirs), which are already open, and which are unanswered on both; “Nothing to reopen” means no guided question is locked in on their draft or answered on this submission, so the student can answer and resubmit now. Once a reopen has come from this submission, the control reads “Nothing to reopen” again until the student submits again, so the submitted copy is never offered twice; an answer the student locks in again on their draft can still be reopened. If the assignment shows the whole transcript at once there is nothing locked, so the control does not appear. It does not touch the submission you are looking at: their earlier answers stay readable here and on their own submission page, and the roster keeps showing this submission until they submit again. On their next load the workspace tells them you reopened the questions, they answer them again in order, and they resubmit; the review page then shows who reopened and when. It is per student and reversible only by the student answering again; if you want every student to see the whole transcript at once instead, that is the Studio's reveal setting, not this.

Sharing designs

Anyone in the classroom can share a saved design from the Shared Designs tab — pick one of your designs and click "Share with classroom". The share is a snapshot bundled with the design's documents, so later edits to the original don't affect it; share again to distribute an updated version. Classmates click "Copy to my designs" to import it — bundled documents land in their library (deduplicated against what's already there), and provider references they don't have configured are cleared, just like a design import.

Classroom Builder & Freeform Builds

Every classroom has its own experiment builder space: the "Experiment builder" button on the classroom page opens the full builder for any member, any time. Runs launched there are funded by the classroom (the classroom allowance is spent first; your own starter tokens absorb any overflow), so it is the place to explore freely on course resources. Exploration runs belong to no assignment and cannot be handed in, and a completed run's page links you back to the classroom builder, so re-running keeps the same funding.

Freeform builds (for students)

A freeform build is the assignment shape that hands you the full experiment builder instead of a fixed workspace. The assignment page's "Open experiment builder" button opens the assignment's own builder: design and run any experiment you like there, funded by the classroom. Only runs launched from this assignment's builder count toward the assignment; runs made in your personal builder or in the classroom's exploration builder can never be handed in, so always start from the assignment page. If your instructor attached a design it is an optional starter, not a lock: it appears in your saved designs the first time you open the builder, yours to edit or ignore. If the assignment has a 📖 Guide, it comes into the builder with you: a Guide button sits in the banner at the top beside "Review & submit" and another in the bottom action bar beside Designs and Data extraction, and both open the same drawer you saw on the assignment page (Expand and Dock both work there: docking narrows the builder's two side rails a little to preserve a usable editor width and makes room for the guide down the right-hand side, so nothing is covered, and the builder remembers a docked guide for that assignment until you close it with the ✕).

Handing in: back on the assignment page, your completed runs appear once at least one finishes, and the analysis questions unlock at the same moment. Pick which completed runs to hand in (the assignment says how many it requires; you may always hand in more, up to six), give each one a name your instructor will see, answer the questions (word minimums, maximums, and per-question paste rules apply as usual), add optional notes, and submit. Submitting freezes your chosen runs and answers together, and resubmitting replaces the whole previous submission.

Attached files. When your instructor has allowed it, the submission card also carries an "Attached files" block, there from your first visit so you can attach design notes before any run completes: choose a file and click Upload. Up to 10 files, 10 MB each and 50 MB in total; images (png, jpg, gif, webp), PDFs, text and markdown, CSV, JSON, and Jupyter notebooks. Each file has a "hand in" checkbox: the ticked files go in with your runs when you submit. A file that has been handed in shows a chip and cannot be removed until a resubmit that leaves it unticked releases it back to your list, where Remove works again; nothing is deleted by a resubmit. Your own submission page shows the attached files exactly as your instructor sees them.

Freeform builds (for instructors)

Turn the shape on in the Assignment Studio's Student agency section with "Students get the full experiment builder"; beside the guided and open workspace shapes it is the third way an assignment can meet students. "Required variants" (1 to 6, default 1) sets how many completed runs each student must hand in. Questions on this shape come in three stages: a before-the-run question is answered at the first launch that asks it and gates that launch (one added later is asked at the student's next launch), design questions (below) are answered beside the elements of the student's own design, and everything else is answered at hand-in. Word limits, paste rules and attached-note minimums work unchanged (students take their notes on each run's own page and attach them at hand-in; the round evaluation can ask for attached notes too), while run-scoped questions and transcript anchors do not apply, and the Studio offers a one-click conversion when an existing question set needs adjusting for the shape. Editable design fields, declared runs, and required transcript notes do not apply either; attach a design and it becomes an optional starter students import, not a lock. You and your co-instructors can open the assignment builder yourselves, before publishing included, and test-run without being metered, though a preview run can never be handed in. On the submission review page the handed-in runs read as tabs under the names the student gave them, with the answers under their own heading; per-question marks and grade draft and release work exactly as on other shapes.

Design questions

For students. A freeform build may ask design questions: questions about the elements of your own design, asked once per element. "Each agent" and "Each round" questions appear once per agent or round you build; the other kinds appear only when you change that setting from its default (memory window, context window, response length, interaction pattern, moderator, or turns). In the assignment's builder each question has an answer box beside its element, and the Design questions panel lists everything asked of your current design. Answers typed there are recorded with the run when you launch and shown at hand-in as what you wrote at launch; the hand-in answer is yours to finish or revise, and it is the one that is graded. Every launched run carries the design questions exactly as they stood when it launched, so an instructor's later edit applies only to later runs. A before-the-run question is answered at the first launch that asks it (your first launch, or your next launch if your instructor added it later), gates that launch, and is then recorded for the assignment; a round evaluation question, when the assignment asks one, is answered at hand-in for each round of the latest run among those you hand in. Each run allows at most 20 question instances with 600 minimum words between them, and at most 20 round evaluations with their own 600 minimum words; a run that would exceed either is refused at launch with the count. A run launched from an assignment's builder is bound to that assignment and cannot be deleted while the assignment exists; it stays in your list as evidence, so start a fresh run rather than trying to remove one.

For instructors. Add design questions under the freeform build's settings in Student agency: pick the trigger, write the question (1,000 bytes, with a live counter and markdown preview), set word bounds (0 to 500) and the paste rule, and tick "Ask an evaluation question for each round" for a per-round evaluation. Before-the-run questions keep their stage on this shape and gate the first launch that asks them; one added after a student's launch is asked at their next launch or answered at hand-in. On the review page every generated answer shows its context (run, element, and the setting's value against the default), the as-written-at-launch text when there is one, and the usual per-question marks. A freeform build is one-way once it has been live: the moment a freeform assignment is saved published, its shape can no longer change back to a workspace (unpublishing does not undo that), because students may have started runs against it, and an assignment that already has submissions cannot become a freeform build. Use Copy assignment for a new shape or a fresh freeform version; a copy starts unlocked and unpublished.

Attached files (instructors). Tick "Students may attach design notes at hand-in" under the freeform build's settings in Student agency to let students attach files (off by default, so existing assignments are unchanged). On the submission review page an "Attached files" card sits above the answers with each file's name, type, size, when it was uploaded and when it was handed in; images show inline and every file has a Download link. Files are stored exactly as uploaded, kept with the submission, and removed with the submission, the assignment, or the classroom; they never expire on their own. Turning the checkbox off later refuses new uploads while files already handed in stay with their submission.

Tokens & Usage

When a run uses one of the platform's shared AI providers, the text it generates costs real money, so student usage is measured in tokens (about four characters of generated text each). Runs against a provider you configured yourself in Settings (your own API key, your own spend) never touch your token balance, and instructor accounts are not metered at all.

For students

Every student account starts with 25,000 tokens of its own, enough to genuinely try the experiment builder. Classroom work runs on tokens your instructor grants: a per-student allowance for the classroom, plus top-ups whenever you need more. Classroom tokens stay in their classroom; your own starter tokens work anywhere. You can see what you have left on Home, under the run button in an assignment workspace, and on the classroom page's Members tab.

When your balance is empty, new launches are blocked; a run that has already started always finishes, even if it takes you a little below zero. The message points you back here: ask your instructor for more, or press "Request more tokens" on the classroom page, which flags you on your instructor's roster so they can top you up. Only one of your metered runs can be in flight at a time, and a batch stops starting new runs when the balance runs out (finished runs are kept, the rest are marked skipped with the reason).

For instructors

The Tokens card on your classroom's Members tab is the whole control surface. Set a per-student allowance and every current member is topped up to it; a student who joins later gets it on joining, and raising the allowance grants everyone the difference. Lowering it never takes tokens back, and nothing ever resets: top-ups are additive, for the whole class or for one student at a time, and a targeted top-up never touches classmates. The spend view shows each student's granted, used, and remaining tokens, with the used figure split honestly between what classroom tokens covered and what the student's own starter tokens absorbed, plus a "Requested more" badge when a student has asked; topping them up clears it. Your own runs, previews, and the classroom TA's replies are not metered, and co-instructors are exempt like you.

Reviewing Variant Changes

On a multi-run submission, select a variant in the run tabs. The Variant changes card reconstructs the student’s editable-field changes from the inputs saved with that submission. It is a read-only review aid and does not change the submission, its grade, or any experiment.

The card opens in Stacked layout. Choose Side by side to compare the reference run and selected variant in adjacent columns, each row carrying its own Changed, Added, or Unchanged marker. The layout choice changes presentation only. It does not edit the stored inputs, and complete text wraps instead of being truncated on narrow screens.

A step later in a chain can show two clearly labeled blocks. Changes from <source run> shows what changed from the run that supplied this variant’s starting values. Changes from Baseline shows the cumulative change from the first run. When the source is Baseline, the two references collapse into one Baseline block.

If a locked or empty run sits immediately before the variant, the source label reaches back to the nearest earlier run with saved inputs. This keeps the label aligned with the values the student actually inherited.

In the Stacked layout, prompt and other text fields put the complete source and current values inside an expandable disclosure, while typed settings such as choices, numbers, toggles, first speaker, and moderator read inline as old value to new value. In Side by side, every field type appears as a row that names the field and its status on its own line, with the reference and current values in adjacent columns beneath it and complete text wrapping instead of truncating. Every row says Changed, Added, or Unchanged; color is only a secondary cue. A no-op variant says that there are no changes from its reference instead of leaving a blank panel.

Each Changes from block folds, and in Side by side each row folds too. Everything starts open; closing a fold parks it while you read the next one, and a closed row still shows its field name and status. Stacked one-line settings do not fold, since there is nothing under them to hide.

Within a changed text field, the words that differ are highlighted in both panes: removed words tinted in the reference value, added words in the current one, so a one-word prompt edit is visible at a glance inside a long prompt. The full text of both values still reads exactly as stored. On a very long or very heavily rewritten value the highlighting steps aside and the two panes render plain, complete as always.

23. Guided Workspace (for students)

Walk me through. New to the workspace? Press "Walk me through" beside the Guide button and a short walkthrough points out each part of the page in order: the stages, your instructor's guide sections, the pre-run questions where they exist, and where you run, analyze, and submit. It walks what your assignment actually has and skips the rest. If the assignment has no guide there is no Guide button and no walkthrough button; type "Show me around" in the assistant instead.

In-transcript questions. Some assignments place questions at specific points in the run — after a particular message or at the end of a round. You answer everything in the questions panel beside the transcript: slim ❓ markers show where each question anchors in the run, and clicking a marker highlights its question in the panel; the question's "↗ Jump to this spot" button takes you back to its place in the transcript. In a gated assignment the transcript reveals one section at a time: read the section, type your answer in the panel, and choose "Submit your answer" — once it meets the word minimum, and stays inside the word maximum when the instructor set one, that answer locks and the next section appears below, without moving you. Where a question has a maximum, the counter under the box reads your count against it ("18 / 50 words"); going over never trims a word of what you wrote, it only holds the answer back until you shorten it yourself. An answer locks only when you press "Submit your answer": nothing locks while you are still typing, and a refresh, a dropped connection, or signing back in brings an answer you have not submitted back exactly where you left it, still editable (for work saved before the fix of 2026-09-24 reached your page, an earlier answer counts as submitted when a later one holds text; if the page cannot load your saved work it says so and shows what this browser holds, with any answer you had started read-only for that visit; reload once you are signed in again and it is yours to edit again). Some instructors instead show the whole transcript at once. Any final set of questions, your notes to the instructor, and the Submit button live in the same panel — nothing hides at the bottom of the page. On narrow screens the panel sits below the transcript.

Hiding the transcript while you write. When the reading is done and the writing starts, press "Hide transcript" above the transcript: the transcript and its contents rail tuck out of the way and the questions panel takes the width at a comfortable reading measure. On a wide screen a "Show transcript" tab also stays docked at the left edge of the page, so you can bring the transcript back without scrolling up. Nothing is covered and nothing is lost while it is hidden: anything that points into the transcript opens it for you first, whether that is a question's "↗ Jump to this spot", a jump from the contents rail, or a link to a particular message. A note you were part-way through writing is waiting where you left it too, even if you answered a question in between. The Compare & submit stage has the same control. It is a per-visit choice, not a setting: the transcript is showing again the next time you open the workspace. On a narrow screen, or when you are zoomed well in, there is no floating tab: the page stays stacked and the same button at the top of the transcript, now reading "Show transcript", brings it back.

Formatting what you write. Every box you write prose in has a small toolbar above it: your analysis answers, a before-the-run prediction, a "Why this choice?" reason, a transcript note, and your notes to the instructor. The buttons are B for bold, I for italics, and • List for a bulleted list. Select the words you want to change and press a button; press one with nothing selected and it drops in an example you can type over. You can also type the markdown by hand (**bold**, *italics*, - item), since the buttons only save you knowing the syntax. Under each box there is a Preview you can open, which shows exactly what your instructor sees, and it is what they see: your formatting is kept when your work is read back, on the submitted work, in the "Committed before this run" box, in the Compare stage, and on your note cards in the transcript. Formatting does not change your word count, which always counts the words you typed. The design slots in Set up have no toolbar: that text configures the experiment rather than being read as writing, so it is shown exactly as typed.

Your notepad. Every note you take in the transcript also lands in your notepad. The "📝 Notepad" button at the top right of the workspace slides open a drawer listing every note you have taken, grouped by the run it came from, each showing its quote, your own comment, and the moment it points at. On a multi-run assignment the notepad keeps growing as you go, so a note you made in the Baseline is still in front of you while you write about a variant, and a staged reveal is no different: new notes stay addable at any time, and each one joins the list. Only one panel covers the page at a time, so opening the notepad closes a "📖 Guide" drawer that is covering it, and opening the guide over the page closes the notepad; a docked guide stays exactly where it is, since it sits beside your work rather than over it. Close the notepad with Escape, the Close ✕ button, or by clicking outside.

Attaching a note to an answer. Click into the answer you want the evidence on, open the notepad, and press that note's Attach button, which names the answer it is about to attach to ("Attach to Baseline Q2"). The note appears as a fixed block under the answer box: the quote first, your comment beneath, and an "↗ Jump to this spot" button that opens the run the note came from and scrolls to the moment. The jump works while that run's transcript is open to you and the moment is still in that run's transcript: if you have re-run that run, the button tells you its transcript is not ready yet and leaves your note attached where it is; if that run is not open yet because an earlier run has to finish first, it says that instead; and if the moment is no longer in the transcript it says so instead of jumping. The note stays readable under your answer either way. You cannot edit the block there; it is the note itself, so editing the note in the transcript updates it. An attached note is evidence rather than writing, so it never counts toward the word minimum or the word maximum: those read only the words you typed in the box. Detach takes a note back off while the answer is still yours to edit. If a question asks for a number of attached notes, a line beside the word counter reads how many you have so far ("1 of 2 notes attached"), and that answer cannot be locked in or submitted until you reach the number. Before-the-run questions never ask for attached notes, since they are answered before there is a transcript to quote.

Deleting a note you have attached. Deleting a note that one of the assignment's current questions is leaning on is refused, and the message names the answers using it ("Used as evidence on Baseline Q1"): detach it there first, then delete it. That protection follows the questions the assignment has right now. If your instructor removes a question after you have answered it, that answer leaves your page and stops holding anything back, and when you submit, the answer goes with the deleted question and the evidence attached to it goes with the answer; your notes themselves stay in the notepad, ready to attach somewhere else. Once an answer locks, its evidence locks with it: the blocks stay exactly where they are and their Detach buttons go dead, while your notepad carries on growing for the rest of the assignment.

The notepad on a freeform build. A freeform build has no workspace, so the notepad splits across two pages. You take notes on a run's own page once it has completed (a run launched from the assignment's builder; its page is the one that showed it running, and you can reopen it from your experiments list on the home page, or later from a note's "↗ Open run"): highlight a passage and press "+ Add note" beside the message, exactly as in the workspace, and every note lands in your notepad. Back on the assignment page, the "📝 Notepad" button opens the notepad grouped by run under the names you gave your runs; click into an answer (an analysis question or a round evaluation), press a note's Attach, and it sits under the answer as a fixed block with "↗ Open run" (that run's page at the note's spot, in a new tab, since the transcript is not on this page) and Detach; a line under the answer counts the notes attached against any minimum the question asks for, and the round evaluation can ask for attached notes too. Your notes travel with the hand-in for the runs you hand in (a note on a run you leave unselected stays in your notepad): the ones you attached read on the review page under your answers, and the rest read in that run's transcript there. Unlike a workspace answer, nothing locks here: your notes stay editable after you hand in, but a hand-in already made keeps the notes as they were, so hand in again to carry an edit; and a note one of your current answers leans on cannot be deleted until you detach it. A note space line in the notepad shows how much of the notes' room you have used.

Some assignments open in a workspace instead of the full builder, a focused, three-step page that walks you from setup to submission. You'll see an "Open workspace" button on the assignment when your instructor has set one up; assignments without one use the regular builder. Some workspaces are open: the instructor left no editable slots, so Set up has nothing to fill in. You review the locked design, run it, and everything else (marking moments, answering questions, submitting) works the same.

The three steps

1. Set up

The experiment is already built by your instructor and shown read-only. If the instructor left any design fields editable, you fill in only those slots (a persona, a condition prompt, the experiment name…), each with a hint and a character counter. Like the analysis answers and the note text, the text slots are typed by hand: pasting is disabled in them unless your instructor allows pasting there, because a persona or condition prompt you fill in is your own thinking too. On a multi-run assignment the instructor can also allow or block pasting for a single run, so a box that requires typing on the Baseline may accept pasting on a later variant. Editable settings (a dropdown, a number, a toggle) show the instructor's hint under their label the same way, when one was written. On a multi-run assignment, the base hint carries to every run unless your instructor customizes that run's hint, in which case that run shows its own wording, which can be a hint on a field that carries none on the other runs. A locked run shows no hints at all, since nothing on it is editable. Editable text boxes start from the design's own words, so you edit what is there instead of retyping it; clear a box and leave it empty and the design's text comes back the next time the page fills the slots. You never have to hunt for what is yours: an "Editable" chip sits on every control you can change, and the subtitle of the Experiment card counts them ("3 settings on this page are yours to choose"). Everything without a chip is your instructor's design, shown so you can read it. In an open workspace there are no slots to fill, so you just review the locked design; there, and on a locked run, nothing is editable, so no chip and no count appear and the page says so in its own words instead. If a round is run by a moderator, its card also shows the moderator's rules read-only ("How the moderator ends the round" and "How the moderator picks speakers") so you can see the instructions your round runs under. When you're ready, click "Run experiment →".

"Check my work" gives quick per-slot feedback: a green "✓ Looks right" or an amber "Fix: …" chip, and a summary line counting how many checks passed. When any check fails, the summary line also says: “Some of your responses differ from the suggested value. If you changed a response intentionally as part of your experiment, you may keep it. Otherwise, change it to the suggested value.” A deliberate change is allowed; the check is a reminder, not a rule. For a text slot it checks only whether the slot is filled in and on-topic for its hint. For a setting your instructor set an expected answer on (a dropdown, a number, a toggle), it simply compares your choice to theirs: the chip tells you whether it matches and points you back to the guide, but it never reveals what the expected answer was. It does not grade your ideas, and it does not block running, so you can run without it.

A heads-up before you run. Editable text boxes start out filled with the design's own words. If a box still reads the text it started with (the design's text, or the hint itself pasted back), or a persona, a condition prompt, or another prose slot (custom instructions, the moderator's ending or speaking instructions) has fewer than five words, a note appears under that box and above the Run button naming it: the agents will run with exactly what is there. It does not stop the run; that is your call, the same as the round-evaluation line. Your instructor sees the same note beside that field on the review page.

Two things an assignment can ask of you. If your instructor turned on "Students must actively choose each editable setting", every editable dropdown, number, and toggle opens on "Choose..." instead of the design's value, and you cannot run until you have picked each one. Picking the same value the design already had is a perfectly valid choice. If they turned on "Students must justify kept preset values", each editable setting also gets a "Why this choice?" box: it is required only when you keep the value the design already had, and optional when you change it. What you type there autosaves with the rest of your work, so it survives a refresh or a move to another device, and your instructor sees it beside the setting it explains when you submit.

2. Run

Watch the simulation stream live, just like the normal live view. If you close the tab and come back, the workspace reconnects to your run.

3. Analyze & submit

Read the finished transcript and mark the key moments: highlight a passage that matters, then click "+ Add note" — the highlighted text is quoted into your note automatically (or type a quote by hand) — and write why it matters. Already have the note open? Highlight the passage and press "Quote it" inside the note to copy it into the quote box. Notes attach to agent turns, moderator messages, and the agents' private self-reflections (open a round's "N private reflections" fold to note those; a fold holding one of your notes stays open and shows a note count). A chip tracks your progress ("2 / 3 moments marked"). Like the typed analysis questions (word counters, and pasting disabled unless your instructor allowed it on a question), the note text is typed by hand — pasting is always disabled in the note box, whatever the questions allow; only the quote field accepts pasted transcript text. Every note you take is also collected in your notepad, and from there you can attach one to an answer as evidence (see Your notepad above). Add optional notes for your instructor and click "Submit".

Running variants (multi-run assignments)

Some assignments declare several runs — typically a locked Baseline plus one or more variants. A run bar at the top of the workspace shows every run (✓ done, current, 🔒 not yet unlocked); you walk each run's own Set up → Run → Analyze cycle, and the next run unlocks once the current one has a completed experiment. A locked run has no editable slots — it runs the instructor's design exactly as authored (your control), so it's one click to Run. Each run keeps its own slot values, notes, and run-specific questions (the required transcript notes count applies to each run separately); you can revisit an earlier run any time, and re-run a variant as often as you like — submitting freezes a snapshot of every run, and resubmitting replaces it. When every run is complete, the final Compare & submit stage opens: read your runs in tabs, take transcript notes on any of them with "+ Add note" (a note taken there belongs to that run, exactly as one taken on its Analyze step, and shows in the notepad and the attached-note pickers the same way), answer the comparison questions, add notes for your instructor, and submit — all runs are frozen together in one submission. If the assignment also offers "Start in builder", note that the copy it makes is the baseline design by itself: the declared runs live in the workspace, so the variants are not in it and the runs you do there do not count towards this assignment.

If your instructor built a guide into the assignment, a "📖 Guide" button floats at the top right of the workspace and of the assignment page, following the page as you scroll. Click it to slide open a drawer on the right with the instructor's step-by-step guidance, grouped under General, Set up, Run, and Analyze & submit, and numbered 1, 2, 3… straight through the groups so you and your instructor can name the same section. Every section stays visible on every step: whichever sections match the step you are on open automatically, the first General section (the overview) opens with them every time when the guide has one, and the rest are collapsed but one click away, so you can peek ahead or look behind whenever you want. That is on purpose, not a leak; the guide is reference material for the whole assignment, not a script that hides what comes next. Move to another step with the drawer open and it repaints for that step. Close it with Escape, the ✕, or by clicking outside. Guide edits show up immediately, so if your instructor updates it mid-assignment you'll see the latest version next time you open it.

Two ways to keep the guide where you want it. Expand in the drawer's header widens it, for a long section you would rather not read in a narrow column. Dock, in the workspace and in a freeform build's Experiment Builder, parks the guide beside your work instead of over it: the page makes room down the right-hand side, nothing is covered, and you can keep typing with the guide open. A docked guide ignores Escape and clicks outside on purpose, since neither of those is you asking to put it away; its ✕ closes it. The workspace and the builder each remember a docked guide for that assignment separately, and each opens it docked again the next time you come back, until you close it there with the ✕. Dock is offered on a wide screen only, and narrowing the window turns a docked guide back into the ordinary drawer over the page (or, while a panel is open above it, such as the builder's Designs or Data extraction drawer or the workspace's notepad, or while a walkthrough is running, closes it and keeps your choice for the next wide visit).

Good to know

  • Your work is saved automatically as you go — slots, notes, and answers survive a refresh or a switch to another device.
  • If your sign-in expires mid-work, the workspace tells you: a red "You've been signed out" banner appears, and everything you keep typing is held safely in your browser. Use the banner's Sign back in link (it opens a new tab), then return — saving resumes and your held work is pushed automatically. The same backup restores your unsaved work if the page crashes or your connection drops.
  • If the workspace ever opens on a copy held in your browser that differs from your saved copy (a save from that browser never reached the server), it asks you which copy to keep and shows both side by side: the step each is at, how many questions each has answered and in how many words, the notes on the transcript, and when each was saved. Nothing saves until you choose. The saved copy is what Delstorm has on record; the browser copy never replaces it on its own. The one exception is your own work from this browser over the very copy the server still holds (a save that failed on this page, say after a sign-out): that is restored and saved for you as before, with a "Restored unsaved work from this browser" note.
  • If Step 2 says a run "did not finish saving", the run itself is over but its record never reached the server (a sign-in that expired while it ran, say). Go back to Step 1 and run it again: your answers, notes and setup are kept, and the new run takes the old one's place.
  • You can't submit until you've marked the required number of moments and answered every question (with enough words, and with as many attached notes as the question asks for). The page tells you what's still missing.
  • The runs you launch here are your real attempt, not a practice mode. You can go back to Step 1 and launch again as often as you need before you submit; each launch runs a new experiment and your earlier attempts stay on record. To practice freely on a design of your own, use the experiment builder (see Classrooms).
  • Submitting freezes a snapshot for your instructor; you can revise and resubmit any time — see Classrooms. If your instructor allows a redo and reopens your analysis questions, a note at the top of the questions panel says so on your next load ("Your instructor reopened your analysis questions for a redo"): the reopened answers are cleared from your draft so you can answer them again in order, your earlier answers stay readable on your submission page, and you submit again when you are done.
  • Runs on the platform's shared providers draw on your token balance, shown under the run button; if a launch is refused because you are out of tokens, your instructor can top you up; see Tokens & Usage.

24. Authoring in the Assignment Studio (for instructors)

Where you edit: the Assignment Studio. Open any assignment and click "Edit in Studio" (/assignments/<id>/studio). The Studio is a sectioned workspace rather than one long form: a left-hand assignment map lists eight sections you move between freely (nothing is a locked wizard) — Purpose & brief (title, instructions, due date, and the resource links students see), Simulation design (attach a design), Student agency (guided mode, editable slots, starters), Runs & variants, Evidence & questions (analysis questions and the reveal policy), Student guide & support, Delivery & run details (run provider and model), and Preview & publish. Every field the assignment supports lives in one of these sections; the three that need an attached design say so and link you to Simulation design. Nothing saves until you press "Save assignment" in the header, which lists any sections with unsaved changes; leaving with edits pending prompts a warning. Publishing is its own terminal step in Preview & publish, where a readiness review separates blocking problems (things the server would reject) from advisory notes. On an already-published assignment, adding or removing a setup-gated question, or changing its identity, stage, or run scope, shows a drift warning: that structural change cancels answers students have already committed, so those attempts can no longer be submitted and the students re-run from the workspace. On a freeform build this is different: a before-the-run question added after a student's launch (a new question, or one newly moved to the before-the-run stage) does not cancel their runs; it is asked at their next launch or answered at hand-in, and the runs stay handable, while a wording or limit edit to a question they already committed to keeps that commitment and asks nothing again. Editing the wording, the minimum words, or the maximum words is always safe, and committed answers keep the exact wording the student saw. The same applies to the editable settings and the runs themselves: on a published assignment, clearing the scaffold, adding or removing an editable setting's run, or renaming, removing or restructuring runs cancels attempts students have in progress, so those runs can no longer be submitted and the students re-run from the workspace.

Anchoring questions to the transcript. Each analysis question can be asked at a position inside the run so students answer it as they read instead of all at the end. Set a round, then an "after message #" counted from 1: message 1 of a round appears once that round has started, message 2 after the second message, and so on. Leave "after message #" blank to ask at the end of the round (no message number). Leave the round blank for an end-of-run (Part 2) question. By default students unlock the transcript section-by-section — they type each anchored answer and choose "Submit your answer" to reveal the next section (nothing unlocks while they’re still typing); tick "show the whole transcript at once" to reveal everything immediately. Keep anchors shallow (the first message, the end of a round) so they hold across different students’ runs — a number deeper than a student’s run falls back to that round’s end.

When an assignment has a design attached (see Classrooms), you can turn it into a workspace: students run the design and answer questions in the focused three-step page. Runs, analysis questions, and the required-notes minimum are available whether or not you make any field editable; they reach students as soon as a design is attached. Making design fields editable (the Student agency section's "Editable design fields" panel) is a separate, optional choice that adds student-filled slots to Set up.

Setting it up

Editable design fields (optional)

In Student agency, tick "Students complete selected design fields themselves before launching" to let students fill in specific fields; leave it off and the design stays fully locked (students still get the workspace whenever you have added runs, questions, or a required-notes value). An assignment with a design but no workspace content at all (no editable fields, runs, questions, or required notes) instead uses the plain builder path: students click Start assignment and open the design in the Experiment Builder.

Start from a built-in starter (optional)

When the attached design ships with a ready-made guided setup, a "Start from a built-in starter" picker appears. Choose one and it fills in the editable slots, the required-notes count, and the analysis questions for you. For example, the intro Adam Smith assignment unlocks just two slots — a one-line experiment rule and a hypothesis — and adds five short analysis questions. Everything stays editable: review it, adjust, and save.

Set the run provider and model

In the Delivery & run details section, pick the Run provider and Model students (and your preview) use to run the workspace. Built-in course designs ship without one, so set this — otherwise launching shows a "no model set" message. Pick a shared provider so every student can run it; leave the provider unset to fall back to your account's default.

Choose the editable fields

You get a control for every field students can safely make editable, grouped as Prompts (the experiment name, each agent's persona or specialty, each round's condition prompt or custom instructions), plus Turn-taking, Length, and Memory and context for the interaction settings. Prompt fields stay text boxes; the interaction settings become typed widgets (a dropdown, a number field, or a checkbox). See Typed Editable Fields for what each becomes. Check the ones to unlock; everything else stays locked. For each checked field set a "Label students see" and an optional markdown hint — every editable field takes one, text boxes and typed widgets alike. The hint renders under the field's label in the workspace wherever the field is editable, and it doubles as feedback: for a text field it is what "Check my work" judges the student's input against (otherwise it judges against the label), and for a typed field with an expected answer it is what a student who misses sees. Text fields also take an optional "Minimum characters" and an "Allow pasting into this box" checkbox; both are text-only. Pasting is disabled in every editable text box by default, so what a student fills in is their own typing; check the box and students may paste into that field, for example when you want them to draft somewhere else first. Typed widgets have no paste rule to set.

Giving one run its own hint. On a multi-run assignment the hint you write here is the base hint: the base hint carries to every run unless you customize that run's hint. Every run row in Runs & variants carries a "Customize hints for this run" panel listing the fields you have ticked here, one editor each, and it is off until you use it: leave the editors alone and nothing extra is stored, so the assignment behaves exactly as it did before the option existed. Each editor starts out following the base hint and shows the base wording, so a later edit to the base still reaches it. The first change you type in an editor makes that field this run's own hint, even when the words end up identical to the base, and "Follow the base hint again" hands it back; clearing an editor does the same, and never hides the hint on that run, because hiding a hint on one run is not something the platform does. A run that customizes anything counts its customized hints beside the panel and opens its editors for you the next time you load the assignment. Unticking a field above drops its customizations from every run. A locked run still hides every hint from students, and you can still write hints on it: they show up if you unlock that run later. Students read the run's own hint where you wrote one and the base hint everywhere else, and "Check my work" quotes whichever hint the run they are on is showing them.

Letting one run allow pasting. Each run row also carries a "Customize pasting for this run" panel listing the editable text fields. Every field starts on "Follow the base setting from Student agency", so the field's own "Allow pasting" checkbox decides and nothing extra is stored. Choose "Allow pasting on this run" or "Block pasting on this run" and that run gets its own rule, whatever the base says. It is the shape Dan asked for: require typing on the Baseline, allow pasting on a later variant once students have drafted the text themselves. Typed widgets have no paste rule, so they are not listed.

Set the annotation requirement

"Required transcript notes" sets how many moments students must mark before they can submit (default 3; set 0 to make it optional).

Declare runs (baseline & variants) — optional

The "Runs" list turns the assignment into a multi-run lab: add up to six labeled runs (e.g. a Baseline and a Variant A), each with an optional markdown intro students see at the top of that run's Set up step. Tick "Locked" and the run has no editable slots — students run your design exactly as authored (the natural choice for a control Baseline). Students complete the runs in order — each unlocks when the previous one has a completed experiment — then a final Compare & submit stage brings their runs together. The editable fields apply to every non-locked run (students fill them per run), and "Required transcript notes" applies per run. Leave the list empty for a normal single-run assignment.

Add analysis questions

Build the typed questions students answer in Step 3. Each has its own text, an optional "Minimum words" and an optional "Maximum words", and you can reorder or remove them. Leave either at 0 for no limit; a maximum must be at least the minimum when you set both. A maximum never trims a student's answer: the whole text stays in the box, the counter reads their count against the limit ("18 / 50 words"), and the answer simply cannot be locked in or submitted until they shorten it. Each question also has an "Allow pasting" checkbox, off by default: unchecked, the answer box requires typing (pasting is disabled, so the answer is the student's own writing); checked, students may paste into that question's answer. Because before-the-run and analysis questions each carry their own checkbox, and each question on a multi-run assignment belongs to one run or to the final comparison, you can require typing in early runs and allow pasting in later ones. The set of questions is frozen onto each submission, so editing them later doesn't change work already handed in.

Formatting question text. Question text is written and displayed as markdown, the same as your instructions and the student guide. The small toolbar above each question box formats the words you have selected: B for bold, I for italics, • List for a bulleted list, and Callout to set a passage off in a highlighted block. You can also type the markdown by hand (**bold**, *italics*, - item, > callout). A Preview under the box shows exactly what students will see, and the same rendering appears in the workspace, in the "Committed before this run" box, and on the submitted work. Students have the same toolbar on every box they write prose in (bold, italics and list; no callout), so the work you review is rendered too: their answers, reasons, transcript notes and notes to you keep whatever formatting they applied, rather than arriving as raw syntax. Their design-slot entries are the exception and stay exactly as typed, since that text configures the experiment rather than being read as writing. Colour and font are deliberately not author-controlled: the callout and the emphasis styles follow the reader's theme, so a question stays legible in both light and dark mode.

When runs are declared, each question gets a Run dropdown: scope it to one run and students answer it during that run's Analyze step (a question anchored to a transcript position must pick a run), or leave it on "Comparison (after all runs)" to ask it in the final Compare & submit stage.

Asking for evidence. Every transcript message carries a copyable citation reference beside the speaker's name, such as R2.T3.Elena for round 2, turn 3, Elena (see Navigating a Transcript). Asking students to quote the reference alongside their claim gives you an answer you can trace straight back to the message it came from, and it tends to move answers from impressions to evidence. The reference travels as plain text: a student reads it off the chip (or copies it, where you allowed pasting), and a reference in an answer stays text rather than becoming a link, so you look it up in the transcript yourself. Worth saying in the question text if you want it, since students will not know to include it otherwise.

Asking for attached notes. Beside the word fields, every analysis question also takes a "Minimum attached notes" value. Leave it at 0 for no requirement, which is what every question that predates this feature has. Set a number and the student must attach that many notes from their notepad to that answer before it can be locked in or submitted, enforced the way a minimum word count is; the ceiling is 20, which is also the most notes one answer can carry. Attached notes are evidence rather than prose, so they never count toward that question's word limits: the student's own words still have to meet those on their own. The field is on after-the-run analysis questions, and on a freeform build the round evaluation takes one too (its minimum freezes with each run at launch). A before-the-run question cannot ask for attached notes, because its answer is committed at launch, when there is no transcript to take notes from yet. Like the word limits, the value is frozen onto each submission as the student submits, so editing it later never changes work already handed in. On a published assignment, raising or lowering it is safe for any answer a student can still edit: they attach one more note from the notepad and submit. An answer a staged reveal has already locked takes no more evidence, so a raised minimum leaves that student one way out: your Reopen answers on their submission review page (above), after which they answer again and resubmit.

Reading the evidence when you grade. Each attached note reads on the review page under the answer it belongs to, quote first and the student's comment beneath, with a jump into that note's own run of the transcript wherever the submitted snapshot still holds that moment. Where it does not, the block has no jump and says why instead: evidence from a run a later resubmit no longer has, or from a moment the stored transcript no longer holds, is still shown and labelled as such rather than dropped. If the student answered a question you have since deleted, that answer and the evidence attached to it are not part of the submission at all, since the answer goes with the question at submit.

Before-the-run (hypothesis) questions. Set a question's "When" dropdown to "Before the run (gates launch)" and it moves to Step 1: students see it under "Before you run" and the run cannot launch until it's answered to its minimum words, and within its maximum words if you set one — predictions are committed before results exist, enforced by the platform. This works on locked runs too (a hypothesis before an untouched Baseline), and in a multi-run assignment each before-the-run question is scoped to its run — so a Variant's hypothesis can be asked after the Baseline has been seen but before the Variant launches. Transcript anchors don't apply to before-the-run questions, and instructors previewing the workspace aren't gated. The answer given at launch is the commitment of record — the submission carries the launch-time answer even if the draft is edited afterwards, and a re-run is a fresh attempt with a fresh commitment. Committed answers stay visible as students work: a "Committed before this run" box sits above the transcript in the Run and Analyze steps, and Compare lists each run's predictions. After launch the Step 1 hypothesis box locks to the committed answer — "Revise for a new run" unlocks it, making clear the revision only applies to a fresh attempt. Each later run is also bound to the exact earlier attempt it launched after: re-running a Baseline after launching a Variant means submitting the original Baseline, or re-running the Variant. Instructors: on a published assignment, adding or removing a before-the-run question, or changing its identity, stage, or run scope, cancels answers students have already committed, so those attempts can no longer be submitted and the students re-run from the workspace. On a freeform build this is different: a before-the-run question added after a student's launch (a new question, or one newly moved to the before-the-run stage) does not cancel their runs; it is asked at their next launch or answered at hand-in, and the runs stay handable, while a wording or limit edit to a question they already committed to keeps that commitment and asks nothing again. Editing the wording, the minimum words, or the maximum words is always safe, and committed answers keep the exact wording the student saw.

Save the assignment and students will see "Open workspace". You get a "Preview workspace" button on the assignment page once the assignment has a workspace (guided or open); it opens exactly what they get. A banner marks it as instructor preview: you can test-run it with your own provider to check the assignment works, but nothing you enter is saved and you can't submit.

Typed Editable Fields (instructors)

The "Editable design fields" panel in Student agency is not limited to prompts. You can mark almost any interaction setting editable, so a guided assignment can teach mechanics (for example, "now choose Free Discussion and set three turns") instead of only prompt writing. Each field you tick becomes a typed widget for the student that accepts only valid values, so a student choice can never break the run.

How much to hand over is a dial, not a switch. The same panel covers the whole range, one assignment at a time. Locked is the default: nothing is ticked, the design runs exactly as you authored it, and students answer your questions. Selectively editable is the middle: tick the specific fields the assignment is about, and the rest stays locked and visible. Fully exposed is the far end: tick everything the panel offers, and students set up the whole design themselves inside your assignment. There is no separate mode to turn on for the last one; it is what the middle becomes when you tick every box. Students never choose their own editable set, on any setting of the dial.

The groups you can open

Prompts

The experiment name, each agent's persona and specialty, each round's name, condition prompt and custom instructions, and the moderator's completion criteria and speaker-selection prompts. These stay text boxes; a round name is a single-line box in the round's header.

Turn-taking

Interaction pattern and speaking order become dropdowns limited to the real options; turns becomes a number field (1 to 20); and the round's first speaker becomes a dropdown of the design's own agents plus a "none" option.

Length

Response length and length type become dropdowns.

Memory and context

Context depth, context overflow, and per-agent memory overflow become dropdowns; context window and per-agent memory window become number fields (0 to 50, where 0 means all past rounds); self reflections becomes a checkbox.

Moderator

Which of the design's agents moderates the round becomes a dropdown of the design's agents plus "none" (an agent-ref, like first speaker); the moderator's completion criteria and speaker-selection prompts stay text boxes.

Model behavior

Each agent's temperature becomes a number field over the same 0 to 1.5 range the Builder's own slider offers, in steps of 0.05. Lower is steadier wording, higher is more surprising.

Structure

A round's participating agents become a set of checkboxes over the design's own agents, where choosing none means every agent takes part. A round's transitions become one editable unit covering its condition rows (go to which round, and when), the agent who decides, the decision guidance, and what happens after the round. Transitions are deliberately not split into separate fields: a student can never end up with a condition whose text and target disagree, because the whole rule is edited together.

Round names are load-bearing here, because a transition points at a round by name. If a student renames a round, every condition pointing at it follows the rename, and a rename that would leave two rounds sharing one name is refused with a message telling them why.

Your "Show branch conditions to students" choice still applies to the rounds you left locked: their conditions show the target and who decides, but not the when-texts or the decision guidance. On a round whose transitions you made editable, students see the whole thing, because you have asked them to write those texts.

Agent-ref fields. First speaker and moderator are agent-ref fields: the student chooses one of the design's own agents, or none (both are optional). A chosen agent must be one of the design's agents, and a chosen moderator cannot also be one of the round's debating agents. Because a moderator is now student-selectable, a moderator-dependent choice (ending the round on the moderator, or letting the moderator pick speakers) is satisfied once the student assigns one.

What the widget accepts. A dropdown offers only the values the runner accepts, a number field enforces its range, an agent picker offers only the design's agents (or none), and a checkbox is a plain yes or no. "Minimum characters", a placeholder, and "Allow pasting" are text-only options: they apply to the prompt fields and are rejected on a number, dropdown, agent picker, or checkbox field, which have no character minimum and no paste rule.

Expected answer (optional). Every typed field you make editable gets an "Expected answer" control beside its label, except participating-agents and transitions fields, which have no single right answer to check against and say so where the control would be: the value you would have chosen. It is read only by "Check my work", which compares the student's choice to yours and answers with a pass or a fix chip. It is advisory only and does not block launching or submitting, and the feedback never shown to students as the answer: a student who misses it sees your field hint if you wrote one, otherwise a neutral nudge back to the guide. Leave it on "No expected answer" and the field is simply not checked. Text fields have no expected answer (there is no single right wording); they keep the on-topic check they always had.

Hint (optional). Every typed field you make editable also takes a markdown hint, exactly like a text field: it renders under the widget's label in the workspace wherever the field is editable, and it is the feedback a student sees on a missed expected answer. Write it as the guidance you would give in person ("debate fits a two-sided question"), not as the answer. On a multi-run assignment the base hint carries to every run unless you customize that run's hint: each run row in Runs & variants has a "Customize hints for this run" panel, off until you use it, whose editors follow the base hint until you edit one. Typed fields sit in that panel beside the text fields, so a dropdown, a number field, or a checkbox can carry its own wording on a single run, and a field whose base hint you left empty can still be given one there. A locked run hides every hint from students either way. Authoring in the Assignment Studio states the whole rule.

How students answer. Two independent checkboxes in Student agency set what you demand of the choosing itself, and both are off by default. "Students must actively choose each editable setting" opens every editable typed widget on "Choose..." rather than the design's value, so a student cannot run by leaving your defaults untouched; choosing the same value the design had still counts. "Students must justify kept preset values" adds a "Why this choice?" box under each editable setting, required only when the submitted value equals the design's, optional when the student changed it. Justifications are submitted with the work and appear on the submission review page beside the field they explain. Editable text boxes are exempt from both: they already demand real writing, and they now start from the design's authored text so students edit rather than retype it.

A guardrail on how much to open. Once four or more fields are editable, the Studio's readiness review adds an advisory note: broad editing can reduce comparability across students. It is a note, not a limit, and it never blocks publishing. If you would rather hand students the full Experiment Builder, that is a different assignment shape, not a variation on this one: attach a design and leave the assignment with no workspace content at all (no editable fields, no runs, no questions, no required-notes minimum). Students then click "Start assignment", the design imports into their own saved designs, and they build and run it in the Builder. Analysis questions are a workspace feature and do not follow them there, so an assignment on that path has no questions to answer; students submit a completed experiment instead.

What stays instructor-only. Provider and model selection stay yours: they are cost and access decisions, not experiment design. So do the raw structural collections, team assignments and agent order. A round's transitions and its participating agents are not in that list any more: both can be made student-editable, transitions as one unit. Design-level "Max rounds executed" and "Show branch conditions to students" also stay yours. Anything instructor-only runs as you set it in the design. To hand students full control of everything including the provider, attach a design with no workspace content and they open it in the Experiment Builder (see Authoring in the Assignment Studio).

The Design Editor (instructors)

When an assignment has a design attached, you edit that design right in the Studio's Simulation design section, with no Builder round-trip and no re-upload. Every agent and every round shows as a clickable line. Click one to open the design editor, a drawer that slides in from the right focused on that card. There are also "Add agent" and "Add round" buttons, and a header "Design editor" button that opens the drawer on the card you last worked on.

The drawer holds the same agent and round cards as the Experiment Builder, so every setting is editable there, plus a small header block for the design name, default provider, and default model. It is non-modal: the page stays usable behind it, the main content reflows to its left on a wide screen, and it goes full width on a narrow one. Closing the drawer with edits you have not applied asks you to confirm before discarding them.

One place saves

Apply, then save

The drawer footer has just two actions. "Apply to Studio" stages your edits into the page's unsaved state and closes the drawer; "Discard" throws them away and restores the last applied version. The drawer never saves on its own. The header "Save assignment" is the single save. If you press it while the drawer still holds unapplied edits, it stops and sends you back to Apply or Discard first.

Edits stage into this assignment only

Saving writes the edited design into this assignment's snapshot. It never changes the saved design you attached from. A structurally invalid design (no agents, no rounds, or a round listing an agent that does not exist) is rejected on save with a plain message. Launch-time concerns such as provider availability are still settled at launch, not here.

Divergence and re-attaching

Because studio edits touch only the snapshot, the design section labels the difference. After you save studio edits it reads "Edited in Studio, diverged from <source design name>", or "Edited in Studio (source design deleted)" when a recorded source design no longer exists, or simply "Edited in Studio" when the assignment has no source design on record, which is how every copied assignment starts. Re-attaching a design from the "Attach design" select replaces the studio edits and clears the label. While a re-attach is chosen but not yet saved, the design editor is locked until you Save assignment, so the design about to be replaced cannot be reopened over the new choice.

Saving a copy to your Designs

After a successful Save assignment, a "Save a copy to my Designs" button appears. It writes the edited design into your own Designs, named with a " (Assignment Studio Build)" suffix. There is exactly one copy per assignment: the first save mints it, and later copy saves update that same design in place rather than minting a second one. The copy identity survives re-attaching a different design. Saving a copy re-syncs the assignment's source to that copy, and the diverged label clears once the snapshot matches the copy.

The student path diagram

Preview & publish carries a read-only "Student path" diagram, and a "View student path" link in Simulation design jumps to it. It is a derived projection of the current design, guided scaffold, declared runs, and analysis questions: it owns no editable state and never writes back, so it always reflects what you have staged. It shows, in order: before-the-run (setup) questions, then run lanes that each walk Set up then Run then Analyze (one lane for a normal assignment, or one lane per declared run), with each round's interaction pattern, moderator, and speaker-order badges. Questions anchored after a round or after a specific message appear at that spot, with message numbers shown starting at 1. A multi-run assignment rejoins at a Compare stage, then Submit. Whether students unlock the transcript step by step (gated) or see it all at once (open reveal) is shown, and a plain linear text list of the same path always ships beside the diagram for screen readers. In an incomplete design the diagram draws only the parts it knows and marks the broken step, never inventing placeholder rounds or agents.

Letting students see the path

A checkbox in Preview & publish, "Students can open the Assignment Studio for this assignment", controls the student-facing view. It is off by default, so students see no such button and reach nothing; instructors always keep full studio access. Turn it on and a student who opens the assignment studio gets only this read-only path view, never the authoring drawer, the prompt library, any write control, or the raw design. It is a separate student page, not the instructor Studio with parts hidden.

25. Tutorials

The Tutorials page — one of the three student front doors — holds short, self-paced lessons that teach the platform one step at a time. They're the gentlest place to start before diving into the full builder.

How they work

  • Deterministic, not an AI chat. A lesson mixes short things to read with "type the answer" checks. A wrong answer just shows a hint and lets you try again — no model is involved, so it's instant and consistent.
  • Pal hosts the lessons. Pal opens each lesson with a short intro, cheers a right answer, encourages after a wrong one, and offers the next lesson when you finish. The hosting is pre-written script, never generated, so tutorials stay free and never call a model.
  • Modules unlock independently. Start any module right away; inside a module, lessons run in order. Each module card shows how many of its lessons you have done. Your progress is saved to your account when you're signed in (so it follows you across devices) and in this browser otherwise.
  • Completion is earned in the sitting. Marking a lesson complete needs that lesson's answer checks passed in the same sitting. If you reload mid-lesson, redo its short checks and finish again; lessons are short on purpose.
  • Self-paced. Re-enter any lesson any time; completing a lesson again never changes its original completion date.

The four modules

Ordered from the basics up: Build (create an agent, write its prompt, add a round, edit round features, and advanced round controls), Run (run a one-round experiment), Analyze (export, annotate, and analyze the transcript), and Submit (turn in an assignment). All lessons are live across the four modules.

Tutorial assignments & certificates

Instructors can assign the tutorials as a classroom assignment. In the Assignment Studio, turn on Tutorial certificate in the Student agency section and pick the required lessons (all selected by default; picking a lesson includes the earlier lessons of its module). The assignment page then shows students each required lesson with its completion state and date, and a "Hand in my certificate" button that enables once every required lesson is complete.

  • Earlier work counts. Lessons finished before the assignment existed satisfy it, and the certificate cites the real original dates.
  • The certificate is the submission. Handing in freezes a certificate of completion: your name as you provided it, your account email, and each required lesson with its completion date. It is built by the server from your saved progress; there is nothing to type. The instructor reviews and grades the certificate like any submission, and a printable page is linked from the assignment after you hand in.
  • Handing in again issues a fresh certificate. The lesson dates stay the original facts; the name and the hand-in date are current.
  • Verification is by account. The name on the face is the one the student provided (their classroom roster name, or their account display name); the account email printed beside it is what an instructor checks a printed copy against.

26. Tips & Best Practices

Start small, then scale up

Test with 2 agents, 1 round, short responses first. Once your design works, add more rounds and agents, and switch to dynamic length so each agent sizes its turn to the moment. This saves time and API costs during iteration.

Invest in personas

The single biggest factor in experiment quality is agent persona specificity. Use document grounding to create rich initial personas, then edit them to sharpen the perspective. A 5-sentence persona produces dramatically better results than a 1-sentence one.

Design round progression

Structure rounds to escalate: Round 1 for initial positions, Round 2 for challenges and rebuttals, Round 3 for synthesis and conclusions. This mirrors real discourse structure and produces more nuanced results than a single long round.

Vary temperatures across agents

A conservative agent (temp 0.3) paired with a creative one (temp 0.8) produces more interesting discourse than agents at uniform temperatures. The conservative agent grounds the conversation while the creative one introduces novel perspectives.

Save designs before launching

Always save your design before running an experiment. If the experiment produces unexpected results, you can reload and adjust without re-entering everything.

Use batch runs for research

If you're studying how agents interact under specific conditions, run the same design 3-5 times with short responses to observe variance before committing to a thorough single run. This helps you identify which experimental designs are worth investing in.

Mix interaction patterns across rounds

Start with Blind Parallel (to capture uninfluenced baseline positions), then switch to Free Discussion (to let agents interact and update their views), then finish with a Debate (to force a structured conclusion). This design reveals how discourse changes agents' positions.

27. Glossary

Agent

An AI participant in an experiment, defined by a persona prompt and backed by an LLM.

Batch Run

Running the same experiment configuration multiple times to observe variance.

Condition Prompt

The text that sets the social scenario for a round — the situation agents are placed in.

Confidence

In evaluation, indicates how a field value was determined: "parser" (regex extraction), "llm" (LLM inference), "both" (confirmed by both), "none" (not found), "not_executed" (on a branching run, the round this field points at did not run), or "recorded" (taken from the routing decision Delstorm recorded on a branching run, never read from the transcript).

Contents Sidebar

The docked table of contents on completed transcripts. Jump to any round, message, or analysis question; in a gated assignment, sections past the current question show as gray 🔒 rows.

Context Depth

Per round, how much of each earlier message agents see: Summary (first 300 characters) or Full (untruncated). Distinct from Context Window, which sets how many past rounds are visible at all.

Data Extraction

Structured fields defined in the builder and pulled out of every run automatically by the evaluator, then exportable as CSV. The design-time counterpart to the on-demand Evaluate page.

Design

A saved experiment configuration (agents, rounds, settings) that can be loaded and reused.

Document Depth

Controls how much of a referenced document agents see: Summary (extraction only), Relevant Excerpts (keyword-matched paragraphs), or Full Text (entire document).

Document Library

Persistent storage for uploaded documents. Documents are analyzed once and can be referenced from any experiment's rounds.

Evaluation

Post-hoc analysis of experiment transcripts using defined fields. Combines deterministic parsing with LLM inference.

Evaluation Field

A structured metric to extract from a transcript: number, text, scale, boolean, or choice.

Evaluation Template

A saved set of evaluation field definitions that can be reused across experiments.

First Speaker

An optional control (for the Sequential and Random speaking orders) that names which agent opens a round. Left blank, the round leads with the first agent in your list.

Hybrid Evaluator

The system that runs deterministic parsing first, then LLM analysis, merging results with confidence tracking.

Interaction Pattern

How agents communicate within a round: free discussion, debate, chain, blind parallel, or custom.

Moderator

A dedicated agent that guides a round — opening it, commenting between turns, and closing it — without participating in the debate. Can also decide when the round ends.

Persona

The system prompt that defines an agent's identity, beliefs, and communication style.

Provider

An LLM backend (e.g., Claude, Ollama, OpenAI) that powers agents. Configured in Settings.

Round

A phase of an experiment defined by a condition prompt, interaction pattern, and settings.

Self-Reflection

A private note an agent writes after a round. Never shown to other agents; an agent's own reflections carry into its later rounds.

Sequential

The default speaking order for Free Discussion and Custom rounds: agents speak one at a time in your agent-list order, each seeing the full conversation so far. Distinct from the Chain interaction pattern, which has its own fixed build-on-the-previous order.

Speaking Order

The round-card control that sets who speaks next in a Free Discussion or Custom round: Sequential (default), Round-robin (all at once), Random, Handoff, or Moderator.

Temperature

An LLM sampling parameter that controls response variability. Lower = more deterministic, higher = more creative.

Turn

One exchange within a round where all participating agents respond. Multiple turns create back-and-forth conversation.

28. Branching & Transitions

By default an experiment runs your rounds top to bottom, once each. Transitions let a round decide what happens next based on what actually happened in it, so one design can take different paths on different runs.

Setting up conditions

Every round card in the Experiment Builder has a Transitions section. Under Conditional transitions, add one condition per possible destination: pick the target round by name, and write a When text describing the circumstances that should send the experiment there ("the Trajectory Score is 3 or above"). Rounds are referenced by their Round name, so a branching design needs its target rounds named. Renaming a round updates every condition that points at it; deleting one asks first and clears the conditions that pointed at it.

Once a round has at least one condition, choose a transition moderator: the agent that reads the round and decides. It can be the round's own moderator, an agent that participates in other rounds, or a dedicated judge; its only extra duty is the decision, and it does not join the round's turn-taking. Decision guidance is optional overall instruction for that judgement ("the outcome depends on practices, not on who wins").

The condition texts and the decision guidance are never shown to the agents taking part. The judge's public announcement is, so it may paraphrase the reason it chose.

After this round

After this round sets what happens when no condition is met, and what happens on a round with no conditions at all. Continue to the next round is the default (and off the last round, the experiment simply ends). End the experiment stops the run there, which is what a branch tail usually wants: without it, the round after "The Iron Cage" in your list would run even though the experiment took the other path. Jump to a round goes somewhere specific every time.

When a round has conditions, the moderator is asked to state a verdict for every condition before it decides: met or not met, with one sentence of evidence from the discussion. When its reply carries those verdict lines and names a round, that decision is followed only if its own verdicts support it and every verdict line could be read; otherwise the named decision is not used for routing, After this round applies, and the announced reason is composed from the verdicts (the same happens when the moderator chooses no condition while marking one met, or names something that is not an option). When the moderator chooses no condition and nothing contradicts it, After this round applies with the moderator's own reason, and if some verdict lines could not be read the transcript says how many. A reply with no verdict lines at all is handled exactly as before: the decision is followed when it names a condition's round, so older designs and models that ignore the request are not broken. The verdicts show under the announcement for everyone who can see the transcript. The moderator is told exactly where the fallback leads before it decides, and the transcript names the destination either way. A decided move reads "Moving to The Iron Cage: ...", while a fallback reads Continuing to "..." (or "Ending the run") followed by the reason, so a run that argues for one branch and takes another is visible at a glance. If the moderator's reply never reached a readable decision, the note says so rather than showing an empty reason.

Loops and the round ceiling

A condition may point backward, or at its own round, so a design can cycle. To keep an unattended run finite, every branching run has a ceiling on the total number of rounds it will execute. Max rounds executed in Experiment Info sets it; leave it blank for the automatic ceiling (three times the number of rounds you authored). The platform caps it at 150 whatever you enter. When a run stops because it hit the ceiling, the transcript says so: "Run ended at its round ceiling."

If a round cannot be reached from the first round by any path, the Studio's readiness review flags it as an advisory. It does not block saving or publishing: an unreachable round may be work in progress.

Writing conditions that route reliably

The transition moderator reads your When texts word for word, is asked for a verdict on every condition, and must then answer with one exact round name, or NEXT when none of them fits. Three habits make that answer land where you meant:

Cover the "not yet" case explicitly. In a cyclical design (say, a Boom and Bust loop that may or may not ripen into a Revolution), the verdict "conditions have not formed yet" deserves its own condition pointing back into the cycle ("If the workers remain fragmented, go to Boom"). A moderator whose honest assessment matches one of your conditions can name that round; a moderator whose assessment falls between conditions is left with NEXT, and NEXT follows After this round, wherever that points.

Point After this round at the safe default. In a cycle, that is usually the earlier round: set the looping round's After this round to Jump to a round and pick the round that repeats the cycle, so an undecided reply repeats the loop rather than exiting it. Left on "Continue to the next round", an undecided reply walks forward into whatever comes next in your list, which in a cycle-then-climax design is exactly the round the moderator just argued the run was not ready for. The Studio's readiness review flags this shape as an advisory: a round with a branch that re-enters a cycle, however the rounds are ordered on the page, while its After this round leaves that cycle.

Phrase conditions as observable evidence, and put thresholds in Decision guidance. "At least two workers explicitly call for seizing production" routes more reliably than "revolutionary consciousness has formed", because the moderator can check it against the transcript rather than interpret it. Where several conditions share a standard, define it once in Decision guidance ("Reformist demands, however organized, are not revolutionary action").

What students see

In an assignment, Show branch conditions to students controls whether the assignment path view spells out the When texts and the decision guidance. It is on by default, because handing students the conditions is usually part of the exercise. Turned off, students still see which rounds a round can lead to, labelled "decided during the run", and the stored condition texts are withheld from their path view; when a round's routing is a student-editable slot the workspace shows the conditions regardless, because the editing widget needs them. What the moderator writes in public is a different channel: its announcement and its verdict evidence are model-written from the discussion and instructed not to restate the conditions, which is an instruction rather than a guarantee.

Reading a branched transcript

Rounds appear in the order they ran, numbered 1, 2, 3 by execution. When the executed order differs from the order you authored, the heading also names the authored position, for example "Round 3: The Bust (design round 2)", and a round that ran more than once is labelled ", second pass (design round 2)". The judge's decision appears between rounds as its own note. Rounds that were not taken are simply absent.

Anchored analysis questions are authored against a design round, so a question attaches to the first time that round actually ran. If the round never ran, the question still appears, at the end of the transcript, marked "This question was set for a round that did not run in this experiment." A gated assignment treats such a question as passed rather than locking the rest of the transcript behind a round that cannot exist.

Extracted data works the same way: a field sourced from a design round reads every execution of that round, and a field pointing at a round that never ran comes back empty with the not_executed confidence badge. Instructors reviewing a submission can open Routing assessment beside the announcement to read the judge's full private reasoning; students never receive it. On your own personal runs, Show routing assessments on the results page fetches the same thing for you.

Ready