# Welcome to Factory Install Droid, delegate your first task, and scale from a single session to an automated SDLC. ## Choose a path Start with a coding task, automate a workflow, or research the platform. Pick the outcome that aligns with your goals. Delegate a task, review the diff, and merge from the App or your terminal. Put code review, QA, and documentation on autopilot across your repos. Deploy, secure, govern, and observe Droid across your organization. ## Start with a surface Droid runs the same everywhere. Pick a surface, install it, and start your first session. Download the desktop app, sign in, and connect it to your codebase. - [Mac (Apple Silicon)](https://app.factory.ai/api/desktop?platform=darwin&architecture=arm64): Download for M-series Macs - [Mac (Intel)](https://app.factory.ai/api/desktop?platform=darwin&architecture=x64): Download for Intel Macs - [Windows (x64)](https://app.factory.ai/api/desktop?platform=win32&architecture=x64): Download for Intel and AMD PCs - [Windows (ARM64)](https://app.factory.ai/api/desktop?platform=win32&architecture=arm64): Download for Snapdragon and ARM PCs Follow the [Factory App quickstart](/factory-app/quickstart) to install the app, connect a repository, and run your first reviewable session. Run one command, then start `droid` in any repository. **macOS/Linux:** `curl -fsSL https://app.factory.ai/cli | sh` **Homebrew:** `brew install --cask droid` **Windows:** `irm https://app.factory.ai/cli/windows | iex` **npm:** `npm install -g droid` Follow the [Droid CLI quickstart](/droid-cli/quickstart) to install the CLI, authenticate, and delegate your first task from the terminal. Nothing to install: [cloud session sync](/droid-cli/settings#cloud-session-sync) and [Droid Computers](/droid-computers/overview) make your sessions available from any browser at app.factory.ai. The [Factory App quickstart](/factory-app/quickstart) covers the same session flow across desktop, web, and mobile. ## Explore the docs - [Product Surfaces](/factory-app/overview): run Droid in the Factory App, the CLI, headless exec, and cloud computers. - [Software Factory](/software-factory/overview): automate triage, code review, QA, docs, and incident response. - [Platform Capabilities](/missions/overview): Missions, model independence, autonomy controls, and harness customization. - [Factory Enterprise](/enterprise): deployment patterns, security, identity, and telemetry for a managed fleet. - [API Reference](/api-reference): drive sessions, computers, wikis, and analytics programmatically. - [Changelog](/changelog/release-notes): release notes and feature maturity. # Individual Plans Compare Factory Pro, Plus, and Max plans for individual developers. Individual plans are for developers who own their own Factory subscription. Pro, Plus, and Max include the Factory App, Droid CLI, and Droid SDK, with usage governed by rolling Rate Limits. ## Individual tiers ## How individual usage works {/* sweep-allow: term-bullets */} - **Rate Limits:** three rolling windows, 5-hour, 7-day, and 30-day. Each runs independently. - **Standard Usage:** consumed first on Individual plans. - **Droid Core:** free open-weight model pool with separate Rate Limits after Standard Usage runs out. - **Extra Usage:** prepaid credits, $10 minimum, no expiration. The toggle stays on until you turn it off or run out. - **Monitor usage:** open **Settings > Usage** or run `/limits` in Droid. ## Missions and BYOK - **Missions** use the same rolling Rate Limits as regular sessions and require Extra Usage to be enabled. If a Mission hits a Rate Limit, it pauses. - **BYOK** is free up to an allowance on all Individual plans. After that, usage is charged according to your specific plan. Run Missions on Droid Core when possible. Core models often deliver comparable results and stretch your plan further. ## FAQ Droid sessions consume Factory Standard Credits based on model usage, along with compute usage for Droid Computers. Yes. Hitting the 5-hour cap does not eat into weekly or monthly any faster, and vice versa. You need headroom in all three to send a request. All three are rolling: 5-hour, 7-day, and 30-day. Each starts ticking when you first use Droid and resets after the window elapses. No. Standard Usage does not roll over month to month. Extra Usage that you purchase does roll over. A set of models Factory has designated as Core, typically smaller, cheaper models that handle most coding tasks well. You can see the current list in **Settings > Models**. The list changes as new models ship. No. It uses the same Standard Usage billing, restricted to the Core model pool so you can keep working after premium-model Rate Limits are exhausted. Extra Usage is prepaid, USD-denominated credit. Buy a dollar amount and it is drawn down as your sessions use models. The current request finishes, then subsequent requests receive a Rate Limit error in the CLI or Factory App. Run `/limits` to toggle Droid Core or Extra Usage and retry immediately. Compare Teams, Business, and Enterprise plans. Browse the model catalog and multipliers that drive usage costs. # Organization Plans Compare Factory Teams, Business, and Enterprise plans for shared usage, governance, security, and support. Teams is self-serve and includes Pro Rate Limits for each seat. Business and Enterprise plans are for teams and companies that need shared usage, administrative controls, onboarding, security terms, and support. Business and Enterprise use custom limits and contract terms rather than the Individual rolling Rate Limit model. ## What organization plans add | Capability | Teams | Business | Enterprise | | :--------- | :---- | :------- | :--------- | | Seats | Up to 10 | Up to 150 | Unlimited | | Shared usage limits | — | Custom limits for team usage | Custom limits for enterprise usage | | Identity and provisioning | — | SSO, SAML/SCIM provisioning | SSO, SAML/SCIM provisioning | | Governance | — | Admin controls for models, autonomy, access, deny lists, and network policy | Full admin controls, sub-organizations, plus custom policy and deployment terms | | Data controls | — | ZDR, audit logging and activity trails | ZDR, audit logging, customer-managed encryption keys, data residency | | Deployment | — | Factory-hosted and managed compute options | Dedicated compute, partitioned inference, and on-premise options | | Support | — | Dedicated onboarding and support | Dedicated Account Manager, Customer Engineer, and priority SLAs | ## Usage and billing On Teams plans, each seat includes Pro Rate Limits. Billing is centralized and paid month-to-month. On Business and Enterprise plans, usage is shared across the workspace and governed by contract terms. It is not governed by the Individual plan Rate Limit model. ## Enterprise Controls Business and Enterprise add controls for platform teams: - [Identity and Access](/enterprise/identity-and-access) for roles, SSO, and provisioning. - [Hierarchical Settings and Org Control](/enterprise/hierarchical-settings-and-org-control) for inherited settings and org-managed policy. - [Network and Deployment](/enterprise/network-and-deployment) for deployment choices and network boundaries. - [Telemetry & Analytics](/enterprise/telemetry) for usage, adoption, and cost reporting. ## FAQ On the Teams plan, each seat has its own individual Pro Rate Limits. Business and Enterprise plans have custom workspace-level terms. Limits, seats, support, and deployment options are scoped to the organization contract. Choose Enterprise when you need unlimited seats, dedicated compute, partitioned inference, on-premise deployment, sub-organizations, customer-managed encryption keys, data residency, or priority support commitments. Discuss Business or Enterprise. Security, deployment, governance, and observability features for Enterprise. # Available Models Models offered natively in the Factory platform. The **Reasoning** column lists the values each model accepts for `reasoningEffort` in [settings](/droid-cli/settings#reasoning-effort) and for `--reasoning-effort` on the command line. The model selector labels them differently, so `xhigh` appears there as **Extra High**. ## Anthropic logoAnthropic | Model | Model ID | Multiplier | Reasoning | | --- | --- | --- | --- | | Claude Fable 5.1\* | `claude-fable-5.1` | 4× | `off`, `low`, `medium`, `high` (default), `xhigh`, `max` | | Claude Fable 5\* | `claude-fable-5` | 4× | `off`, `low`, `medium`, `high` (default), `xhigh`, `max` | | Claude Opus 5 | `claude-opus-5` | 2× | `off`, `low`, `medium`, `high` (default), `xhigh`, `max` | | Claude Opus 5 Fast | `claude-opus-5-fast` | 4× | `off`, `low`, `medium`, `high` (default), `xhigh`, `max` | | Claude Opus 4.8 | `claude-opus-4-8` | 2× | `off`, `low`, `medium`, `high` (default), `xhigh`, `max` | | Claude Opus 4.8 Fast | `claude-opus-4-8-fast` | 4× | `off`, `low`, `medium`, `high` (default), `xhigh`, `max` | | Claude Opus 4.7 | `claude-opus-4-7` | 2× | `off`, `low`, `medium`, `high` (default), `xhigh`, `max` | | Claude Opus 4.6 | `claude-opus-4-6` | 2× | `off`, `low`, `medium`, `high` (default), `max` | | Claude Opus 4.5 | `claude-opus-4-5-20251101` | 2× | `off` (default), `low`, `medium`, `high` | | Claude Sonnet 5 | `claude-sonnet-5` | 0.8× | `off`, `low`, `medium`, `high` (default), `xhigh`, `max` | | Claude Sonnet 4.6 | `claude-sonnet-4-6` | 1.2× | `off`, `low`, `medium`, `high` (default), `max` | | Claude Sonnet 4.5 | `claude-sonnet-4-5-20250929` | 1.2× | `off` (default), `low`, `medium`, `high` | | Claude Haiku 4.5 | `claude-haiku-4-5-20251001` | 0.4× | `off` (default), `low`, `medium`, `high` | \* Anthropic requires all Mythos-class models comply with 30 day data retention for trust and safety, please see Anthropic's [data retention practices for Mythos-class models](https://support.claude.com/en/articles/15425996-data-retention-practices-for-mythos-class-models). Organization admins opt in to the Anthropic Data Retention Policy in model access settings. ## OpenAI logoOpenAI | Model | Model ID | Multiplier | Reasoning | | --- | --- | --- | --- | | GPT-5.6 Sol | `gpt-5.6-sol` | 2× | `none`, `low`, `medium` (default), `high`, `xhigh`, `max` | | GPT-5.6 Sol Fast | `gpt-5.6-sol-fast` | 4× | `none`, `low`, `medium` (default), `high`, `xhigh`, `max` | | GPT-5.6 Terra | `gpt-5.6-terra` | 0.8× | `none`, `low`, `medium` (default), `high`, `xhigh`, `max` | | GPT-5.6 Luna | `gpt-5.6-luna` | 0.08× | `none`, `low`, `medium` (default), `high`, `xhigh`, `max` | | GPT-5.5 | `gpt-5.5` | 2× | `low`, `medium` (default), `high`, `xhigh` | | GPT-5.5 Fast | `gpt-5.5-fast` | 5× | `low`, `medium` (default), `high`, `xhigh` | | GPT-5.5 Pro | `gpt-5.5-pro` | 12× | `medium` (default), `high`, `xhigh` | | GPT-5.4 | `gpt-5.4` | 1× | `low`, `medium` (default), `high`, `xhigh` | | GPT-5.4 Fast | `gpt-5.4-fast` | 2× | `low`, `medium` (default), `high`, `xhigh` | | GPT-5.4 Mini | `gpt-5.4-mini` | 0.3× | `low`, `medium`, `high` (default), `xhigh` | | GPT-5.4 Mini Fast | `gpt-5.4-mini-fast` | 0.6× | `low`, `medium`, `high` (default), `xhigh` | | GPT-5.3-Codex | `gpt-5.3-codex` | 0.7× | `low`, `medium` (default), `high`, `xhigh` | | GPT-5.3-Codex Fast | `gpt-5.3-codex-fast` | 1.4× | `low`, `medium` (default), `high`, `xhigh` | | GPT-5.2 | `gpt-5.2` | 0.7× | `off`, `low` (default), `medium`, `high`, `xhigh` | ## Google logoGoogle | Model | Model ID | Multiplier | Reasoning | | --- | --- | --- | --- | | Gemini 3.1 Pro | `gemini-3.1-pro-preview` | 0.8× | `low`, `medium`, `high` (default) | | Gemini 3.7 Flash | `gemini-3.7-flash` | 0.3× | `low`, `medium`, `high` (default) | | Gemini 3.6 Flash | `gemini-3.6-flash` | 0.6× | `low`, `medium`, `high` (default) | | Gemini 3.5 Flash | `gemini-3.5-flash` | 0.6× | `minimal`, `low`, `medium`, `high` (default) | | Gemini 3 Flash | `gemini-3-flash-preview` | 0.2× | `minimal`, `low`, `medium`, `high` (default) | Promotional pricing. Gemini 3.7 Flash bills at 0.3× through January 1, 2027, then returns to 0.6×. ## xAI logoxAI | Model | Model ID | Multiplier | Reasoning | | --- | --- | --- | --- | | Grok 4.6 | `grok-4.6` | 0.8× | `low`, `medium`, `high` (default), `xhigh` | | Grok 4.5 | `grok-4.5` | 0.8× | `low`, `medium`, `high` (default) | ## Factory logoDroid Core (Open Models) | Model | Model ID | Multiplier | Reasoning | | --- | --- | --- | --- | | Inkling | `inkling` | 0.4× | `off`, `minimal`, `low`, `medium`, `high` (default), `xhigh`, `max` | | GLM-5.3 | `glm-5.3` | 0.56× | `low`, `high`, `max` (default) | | GLM-5.2 | `glm-5.2` | 0.56× | `off`, `high` (default), `max` | | GLM-5.2 Fast | `glm-5.2-fast` | 0.84× | `off`, `high` (default), `max` | | Kimi K3 | `kimi-k3` | 1.2× | `off`, `low`, `high` (default), `max` | | Kimi K2.7 Code | `kimi-k2.7-code` | 0.38× | `off`, `high` (default) | | Kimi K2.6 | `kimi-k2.6` | 0.4× | `off`, `high` (default) | | Nemotron 3 Ultra | `nemotron-3-ultra` | 0.24× | `off`, `high` (default) | | DeepSeek V4 Flash 0731 | `deepseek-v4-flash-0731` | 0.176× | `off`, `low`, `high` (default), `max` | | DeepSeek V4 Pro | `deepseek-v4-pro` | 0.528× | `off`, `low`, `high` (default), `max` | | MiniMax M3 | `minimax-m3` | 0.12× | `high` (default) | | MiniMax M2.7 | `minimax-m2.7` | 0.12× | `high` (default) | | Kimi K2.5 | `kimi-k2.5` | 0.25× | `off`, `high` (default) | | GLM-5.1 | `glm-5.1` | 0.55× | `off`, `high` (default) | Deprecated. MiniMax M2.7, Kimi K2.5 and GLM-5.1 remain available for now and will be removed in a future release. ## Custom models Configure custom models through [Custom Models (BYOK)](/model-independence/byok). Configure custom model endpoints and bring your own API keys. Let Factory route requests across providers automatically. # Contact Support Find the right channel to get help with Factory, depending on what you need. If you need help with Factory, the team and community are here for you. Pick the channel that matches your request so you reach the right people fast. ## Choose a channel Account, billing, and technical issues that need the Factory team. Questions, feedback, and conversation with the community and team. Reproducible bugs and feature requests, tracked in the open. ## Email support Use email for anything specific to your account that should not be public: - Account access, organization, and membership - Billing, invoices, and plan changes - Technical problems that need someone to look at your account Reach the team at [support@factory.ai](mailto:support@factory.ai). ## Community on Discord - How-to questions and usage tips - Product feedback and ideas - Connecting with other users and the team [Open the Factory Discord invite](https://discord.gg/zuudFXxg69). ## Bugs and feature requests - Reproducible bugs in Droid - Feature requests and enhancement ideas Search the existing issues first, then open a new one if your report is not already there. ## Write a report we can act on Whether you email us or open a GitHub issue, a few details help us resolve it faster: {/* sweep-allow: term-bullets */} - **What happened** and what you expected to happen instead - **Steps to reproduce**, including the exact command or action - **Version and environment**: operating system, terminal or IDE, and the Droid version - **Evidence**: relevant error messages, logs, or screenshots - **Impact**: how often it happens and whether it blocks your work Run `droid --version` to confirm which version you are on before reporting a bug. Found a security vulnerability? Do not post it publicly. Report it privately to [security@factory.ai](mailto:security@factory.ai). Running Droid in a regulated or enterprise environment? See Factory for Enterprise. # Factory App Use the Factory App for a visual Droid workspace with local machine access, project switching, and reviewable sessions. Use the Factory App when you want a visual workspace for Droid sessions, project switching, model controls, machine connections, and review state. Preview whatever Droid makes, from code to documents to live sites, and comment right where a change belongs. The app keeps local, web, and mobile review surfaces connected to the same Factory runtime. ## When to use the app - You want to review Droid sessions, diffs, and project state outside a terminal. - You switch between projects or machines often. - You need a guided path to local machine access and Factory-managed cloud computers. - You want web and mobile review surfaces for work that starts on your workstation. ## What the Factory App brings to your workflow The app is a dedicated workspace with Droid built in, organized around reviewing and refining whatever a session produces. Documents, presentations, spreadsheets, PDFs, live websites, and full code diffs render beside the session that produced them. Design mode selects an element and says what to change; inline code review comments on a diff line by line. Droid acts on the feedback in place. Work directly on your local filesystem with your full development environment, informed by project context, documentation, and team knowledge. Filter and group sessions across projects and machines, synchronized with Jira, Notion, Slack, Linear, PagerDuty, and MCP tools. Connect Slack, Linear, and other tools. The terminal-first surface for the same platform. # Factory App Quickstart Install the Factory App, connect your codebase, and start your first reviewable Droid session. Use this quickstart to install the Factory App, connect a codebase, and start a first reviewable Droid session. ## Step 1: download and install - [Mac (Apple Silicon)](https://app.factory.ai/api/desktop?platform=darwin&architecture=arm64): Download for M-series Macs - [Mac (Intel)](https://app.factory.ai/api/desktop?platform=darwin&architecture=x64): Download for Intel Macs - [Windows (x64)](https://app.factory.ai/api/desktop?platform=win32&architecture=x64): Download for Intel and AMD PCs - [Windows (ARM64)](https://app.factory.ai/api/desktop?platform=win32&architecture=arm64): Download for Snapdragon and ARM PCs - **Mac**: Open the `.dmg` file and drag Factory to Applications - **Windows**: Run the installer Launch the Factory App and sign in with your account. ## Step 2: connect to your codebase Set your working directory to your project folder. The Factory App connects directly to your local machine. Choose your preferred AI model from the model selector. You can change this anytime. ## Step 3: start a reviewable session Start with a context-gathering request: ```text Analyze this codebase and explain the architecture, key entry points, and test setup ``` Then ask for one small, verifiable change: ```text Add error handling to the user authentication flow and show me the diff before applying it ``` Droid analyzes your codebase, proposes changes, and shows what will be modified before applying anything. ## Step 4: connect your engineering systems The Factory App becomes more useful when Droid can read the systems where work starts. Connect team context through: Connect tools and internal services through Model Context Protocol servers. Let Droid read issues and turn product requests into reviewable implementation work. Pull thread context into triage, incident response, and follow-up tasks. ## Step 5: scale into a Software Factory Once Droid is part of your daily work, move repeatable steps into persistent automations. Software Factory connects agents across your delivery lifecycle: triage, code-gen, validation, documentation, and monitoring. See how the SDLC stages fit together. Run repeatable engineering workflows on a schedule or trigger. Coordinate planned, delegated, and validated multi-agent work. So Droid follows your repository conventions. Before adding more users or shared usage. # Worktrees in the Factory App Learn how to run Factory App sessions in separate Git worktrees while keeping your main project folder unchanged. Use a worktree when you want Droid to work on another task without changing the files in your main project folder. The Factory App creates a separate working directory for the session, prepares it, and keeps it isolated from your other work. Factory-managed worktrees are available for Git projects on your local machine and connected Droid Computers. Worktrees require a Git repository. For folders that are not Git repositories, start the session without a worktree. ## What is a worktree? A Git repository can have more than one working directory at the same time. Each directory has its own files and current branch, while all of them share the same commit history. You can think of your main project folder as your regular workspace and a worktree as a second workspace for a separate task. Changes in one workspace do not appear as uncommitted changes in the other. This allows you to make file changes on the same repository in parallel sessions, without the changes interfering with each other. ### Key terms A working directory that contains the repository files for a specific branch or commit. The project directory you selected in Factory. This is usually the checkout you already use in your editor and terminal. A separate checkout of the same Git repository. It has its own working files but shares commits and branches with the main checkout. The branch Factory uses as the starting point for a new worktree. Optional instructions that prepare a new worktree, give Droid initial context, or clean up services later. ## Create your first worktree Start a new session and choose the Git repository you want Droid to work on. Open the **Worktree** selector in the composer. Choose **New worktree**. Use the branch selector to choose the base branch. Factory creates a new branch from it by default, without switching branches in your main checkout. Open the panel beside **New worktree** if you want to keep the worktree permanently, use an existing branch, or run a setup profile. You can leave the defaults in place for your first session. Factory creates the checkout, runs the selected setup profile, and starts Droid inside it. The session shows the creation progress and worktree path. To work in the selected project directory instead, choose **Start without worktree**. ## Choose how long to keep it Every Factory-managed worktree has a lifecycle: | | Ephemeral | Persistent | | ----------------------- | -------------------------------------- | ------------------------------------ | | Best for | A task you expect to finish soon | An environment you plan to revisit | | Sessions | Usually one | Can contain multiple | | Automatic cleanup | Eligible | Never | | Available as a project | No | Yes | Choose **Ephemeral** for most isolated tasks. Choose **Persistent** when you want a long-lived checkout that appears in the project selector for future sessions. Factory remembers your last lifecycle choice for the project. ## Choose how the branch works By default, Factory creates a new branch from the base branch you selected. This is the simplest option because the new worktree owns its branch from the start. To continue work on a branch that already exists, open the branch mode control and choose **Use existing branch**. The branch cannot be: - The repository's default branch. - Checked out in your main checkout or another worktree. If you are unsure which branch mode to use, keep the default. Git allows a branch to be checked out in only one worktree at a time. ## Prepare new worktrees automatically A new checkout may need dependencies, generated files, or project-specific instructions before Droid can begin. A setup profile runs those tasks each time Factory creates a worktree for the project. In the desktop app, open **Settings → Worktrees** and select a project under **Setup profiles**. Select **Add profile** and enter a name. Add at least one setup script, initial prompt, or cleanup script. In the new-session composer, open the panel beside **New worktree** and select the profile. Factory preselects the last profile you used for that project. Each profile can contain: Runs from the root of the new worktree before the session starts. Use it to install dependencies or prepare local tooling. Gives Droid project-specific instructions before it receives your session prompt. Your prompt waits until this instruction finishes. Runs from the worktree root before Factory removes the checkout. Use it to stop services or remove external resources created during setup. Setup and cleanup scripts can use `REPO_ROOT_PATH` to find the main checkout. Factory removes credentials and Factory-specific environment variables before running profile scripts. ### Share a profile with your repository Profiles created in the app stay on the selected machine. To give everyone in the repository the same profile, commit a `.yaml` or `.yml` file under `.factory/worktree-setups/`: ```yaml title=".factory/worktree-setups/node.yaml" name: Node.js setup script: | set -euo pipefail pnpm install initial_prompt: Read the repository instructions before starting. cleanup_script: | docker compose down ``` The supported fields are `name`, `script`, `initial_prompt`, and `cleanup_script`. Shared profiles are read-only in the app. Edit their YAML files in the repository to change them. ## Include local files that Git ignores A worktree starts with the files tracked by Git. Local files excluded by `.gitignore`, such as `.env.local`, are not present automatically. To copy selected ignored files into each new Factory-managed worktree, create `.worktreeinclude` in the root of your main checkout. List one repository-relative path or directory per line: ```text # Local environment and tool configuration .env.local local/ ``` Blank lines and lines that begin with `#` are ignored. Directories are copied recursively. Factory copies matching files before running the setup profile, so the setup script can use them. Only paths already ignored by Git are eligible. Factory skips: - Tracked or unignored paths. - Missing files and directories. - Absolute paths and paths outside the repository. - Symbolic links. - Files that already exist in the new worktree. Review each path before adding it. Files ignored by Git can still contain credentials or other sensitive data. ## Manage and remove worktrees In the desktop app, open **Settings → Worktrees** to: - Change the worktree directory. The default is `~/.factory/worktrees`. - Set the number of ephemeral worktrees Factory keeps. The default limit is 15. - Browse managed worktrees by repository and machine. - Check each worktree's path, size, lifecycle, and active sessions. - Delete a worktree you no longer need. You can also delete a worktree from its sidebar menu. Before you confirm, Factory shows any uncommitted changes, untracked files, and unpushed commits it found. Removing the local branch or remote branch is a separate choice. When you archive the last session in a worktree, Factory asks whether to delete the worktree. If you confirm, Factory runs its cleanup script, removes the checkout, and archives the associated sessions. If you cancel, Factory keeps both the session and the worktree. ### How automatic cleanup works When Factory prepares a new worktree, it checks whether the new worktree would exceed your ephemeral worktree limit. If so, Factory starts cleanup with the least recently used ephemeral worktree and skips that worktree if it has: - An active or recently used session. - Uncommitted changes or untracked files. - An open pull request. - Unpublished commits, unless the related pull request is merged. If Factory skips a worktree, it checks the next one until it removes enough worktrees to meet the limit. Persistent worktrees do not count toward the limit. ### How orphan maintenance works Factory also runs a background maintenance sweep to reconcile sessions, worktrees, and Git metadata. The sweep runs after you archive a session and periodically when Factory refreshes the session list. It completes three passes in order: 1. **Archive sessions with missing worktrees.** If a session points to a worktree directory that no longer exists, Factory archives the session. If the configured worktree root is unavailable, Factory leaves the session unchanged. 2. **Remove orphaned worktrees.** Factory looks for managed worktrees that no session references. It preserves worktrees that are in use, persistent, referenced by any session, or waiting for a cleanup retry. Factory reloads the session list immediately before removal and preserves uncommitted changes or untracked files. This pass does not block on unpublished commits or pull request status because it removes only the worktree directory, not its branch. 3. **Clear leftover metadata.** Factory removes empty managed worktree group directories and asks Git to prune stale worktree records for repositories found in session metadata. Automatic cleanup removes the checkout, not its Git branch. ## Frequently asked questions ### Why is my `.env` file missing? Git worktrees contain tracked files. If Git ignores the file, add its path to [`.worktreeinclude`](#include-local-files-that-git-ignores) so Factory copies it into future worktrees. ### Does deleting a worktree delete its branch? Not automatically. Factory treats removal of the checkout, local branch, and remote branch as separate actions. Automatic cleanup removes only the checkout. ### What happens if I restore a session after deleting its worktree? If the worktree checkout no longer exists, restoring the archived session resumes it from the main checkout (the project directory you selected in Factory) instead of recreating the deleted worktree. ### Where does Factory store worktrees? Factory uses `~/.factory/worktrees` by default. Change the location under **Settings → Worktrees**. Configure personal defaults for Factory App sessions. Start isolated worktree sessions from the Droid CLI. Install the app and start a reviewable Droid session. Give every worktree session your repository instructions. # App settings Manage personal session defaults in the Factory App, including default model, reasoning effort, interaction mode, autonomy level, and spec mode overrides. The Factory App lets you set personal session defaults so every new session starts the way you want. Your choices are **capped by org policy**: if your organization enforces hard controls, your preferences take effect only within the boundaries those controls allow. See [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control) for the org side. ## Session defaults Session defaults are applied when a new session starts. In the Factory App, use the mode selector for **Normal Mode**, **Spec Mode**, or **Mission Mode**, and the Autonomy selector for `Auto Off`, `Auto Low`, `Auto Medium`, or `Auto High`. ### Default model Set `model` to a [model ID from the catalog](/models). This is the model new sessions use unless you pick a different one from the model selector. For custom models, see [Custom Models (BYOK)](/model-independence/byok). ### Reasoning effort `reasoningEffort` adjusts how much structured thinking the model performs before replying. The available levels and the default are **set by each model**, so the [Available Models](/models) table is the source of truth for what a given model accepts. Across the catalog they range from least to most deliberation: {/* sweep-allow: term-bullets */} - **`off` / `none`**: no structured reasoning (fastest). - **`minimal`**, **`low`**, **`medium`**, **`high`**: progressively more deliberation. - **`extra high`**, **`max`**: the deepest reasoning, where a model offers it. ### Interaction mode `sessionDefaultSettings.interactionMode` sets whether new sessions start in **Normal Mode** (`auto`) or **Spec Mode** (`spec`). In the app, the mode selector switches between the two. See [Interaction Modes](/autonomy-and-safety/specification-mode) for what each mode does. ### Autonomy level `sessionDefaultSettings.autonomyLevel` sets the default [Autonomy Level](/autonomy-and-safety/auto-run) for new sessions: `off` keeps manual approvals; `low`, `medium`, and `high` pre-authorize work at or below that risk level. In the app, the Autonomy selector chooses `Auto Off`, `Auto Low`, `Auto Medium`, or `Auto High`. ### Spec mode overrides When sessions start in Spec Mode, you can use a different model and reasoning effort for planning: - `sessionDefaultSettings.specModeModel`: the model used for spec mode planning. See [Models](/models) for available IDs. - `sessionDefaultSettings.specModeReasoningEffort`: reasoning effort for the spec mode model. Same values as `reasoningEffort`. ## Cloud session sync When `cloudSessionSync` is on, sessions are mirrored to the Factory App so you can revisit conversations from any browser. Your organization can control whether cloud sync is enabled at the org level; see [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control). ## How your preferences interact with org policy Your personal defaults are **session defaults**, not hard controls. They use the most-local-wins precedence: your choice overrides an org default, and a runtime override (such as a `--settings` flag or an in-app selector change) overrides your saved default. But org hard controls always cap what you can choose: - `modelPolicy` (`allowedModelIds`, `blockedModelIds`) decides which models exist at all. You can pick any allowed model, but blocked models are not available. - `maxAutonomyLevel` caps the effective `autonomyLevel` no matter who set it. If the org sets `medium`, your `high` default is clamped to `medium`. In short: your org decides the **boundaries**, and you pick your **preferred defaults inside those boundaries**. For the full hierarchy, merge semantics, and org schema, see [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control). Full settings.json reference for the Droid CLI. Org-managed settings hierarchy, merge semantics, and schema. Normal, Spec, and Mission Mode with Autonomy Level. How Autonomy Level controls approval prompts. # Droid CLI Use the Droid CLI to run Factory from your terminal, editor, tests, and Git workflow. Use the Droid CLI when you want Droid beside your shell, editor, tests, and Git workflow. It gives the same Factory runtime through an interactive terminal UI with project context, approvals, MCP tools, Missions, and headless execution. ## When to use the CLI - You want Droid to work in the same terminal where you run Git, tests, and scripts. - You prefer keyboard-first review of plans, tool calls, and diffs. - You need quick access to shell commands, slash commands, and project-local context. - You want a path from interactive work to `droid exec` automation. **Quick tip:** Press ! to toggle bash mode and run shell commands directly without AI interpretation. Press Esc to return to normal mode. See the [CLI Reference](/droid-cli/cli-reference#bash-mode) for details. ## Key capabilities The full Factory platform is available from the terminal. Launch multi-agent [Missions](/missions/overview) with `/missions` or `droid exec --mission`, delegate scoped tasks to [Custom Droids](/harness/subagents) via `/droids`, and package reusable procedures as [Skills](/harness/skills) with `/skills` or `/create-skill`. To extend the harness, connect external tools and data sources over [MCP](/harness/mcp) (`droid mcp add` or `/mcp`), run shell commands around agent lifecycle events with [Hooks](/harness/hooks) (`/hooks`), and install bundles of commands, droids, skills, and hooks as [Plugins](/harness/plugins) (`droid plugin install` or `/plugins`). When a workflow stabilizes, take it headless with [Exec Mode](/droid-exec/overview): `droid exec` runs in scripts and CI/CD pipelines with structured input and output formats. ## What `droid` brings to your workflow Droid plans, implements, tests, and reviews changes end to end while approval workflows keep you in control. It works from deep codebase understanding: project files, repository instructions, and connected knowledge inform every answer and scoped change. Connect Jira, Notion, Slack, Linear, PagerDuty, and MCP tools to keep development synchronized with team systems, and move repeatable local workflows into Droid Exec for scripts and CI/CD pipelines. Configure project-specific guidance and conventions. Package reusable procedures as self-contained workflows. # Droid CLI Quickstart Install the Droid CLI, start an interactive terminal session, and delegate your first reviewable coding task. Use this quickstart to install the Droid CLI, start an interactive terminal session, and delegate a first scoped task. You will see how Droid reads your codebase, proposes changes, and waits for review before editing. ## Before you begin Make sure you have: A terminal open in a code project A Git repository (recommended for the full workflow demonstration) ## Step 1: Install and start Droid **macOS/Linux:** `curl -fsSL https://app.factory.ai/cli | sh` **Homebrew:** `brew install --cask droid` **Windows:** `irm https://app.factory.ai/cli/windows | iex` **npm:** `npm install -g droid` **Linux users:** Ensure `xdg-utils` is installed for proper functionality. Install with: `sudo apt-get install xdg-utils` Then navigate to your project and start the Droid CLI. You'll see Droid's welcome screen in a full-screen terminal interface. If prompted, sign in via your browser to connect to Factory's development agent. ## Step 2: understand your codebase Start by asking Droid to map the project before it edits anything: ```text > analyze this codebase and explain the overall architecture ``` ```text > where are the main entry points, test commands, and code conventions? ``` Droid reads the files it needs, summarizes the architecture, and identifies the checks it should run before making changes. ## Step 3: run your first reviewable change Ask for one small, verifiable change: ```text > add structured logging to the app entry point and replace the existing console calls ``` Droid will: 1. Inspect the current implementation. 2. Propose a plan when the task needs one. 3. Show the exact diff. 4. Wait for approval before editing. 5. Run the relevant validators when you ask it to finish the task. Review the diff and run the relevant tests before accepting the change. This propose-and-approve loop is how you keep control as you delegate more work. ## Step 4: pick a goal to go deeper Use Specification Mode so Droid writes a plan before it starts a multi-step implementation. Review pull requests and local diffs for correctness, security, and edge cases. Move repeatable workflows into Droid Exec for scripts, CI, scheduled jobs, and pull request checks. Use Factory Missions when a task needs planning, delegation, validation, and checkpoints. ## Step 5: handle version control Droid makes Git operations conversational and intelligent: ```text > review my uncommitted changes and suggest improvements before I commit ``` ```text > create a well-structured commit with a descriptive message following our team conventions ``` ## Essential controls Here are the key interactions you'll use daily: | Action | What it does | How to use | | ----------------- | ----------------------------------------------------- | -------------------------------------------- | | Send message | Submit a task or question | Type and press Enter | | Multi-line input | Write longer prompts | for new lines | | Approve changes | Accept proposed modifications | Accept change in the TUI | | Reject changes | Decline proposed modifications | Reject change in the TUI | | Switch modes | Toggle between modes | | | Bash mode | Run shell commands directly without AI interpretation | Press ! on an empty input; Esc to return | | Transcript view | Toggle detailed transcript with full message details | | | Mission Control | Toggle Mission Control overlay (orchestrator sessions)| | | Close overlay | Dismiss the active overlay or menu | Escape | | Scroll transcript | Navigate through previous turns | / | | View shortcuts | Open the keyboard shortcuts pane (Esc to close) | Press ? | | Exit session | Leave droid | or type `exit` | ### Useful slash commands Quick shortcuts to common actions: | Command | What it does | | :------ | :----------- | | `/review` | Start the code review workflow ([learn more](/software-factory/code-review)). | | `/settings` | Configure Droid behavior, models, output style, and preferences. | | `/model` | Switch between AI models mid-session. | | `/sessions` | List and resume previous sessions. | | `/fork` | Copy the current session into a new session. You stay in the original; the receipt prints a `droid --resume ` command to open the fork elsewhere. | | `/compress` | Compress the current session to free up context. | | `/missions` | Open the Missions menu to launch multi-agent projects. | | `/droids` | Manage custom Droids (specialized subagents). | | `/skills` | Manage prompt-based skills (reusable procedures). | | `/hooks` | Manage tool execution hooks for lifecycle automation. | | `/plugins` | Manage plugins and marketplaces. | | `/mcp` | Manage Model Context Protocol servers. | | `/account` | Open your Factory account settings in the browser. | | `/billing` | View and manage your billing settings. | | `/help` | See all available commands. | [Learn how to create custom slash commands →](/harness/custom-slash-commands) Turns a repository into browsable architecture and setup documentation. Starts `/review` for uncommitted changes, commits, and branch diffs. Scores whether a repository is ready for larger autonomous work. Teaches Droid repository-specific conventions before it edits code. # Settings Configure how droid behaves and integrates with your workflow. This is the reference for `settings.json`, the personal CLI configuration file. It covers the settings you control as an individual developer. Settings that are typically managed by an organization (model policy, network restrictions, autonomy ceilings, and similar hard controls) are listed in a [clearly labeled section](#enterprise-and-org-level-settings) below and documented in full in [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control). For personal session defaults managed in the Factory App, see [App settings](/factory-app/settings). ## Accessing settings To configure droid settings: 1. Run `droid` 2. Enter `/settings` 3. Adjust your preferences interactively Changes take effect immediately and are saved to your settings file. ## Where settings live | OS | Location | | ------------- | -------------------------------------- | | macOS / Linux | `~/.factory/settings.json` | | Windows | `%USERPROFILE%\.factory\settings.json` | If the file doesn't exist, it's created with defaults the first time you run **droid**. ### Local overrides You can create a `settings.local.json` alongside `settings.json` in any `.factory/` folder: - `~/.factory/settings.local.json` (user-level) - `/.factory/settings.local.json` (project-level) Local overrides merge on top of the corresponding `settings.json` at the same level and follow the same hierarchy precedence. Add `settings.local.json` to `.gitignore` if you want to keep machine-specific preferences out of version control. ## Legacy Droid YAML configuration `.droid.yaml` was an older project configuration surface. Use the current `.factory/` files instead: - Use `settings.json` and `settings.local.json` for Droid preferences and local overrides. - Use [AGENTS.md](/harness/agents-md) for repository instructions, conventions, and validation commands. - Use [MCP servers](/harness/mcp), [hooks](/harness/hooks), and [skills](/harness/skills) for integrations, automation, and reusable workflows. ## Available settings | Setting | Options | Default | Description | | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------- | -------------------------------------------------------------------------- | | `model` | Any [available model ID](/models) | Product default | The default AI model used by droid | | `reasoningEffort` | Model-dependent: `none`, `dynamic`, `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max` (see [per-model options](/models)) | Model-dependent default | Controls how much structured thinking the model performs. | | `outputStyle` | `default`, `concise`, or a [custom style](/droid-cli/output-styles) selected in `/settings` | `default` | Controls how Droid structures and writes interactive CLI responses. | | `sessionDefaultSettings.interactionMode` | `auto`, `spec` | `auto` | Sets whether new sessions start in Auto or Spec Mode. | | `sessionDefaultSettings.autonomyLevel` | `off`, `low`, `medium`, `high` | `off` | Sets the default [Autonomy Level](/autonomy-and-safety/auto-run) for new sessions. | | `cloudSessionSync` | `true`, `false` | `true` | Mirror CLI sessions to Factory web. | | `diffMode` | `github`, `unified` | `github` | Choose between split GitHub-style diffs and a single-column view. | | `completionSound` | `off`, `bell`, `fx-ok01`, `fx-ack01`, or custom file path | `fx-ok01` | Audio cue when a response finishes. | | `awaitingInputSound` | `off`, `bell`, `fx-ok01`, `fx-ack01`, or custom file path | `fx-ack01` | Audio cue when droid is waiting for user input. | | `soundFocusMode` | `always`, `focused`, `unfocused` | `always` | When to play sound notifications. | | `commandAllowlist` | Array of commands | Safe defaults provided | Commands that run without extra confirmation. | | `commandDenylist` | Array of commands | Restrictive defaults provided | Commands that always require confirmation. | | `commandBlocklist` | Array of commands | `[]` (none) | Commands that can never run, even when approvals are skipped. | | `includeCoAuthoredByDroid` | `true`, `false` | `true` | Automatically append the Droid co-author trailer to commits. | | `enableDroidShield` | `true`, `false` | `true` | Enable secret scanning and git guardrails. | | `hooksDisabled` | `true`, `false` | `false` | Globally disable all hooks execution. | | `disabledSkills` | Array of skill names | `[]` | Disable discovered skills without deleting their files. | | `ideAutoConnect` | `true`, `false` | `false` | Auto-connect to IDE from external terminals. | | `showThinkingInMainView` | `true`, `false` | `false` | Display AI thinking/reasoning blocks in the main chat view. | | `customModels` | Array of model configs | `[]` | Custom model configurations for BYOK. See [BYOK docs](/model-independence/byok). | | `blockOnMcpLoad` | `true`, `false` | `false` | Wait for MCP servers to finish loading before starting an agent turn. | ### Model Set `model` to a [model ID from Factory-Managed Inference](/models). For custom models, see [Custom Models (BYOK)](/model-independence/byok). ### Reasoning effort `reasoningEffort` adjusts how much structured thinking the model performs before replying. The available levels and the default are **set by each model**, so the [Available Models](/models) table is the source of truth for what a given model accepts. Across the catalog they range from least to most deliberation: {/* sweep-allow: term-bullets */} - **`off` / `none`**: no structured reasoning (fastest). - **`dynamic`**: use the model's dynamic reasoning behavior. - **`minimal`**, **`low`**, **`medium`**, **`high`**: progressively more deliberation. - **`xhigh`**, **`max`**: the deepest reasoning, where a model offers it. Each model accepts only the subset shown in its [Available Models](/models) row, and its default sits within that subset (for example, Claude Opus 4.8 defaults to `high`, while Claude Sonnet 4.5 defaults to `off`). ### Output style Choose **Default**, **Concise**, or a custom style from `/settings`. Select custom styles in the TUI instead of editing their saved values by hand. Droid discovers them from user, project, and folder `.factory/output-styles/` directories and enabled plugins. See [Output styles](/droid-cli/output-styles) for the file format, scope, and session behavior. ### Autonomy level Use `sessionDefaultSettings.interactionMode` to choose whether new sessions start in Auto or Spec Mode, and `sessionDefaultSettings.autonomyLevel` to set the default [Autonomy Level](/autonomy-and-safety/auto-run). `off` keeps manual approvals; `low`, `medium`, and `high` pre-authorize work at or below that risk level. `sessionDefaultSettings.autonomyMode` is deprecated and retained for older configurations. ### Diff mode Control how droid displays code changes: - **`github`**: Side-by-side, higher fidelity render (recommended). - **`unified`**: Traditional single-column diff format. ### Cloud session sync When this switch is on, every CLI session is mirrored to Factory web so you can revisit conversations in the browser: - **`true`**: Sync sessions to the web app. - **`false`**: Keep sessions local only. ### Sound notifications Configure audio feedback for droid events: **Completion sound** (`completionSound`) - plays when a response finishes: Built-in completion sound (default) - soft success bloop. Alternative built-in sound effect - tactile ripple feedback. Use the system terminal bell. No sound notifications. Provide a file path to your own sound file (e.g., `"/path/to/sound.wav"`). **Awaiting input sound** (`awaitingInputSound`) - plays when droid is waiting for user input. Same options as completion sound, defaults to `fx-ack01`. **Sound focus mode** (`soundFocusMode`) - controls when sounds play: Play sounds regardless of window focus (default). Only play sounds when the terminal is focused. Only play sounds when the terminal is not focused. Access sound settings via `/settings` or → **Settings** in the TUI. ### Hooks The `hooksDisabled` setting provides a global toggle to disable all hooks execution without removing your hook configurations: - **`false`**: Hooks are enabled and will execute normally (default) - **`true`**: All hooks are disabled globally You can also toggle this from the `/hooks` menu or `/settings`. Set `showHookOutput` to `true` when you want hook stdout and stderr to appear in the session transcript for debugging. ### Disabled skills `disabledSkills` is a ledger of skill names that Droid should not invoke. Configure it from the `/skills` manager or edit the user or project settings file directly: ```json { "disabledSkills": ["deploy-production", "summarize-diff"] } ``` User and project arrays are combined, so a skill disabled at either level remains disabled. Stale entries are safe: a name has no effect until Droid discovers a matching skill. See [Skills](/harness/skills#manage-skills) for the manager, source precedence, and `SKILL.md` controls. ### IDE auto-connect The `ideAutoConnect` setting controls whether droid automatically connects to your IDE when running from external terminals (outside the IDE's built-in terminal): - **`false`**: Only auto-connect when running inside IDE terminal (default) - **`true`**: Auto-connect to IDE from any terminal ## Command allowlist, denylist & blocklist Use these settings to control which commands droid can execute automatically and which it must never run: Commands in this array are treated as safe and run without additional confirmation, regardless of autonomy prompts. Include only low-risk utilities you rely on frequently (for example `ls`, `pwd`, `dir`). Commands in this array always require confirmation and are typically blocked because they are destructive or unsafe (for example recursive `rm`, `mkfs`, or privileged system operations). A denied command can still be run if you explicitly approve it. Commands in this array can **never** run. Unlike the denylist, there is no prompt and no way to approve them: the block applies even under full autonomy, auto-run, or `--skip-permissions-unsafe`. droid also resolves the actual program being invoked, so a blocked command cannot be slipped through with a wrapper shell, an absolute path, quoting tricks, or command substitution. Commands that appear in both the allowlist and denylist default to the denylist behavior. The blocklist always takes precedence over both. Any command that is in none of the lists falls back to the autonomy level you selected for the session. ### Example allow/deny/block configuration ```json { "commandAllowlist": ["ls", "pwd", "dir"], "commandDenylist": ["rm -rf /", "mkfs", "shutdown"], "commandBlocklist": ["shutdown", "mkfs", "curl"] } ``` Use `commandBlocklist` for commands that should be hard-stopped regardless of approvals; use `commandDenylist` for commands that are allowed only after explicit confirmation. Review and update these arrays periodically to match your workflow and security posture, especially when sharing configurations across teams. ## Session defaults Defaults applied when a new session starts. See also `sessionDefaultSettings.interactionMode` and `sessionDefaultSettings.autonomyLevel` in the table above. | Setting | Type | Options | Default | Description | | ------------------------------------------------ | ------ | ---------------------------------------- | -------------- | ---------------------------------------------------------- | | `sessionDefaultSettings.specModeModel` | string | Any [available model ID](/models) | Inherits model | Override the model used when sessions start in Spec Mode. | | `sessionDefaultSettings.specModeReasoningEffort` | string | Model-dependent (see [/models](/models)) | Model default | Reasoning effort applied to the spec model. | ## Display and UI Tune how droid renders content in the terminal. How tool results are rendered in the transcript. Options: `expanded`, `compact`. Show the live token usage indicator at the bottom of the input. Control how the todo list is shown in the terminal UI. Animate the droid logo on startup. Options: `once`, `always`, `off`. Color theme used by the TUI. Accepts a theme ID (see `/themes`). Force droid's theme to override the terminal's color scheme. Enable Nerd Font glyphs in the UI (requires a Nerd Font in your terminal). ## Additional sound and notification settings Extends the [Sound notifications](#sound-notifications) section with toggles for the bell, per-event focus modes, and subagent activity. | Setting | Type | Options | Default | Description | | -------------------------------- | ------- | -------------------------------------- | --------- | -------------------------------------------------------------------- | | `subagentSounds` | string | `on`, `off` | `off` | Play sounds for subagent lifecycle events (start, complete, error). | | `completionSoundFocusMode` | string | `always`, `focused`, `unfocused` | Inherits `soundFocusMode` | Override focus behavior for completion sounds. | | `awaitingInputSoundFocusMode` | string | `always`, `focused`, `unfocused` | Inherits `soundFocusMode` | Override focus behavior for awaiting-input sounds. | ## Mission settings Configure [Missions](/missions/overview), multi-agent orchestration runs. Default model used by mission worker subagents. Accepts any [available model ID](/models). Reasoning effort for mission workers. Options are model-dependent (see [/models](/models)). Model used by mission validators (scrutiny / user-testing workers). Accepts any [available model ID](/models). Reasoning effort for validation workers. Options are model-dependent (see [/models](/models)). Skip scrutiny validation milestones during missions. Skip user-testing validation milestones during missions. Model used by the mission orchestrator. Accepts any [available model ID](/models). Reasoning effort for the mission orchestrator. Options are model-dependent (see [/models](/models)). Prevent the OS from sleeping while a mission is running. ## Subagent settings Configure the autonomy and models used by subagents spawned by the Task tool. These appear in the **Subagents** tab of `/settings`. | Setting | Type | Options | Default | Description | | ----------------------- | ------ | ----------------------------------------- | --------- | ------------------------------------------------------------------------------------ | | `subagentAutonomyLevel` | string | `inherit`, `off`, `low`, `medium`, `high` | `inherit` | [Autonomy Level](/autonomy-and-safety/auto-run) applied to subagents spawned by the Task tool. | ### Subagent autonomy level `subagentAutonomyLevel` controls how much subagents can do without approval: {/* sweep-allow: term-bullets */} - **`inherit`**: Subagents use the parent session's autonomy level (default). - **`off`**: Subagents require manual approval for every action. - **`low`**, **`medium`**, **`high`**: Pre-authorize subagent work at or below that risk level. An explicit level is still clamped to the org-managed `maxAutonomyLevel`, so it can never exceed the enterprise cap. Mission workers do not use this setting. ### Subagent model settings Subagents are routed by a complexity tier (`light`, `medium`, `heavy`). Each tier can pin its own model and reasoning effort under `subagentModelSettings`. | Setting | Type | Options | Default | Description | | --------------------------------------------- | ------ | ---------------------------------------------------------------------------------------- | -------------------- | -------------------------------------------------------------------- | | `subagentModelSettings.lightModel` | string | Any [available model ID](/models) or `inherit` | `inherit` | Model for light-complexity subagents (e.g. the built-in `explorer`). | | `subagentModelSettings.lightReasoningEffort` | string | Model-dependent: `none`, `dynamic`, `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max` | Parent/model default | Reasoning effort for light-complexity subagents. | | `subagentModelSettings.mediumModel` | string | Any [available model ID](/models) or `inherit` | `inherit` | Model for medium-complexity subagents (e.g. the built-in `worker`). | | `subagentModelSettings.mediumReasoningEffort` | string | Model-dependent: `none`, `dynamic`, `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max` | Parent/model default | Reasoning effort for medium-complexity subagents. | | `subagentModelSettings.heavyModel` | string | Any [available model ID](/models) or `inherit` | `inherit` | Model for heavy-complexity subagents. | | `subagentModelSettings.heavyReasoningEffort` | string | Model-dependent: `none`, `dynamic`, `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max` | Parent/model default | Reasoning effort for heavy-complexity subagents. | How the tiers are chosen: - The built-in `worker` subagent runs at `medium` complexity and `explorer` at `light`. Custom droids and the Task tool can request a specific tier. - **`inherit`** (the default) means the subagent uses the spawning session's active model and reasoning effort. Upgrading the parent session's model does not retroactively change a tier pinned to a specific model. - A custom droid with an explicit `model` in its configuration uses that model directly and bypasses complexity routing. - Without an explicit reasoning effort, a routed tier preserves the parent's active effort when it uses the same model; if it switches models, it uses the routed model's default. Explicit efforts are clamped to what the selected model supports. ## Context and compaction Controls when and how droid compacts the conversation to stay within the model's context window. Token threshold that triggers automatic compaction of the current session. Per-model overrides for `compactionTokenLimit`, as a `{ "": number }` map. Which model performs compaction: `same` uses the current session model, or specify a model ID. Per-model fallback routing when a selected model is unavailable, as a `{ "": "" }` map. ## Spec mode settings Controls the persistent spec store created by Spec Mode. | Setting | Type | Options | Default | Description | | ----------------- | ------- | ------------------------ | ----------------------------- | -------------------------------------------------------------------------- | | `specSaveDir` | string | Directory path | `~/.factory/specs` | Directory where saved specs are written. Supports `~` expansion. | ## Infrastructure System-level settings for status line, worktrees, and request timeouts. Custom status line configuration. Accepts `{ "command": string, "padding"?: number, "maxRows"?: number }`. The `command` is executed and its stdout rendered above the input. Configure interactively with `/statusline`. Default parent directory for git worktrees created with `--worktree` / `-w`. Timeout in milliseconds for individual LLM requests before they are aborted. Timeout in milliseconds for subagent inactivity before Droid treats the worker as stalled. Allow this machine to be used for remote Droid Computer access. ## Enterprise and org-level settings **These keys are org-managed, not personal preferences.** They are typically configured by an organization Manager or Owner and pushed to members through the Factory App or a managed `settings.json`. They appear here so the `settings.json` schema is complete in one place, but individual users should not expect to change them. See [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control) for hierarchy, merge semantics, and admin enforcement details. Usage alert preferences are Factory App controls; individual users control their own threshold alerts. | Setting | Type | Options | Default | Description | | ------------------------------------ | ------- | ---------------------------------------------------------------------------------------------------------------------- | ------------ | ------------------------------------------------------------------------------------------------------------------------------------ | | `maxAutonomyLevel` | string | `off`, `low`, `medium`, `high` | `high` | **Enterprise.** Maximum [Autonomy Level](/autonomy-and-safety/auto-run) any session may use. Higher levels selected by users are clamped. | | `webSearchDisabled` | boolean | `true`, `false` | `false` | **Org-only.** Set `true` to remove built-in Web Search from the CLI. User, project, and folder values are ignored; Web Fetch is unchanged. See [managed tool controls](/enterprise/hierarchical-settings-and-org-control#autonomy-safety-and-commands). | | `subagentAutonomyLevel` | string | `low`, `medium`, `high`, `inherit` | `inherit` | **Enterprise.** Default autonomy ceiling for subagent workers. | | `modelPolicy` | object | `{ "allowedModelIds"?: string[], "blockedModelIds"?: string[], "allowCustomModels"?: boolean, "allowedBaseUrls"?: string[] }` | unset | **Enterprise.** Restrict which models members can select, control whether custom models are permitted, and allowlist base URLs for custom model providers. | | `mcpPolicy` | object | `{ "enabled"?: boolean, "allowlist"?: string[] }` | unset | **Enterprise.** Control whether MCP servers can run and which servers are permitted. See [MCP](/harness/mcp). | | `mcpAutonomyOverrides` | object | Per-server and per-tool autonomy map | unset | **Enterprise.** Set default autonomy levels for specific MCP servers or tools. | | `mcpAutonomyUrlOverrides` | array | URL pattern autonomy rules | unset | **Enterprise.** Set autonomy defaults for MCP tools that target matching URLs. | | `missionPolicy` | object | `{ "restrictedAccess"?: boolean, "allowedUserIds"?: string[] }` | unset | **Enterprise.** Restrict who can launch [Missions](/missions/overview). | | `networkPolicy` | object | `{ "allowedIps": string[] }` | unset | **Enterprise.** Restrict outbound network access from droid sessions to the specified IPs or CIDR ranges in `allowedIps`. | | `sandbox` | object | `{ "enabled"?: boolean, "mode"?: string, "filesystem"?: object, "network"?: object }` | unset | Sandbox configuration controlling filesystem and network isolation for tool execution. See [Sandbox](/autonomy-and-safety/sandbox). | | `restrictMemberVisibility` | boolean | `true`, `false` | `false` | **Org-level.** Hide other org members from non-admin users. | | `restrictApiKeyCreationToManagers` | boolean | `true`, `false` | `false` | **Org-level.** Only org managers may create API keys. | | `sessionRetentionDays` | number | `14`–`365` | Org default | **Org-level.** How long synced session history is retained before deletion. | | `wikiCloudSync` | boolean | `true`, `false` | `true` | **Org-level.** Sync generated [Wiki](/software-factory/wiki/overview) content to Factory cloud. | | `voiceDictationEnabled` | boolean | `true`, `false` | Org default | **Org-level.** Control whether voice dictation is available as an input method to organization members. | | `disableWeeklyUsageSummary` | boolean | `true`, `false` | `false` (summary on) | **Enterprise/org-level.** Opt out of the weekly usage-cap summary email sent to Managers and Owners for users and active service accounts at or above 80% of their monthly credit cap. Set `true` to suppress. | | `managedComputersEnabled` | boolean | `true`, `false` | `false` | **Enterprise.** Enable Factory-managed remote computers for the org. | | `managedComputersAllowedEmails` | array | Email allowlist | unset | **Enterprise.** Limit Factory-managed Droid Computers to specific members. | | `byomComputersEnabled` | boolean | `true`, `false` | `false` | **Enterprise.** Allow members to register their own machines via `droid computer register` (BYOM). | | `byomComputersAllowedEmails` | array | Email allowlist | unset | **Enterprise.** Limit BYOM Droid Computer registration to specific members. | | `factoryRouterGuidance` | string | Routing guidance text | unset | **Enterprise.** Organization-wide guidance for Factory Router model selection. | | `factoryRouterRules` | array | Routing rule objects | unset | **Enterprise.** Conditional Factory Router guidance rules. | | `allowManagedHooksOnly` | boolean | `true`, `false` | `false` | **Enterprise.** Require hooks to come from managed settings rather than local configuration. | | `disableAutoUpdate` | boolean | `true`, `false` | `false` | **Enterprise.** Disable automatic Droid CLI updates where managed deployment controls updates. | ## Factory App usage alert preferences The following preference is stored in each user's Factory App profile. It cannot be configured in `settings.json`. | Preference | Type | Options | Default | Description | | ------------------------- | ------- | --------------- | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | | `disableUsageLimitAlerts` | boolean | `true`, `false` | `false` (alerts on) | Opt out of the emails sent when you reach 80% and 100% of your monthly credit cap. Set `true` to stop these alerts. | ## Example configuration ```json { "model": "claude-opus-4-7", "reasoningEffort": "low", "outputStyle": "concise", "diffMode": "github", "cloudSessionSync": true, "completionSound": "fx-ok01", "awaitingInputSound": "fx-ack01", "soundFocusMode": "always" } ``` Editor-specific setup. Command flags and options. Customize how Droid structures and writes interactive responses. Add custom models and API keys. # Output styles Control how Droid structures and writes responses in interactive CLI sessions. Output styles add response-writing instructions to interactive Droid CLI sessions. Choose a built-in style or create a Markdown file for a personal, project, or shared team style. ## Choose an output style 1. Start an interactive session with `droid`. 2. Enter `/settings`. 3. Select **Output style**. 4. Choose a built-in or custom style. Droid includes two built-in styles: Use Droid's standard response style. Return short, direct responses without unnecessary detail. The picker shows where each custom style came from, such as a user, project, or plugin. Your selection is saved as the `outputStyle` setting in `~/.factory/settings.json`. ## Create a custom style Create an `output-styles` directory at the scope where you want the style to be available: | Scope | Location | | :---- | :------- | | User | `~/.factory/output-styles/.md` | | Project | `/.factory/output-styles/.md` | | Folder | `//.factory/output-styles/.md` | | Plugin | `/output-styles/.md` | Droid loads direct `.md` children of each `output-styles` directory. Nested files and other file types are ignored. Each file contains optional YAML frontmatter followed by the instructions Droid should apply: ```markdown title="~/.factory/output-styles/review-notes.md" --- name: Review Notes description: Put findings before the summary --- Start with actionable findings, ordered by severity. Keep the summary to two sentences. ``` `name` appears in the picker and defaults to the filename without `.md`. `description` appears below the name in the picker. Keep the Markdown body focused on response presentation. Repository instructions, tool permissions, and user requests still apply. Droid uses the style content available when a session sends its first request. Start a new session to use edits to that style. Selecting a different style takes effect on the next request. ## Scope and precedence Output styles apply only to interactive Droid CLI sessions. They do not change responses from `droid exec`, Automations, Missions, subagents, the Factory App, or web sessions. Droid combines styles from the active settings hierarchy and enabled plugins. Higher-priority sources win when the same style is provided more than once. The built-in `Default` and `Concise` styles are reserved and cannot be replaced. When Droid asks you to trust a working folder, review its style files before accepting. Project and folder styles, including styles from project-scoped plugins, remain unavailable until the folder is trusted. ## Troubleshooting Confirm the style is a direct `.md` child of `output-styles/` with a non-empty body. Run `/diagnostics` to see the source path and validation error, including when a style exceeds the size limit. Diagnostics do not include the style's instruction body. Confirm the style still exists, passes validation, comes from a trusted folder, and belongs to an enabled plugin. Droid keeps an unavailable selection but uses **Default** until the style becomes available again. Configure the saved output style and other Droid preferences. Package output styles with reusable Droid extensions. Review interactive commands and headless output formats. See how organization, folder, project, and user sources resolve. # Droid CLI Reference Complete reference for the Droid CLI, including commands and flags ## Installation **macOS/Linux:** `curl -fsSL https://app.factory.ai/cli | sh` **Homebrew:** `brew install --cask droid` **Windows:** `irm https://app.factory.ai/cli/windows | iex` **npm:** `npm install -g droid` The Droid CLI operates in two modes: - **Interactive (`droid`)** - Chat-first REPL with slash commands - **Non-interactive (`droid exec`)** - Single-shot execution for automation and scripting ### Installing a specific version Pin a version when upgrades must be reviewed and applied intentionally, such as in CI runners, provisioned development environments, or other reproducible systems. #### npm Specify an exact package version: ```bash npm install -g droid@0.174.0 ``` The binaries published through npm have auto-updates disabled at build time, so the installed version remains pinned. To upgrade, install the new version explicitly with npm. #### Built-in updater After installing Droid through the standalone installer, select a version with: ```bash droid update --version 0.174.0 ``` This command can also roll back to an older version. Use `droid update --check` to check for an available update without installing it. Standalone installations enable automatic updates by default. After selecting a version, follow the [auto-update guidance](#auto-updates) to keep it pinned. #### Direct binary download Release binaries and their SHA256 checksums use the following paths: ```text https://downloads.factory.ai/factory-cli/releases//// https://downloads.factory.ai/factory-cli/releases////.sha256 ``` - ``: `linux`, `darwin`, or `windows` - ``: `x64`, `arm64`, or `x64-baseline` - ``: `droid` on Linux and macOS, or `droid.exe` on Windows For example, download and verify the Linux x64 binary: ```bash VERSION=0.174.0 BASE_URL="https://downloads.factory.ai/factory-cli/releases/${VERSION}/linux/x64" curl -fsSLO "${BASE_URL}/droid" curl -fsSLO "${BASE_URL}/droid.sha256" echo "$(awk '{print $1}' droid.sha256) droid" | sha256sum --check chmod +x droid ``` Move the verified binary to a directory on `PATH`. The download URL layout is an implementation detail, so always pin a version and verify its checksum. ## Droid CLI commands | Command | Description | Example | | :----------------------------------- | :--------------------------------------------------------------------- | :--------------------------------------------- | | `droid` | Start interactive REPL | `droid` | | `droid "query"` | Start REPL with initial prompt | `droid "explain this project"` | | `droid --resume [sessionId]` | Resume a session (defaults to last modified). Alias: `-r` | `droid --resume` | | `droid --fork ` | Fork and resume a session in a new copy | `droid --fork session-abc123` | | `droid exec "query"` | Execute task without interactive mode | `droid exec "summarize src/auth"` | | `droid exec -f prompt.md` | Load prompt from file | `droid exec -f .factory/prompts/review.md` | | `cat file \| droid exec` | Process piped content | `git diff \| droid exec "draft release notes"` | | `droid exec -s "query"` | Resume existing session in exec mode | `droid exec -s session-123 "continue"` | | `droid exec --list-tools` | List available tools, then exit | `droid exec --list-tools` | | `droid search "query"` | Search across local sessions (messages, documents, tool results). Alias: `droid find` | `droid search "auth bug"` | | `droid mcp add ` | Add an MCP server | `droid mcp add api https://api.example.com/mcp --type http` | | `droid mcp remove ` | Remove an MCP server | `droid mcp remove linear` | | `droid plugin install ` | Install a plugin. Alias: `droid plugin i` | `droid plugin install droid-control@factory-plugins` | | `droid plugin uninstall ` | Uninstall a plugin. Alias: `droid plugin remove` | `droid plugin uninstall droid-control@factory-plugins` | | `droid plugin update ` | Update a plugin to the latest version | `droid plugin update droid-control@factory-plugins` | | `droid plugin list` | List installed plugins | `droid plugin list` | | `droid plugin marketplace` | Manage plugin marketplaces | `droid plugin marketplace` | | `droid computer register [name]` | Register this machine as a Bring-Your-Own-Machine (BYOM) computer | `droid computer register laptop` | | `droid computer remove` | Remove this machine's BYOM registration | `droid computer remove` | | `droid computer list` | List registered BYOM computers | `droid computer list` | | `droid computer ssh ` | SSH into a registered BYOM computer | `droid computer ssh laptop` | | `droid computer port-forward ` | Forward local ports to a computer over the relay | `droid computer port-forward laptop 8080:80` | | `droid daemon` | Run the Factory daemon server | `droid daemon` | | `droid update` | Manually update the CLI to latest version | `droid update` | ## Droid CLI flags Customize droid's behavior with command-line flags: | Flag | Description | Example | | :-------------------------------- | :----------------------------------------------------------------- | :----------------------------------------------------------- | | `-f, --file ` | Read prompt from a file | `droid exec -f plan.md` | | `-m, --model ` | Select a specific [model ID](/models) | `droid exec -m claude-opus-4-7` | | `-s, --session-id ` | Continue an existing session | `droid exec -s session-abc123` | | `--auto ` | Set [autonomy level](#autonomy-levels) (`low`, `medium`, `high`) | `droid exec --auto medium "run tests"` | | `--restrict-tools ` | Restrict the run to only the specified tools (comma or space separated) | `droid exec --auto low --restrict-tools ApplyPatch,Execute` | | `--additional-tools ` | Force-enable additional tools beyond the defaults (comma or space separated) | `droid exec --additional-tools ApplyPatch` | | `--disabled-tools ` | Disable specific tools for this run | `droid exec --disabled-tools execute-cli` | | `--disable-builtin-skills` | Disable Factory-provided builtin skills in interactive or exec sessions while preserving other skill sources | `droid --disable-builtin-skills` | | `--list-tools` | Print available tools and exit | `droid exec --list-tools` | | `-o, --output-format ` | Output format (`text`, `json`, `stream-json`, `stream-jsonrpc`) | `droid exec -o json "document API"` | | `--input-format ` | Input format for multi-turn sessions. Use `stream-jsonrpc`; the older `stream-json` mode is deprecated. Must match `--output-format`. | `droid exec --input-format stream-jsonrpc -o stream-jsonrpc` | | `-r, --resume [sessionId]` | Resume a previous session. In interactive mode, `-r` is `--resume`; in `droid exec`, `-r` is `--reasoning-effort`. | `droid -r` | | `-r, --reasoning-effort ` | Override reasoning effort; valid levels are model-dependent (see [/models](/models)). In `droid exec`, `-r` maps to this flag. | `droid exec -r high "debug flaky test"` | | `--spec-model ` | Use a different [model ID](/models) for specification planning | `droid exec --spec-model claude-opus-4-7` | | `--spec-reasoning-effort ` | Override reasoning effort for spec mode | `droid exec --use-spec --spec-reasoning-effort high` | | `--use-spec` | Start in specification mode (plan before executing) | `droid exec --use-spec "add user profiles"` | | `--skip-permissions-unsafe` | Skip all permission prompts (use with extreme caution) | `droid exec --skip-permissions-unsafe` | | `--cwd ` | Execute from a specific working directory | `droid exec --cwd ../service "run tests"` | | `-w, --worktree [name]` | Run the session in an isolated [git worktree](#git-worktrees) | `droid --worktree fix-bug` | | `--tag ` | Session tag (name or JSON, repeatable) | `droid exec --tag code-review` | | `--log-group-id ` | Log group ID for filtering logs | `droid exec --log-group-id grp-123` | | `--fork ` | Fork and resume an existing session into a new copy | `droid exec --fork session-abc123` | | `--mission` | Run `droid exec` in [Mission Mode](/missions/overview) (multi-agent orchestration). Requires `--auto high` or `--skip-permissions-unsafe`. | `droid exec --mission --auto high -f mission.md` | | `--worker-model ` | Model used for mission worker agents | `droid exec --mission --worker-model claude-sonnet-4-6` | | `--worker-reasoning-effort ` | Reasoning effort for mission worker agents (model-dependent; see [/models](/models)) | `droid exec --mission --worker-reasoning-effort medium` | | `--validator-model ` | Model used for mission validator agents | `droid exec --mission --validator-model claude-opus-4-7` | | `--validator-reasoning-effort ` | Reasoning effort for mission validator agents | `droid exec --mission --validator-reasoning-effort high` | | `--append-system-prompt ` | Append custom text to the end of the system prompt | `droid --append-system-prompt "Always run tests."` | | `--append-system-prompt-file ` | Append the contents of a file to the end of the system prompt | `droid --append-system-prompt-file .factory/system.md` | | `-v, --version` | Display CLI version | `droid -v` | | `-h, --help` | Show help information | `droid --help` | Use `--output-format json` for scripting and automation, so you can parse droid's responses programmatically. `--output-format` controls how `droid exec` serializes results. To change how Droid writes responses in an interactive CLI session, use [output styles](/droid-cli/output-styles). ## Autonomy levels `droid exec` uses tiered autonomy to control what operations the agent can perform. Only raise access when the environment is safe. | Level | Intended for | Notable allowances | | :-------------------------- | :----------------------- | :------------------------------------------------------------ | | _(default)_ | Read-only reconnaissance | File reads, git diffs, environment inspection | | `--auto low` | Safe edits | Create/edit files, run formatters, non-destructive commands | | `--auto medium` | Local development | Install dependencies, build/test, local git commits | | `--auto high` | CI/CD & orchestration | Git push, deploy scripts, long-running operations | | `--skip-permissions-unsafe` | Isolated sandboxes only | Removes all guardrails (use only in disposable containers) | **Examples:** ```bash # Default (read-only) droid exec "Analyze the auth system and create a plan" # Low autonomy - safe edits droid exec --auto low "Add JSDoc comments to all functions" # Medium autonomy - development work droid exec --auto medium "Install deps, run tests, fix issues" # High autonomy - deployment droid exec --auto high "Run tests, commit, and push changes" ``` `--skip-permissions-unsafe` removes all safety checks. Use **only** in isolated environments like Docker containers. ## Model IDs Use any [available model ID](/models) with `-m, --model` or `--spec-model`. For custom models, see [Custom Models (BYOK)](/model-independence/byok). ## Interactive mode features ### Keyboard shortcuts The interactive REPL supports a rich set of keyboard shortcuts for navigation, overlays, and input control: | Shortcut | Action | | :------- | :----- | | | Cancel the current operation / interrupt the agent. Press twice quickly to exit | | | Suspend the process (Unix only). Resume with `fg` | | | Toggle the detailed transcript view (full message details) | | | Toggle the Mission Control overlay (orchestrator sessions only) | | | Cycle through available models (when typing in chat input) | | | Cycle through autonomy levels (when typing in chat input) | | | Toggle the `/btw` scroll view (side-question history) | | | Toggle the changelog display (dismiss / restore) | | | Toggle the approval details view (Option+E on macOS) | | | Paste an image from the clipboard as an attachment | | Tab | Cycle through the model's available reasoning-effort levels | | | Toggle Normal Mode and Spec Mode | | @ | File path autocomplete: typing `@` triggers fuzzy file search | | Up / Down | Navigate input history (cycle through previously submitted messages) | | Double Escape | Second press clears the input draft; a third Escape opens the rewind menu | | ? | Open the scrollable keyboard shortcuts pane (when the input is empty); close it with Esc | | | Open the keyboard shortcuts pane (works even when the input has content) | | | Move the cursor to the start of the line | | | Delete the word before the cursor | | | Delete from the cursor to the end of the line | | | Delete from the cursor to the start of the line | | | Clear all attached images, or forward-delete the next character if none attached | | | Open the fork tree for the highlighted session (in the `/sessions` list view) | | | Rename the highlighted session (in the `/sessions` list view) | | | Archive the highlighted session (in the `/sessions` list view); restores it on the Archived tab | | | Scroll the transcript up (navigate any turn) | | | Scroll the transcript down (navigate any turn) | | | Scroll the transcript up to the previous user turn | | | Scroll the transcript down to the next user turn | | Escape | Close the active pane, overlay, or menu | | | Insert a newline in the chat input (multiline editing) | | ! | Toggle [bash mode](#bash-mode) (when the input is empty) | Run `/terminal-setup` once to configure your terminal so reliably produces a newline in the chat input. ### Bash mode Press ! when the input is empty to toggle bash mode. In bash mode, commands execute directly in your shell without AI interpretation, useful for quick operations like checking `git status` or running `npm test`. {/* sweep-allow: term-bullets */} - **Toggle on:** Press ! (when input is empty) - **Execute commands:** Type any shell command and press Enter - **Toggle off:** Press Esc to return to normal AI chat mode The prompt changes from `>` to `$` when bash mode is active. ### Mermaid diagram rendering Droid automatically renders Mermaid diagram code blocks as ASCII art directly in the terminal, with no external viewer or browser required. When a response contains a fenced ```` ```mermaid ```` block of a supported diagram type, the diagram is drawn inline in the transcript. **Supported diagram types:** - `flowchart` (and `graph`) - `sequenceDiagram` - `stateDiagram` - `classDiagram` - `erDiagram` Unsupported diagram types (for example `gantt`, `pie`, `mindmap`, `timeline`, `journey`, `gitGraph`) fall back to displaying the raw Mermaid source and a link to view the diagram externally. ### Markdown and math rendering Droid renders supported inline Markdown styles in responses, including formatting inside table cells. It also converts inline and display TeX into Unicode math for terminal output. Use `\(...\)` or `$...$` for inline math. Use `\[...\]` or `$$...$$` for display math. ### Slash commands Available when running `droid` in interactive mode. Type the command at the prompt: | Command | Description | | :---------------------------- | :------------------------------------------------------------- | | `/account` | Open Factory account settings in browser | | `/archive` | Archive the current session (restore it from the `/sessions` Archived tab) | | `/billing` | View and manage billing settings | | `/btw ` | Ask a side question without polluting the main transcript | | `/bug [title]` | Create a bug report with session data and logs | | `/clear` | Clear conversation context, keep current model & autonomy | | `/commands` | Manage custom slash commands | | `/compress [prompt]` | Compress session and move to new one with summary | | `/context` | Show context window usage breakdown with progress bar | | `/copy` | Copy prompts, responses, turn ranges, or session ID | | `/cost` | Show usage statistics | | `/create-skill` | Create a reusable skill from current session | | `/cd ` | Change session working directory (alias for `/cwd`) | | `/cwd ` | Change session working directory | | `/diagnostics` | Show settings configuration errors | | `/droids` | Manage custom droids | | `/missions` | Enter Mission Mode | | `/fast` | Enable fast mode for current model (`/fast off` to disable) | | `/favorite` | Mark current session as a favorite | | `/fork` | Copy current session into a new session; you stay in the original and get a `droid --resume ` command for the fork | | `/help` | Show available slash commands | | `/hooks` | Manage lifecycle hooks | | `/ide` | Configure IDE integrations | | `/install-code-review` | Set up automated code review | | `/install-slack-app` | Install/connect Slack integration | | `/language ` | Switch TUI display language | | `/limits` | Manage token usage limits and overage preferences | | `/login` | Sign in to Factory | | `/logout` | Sign out of Factory | | `/mcp` | Manage Model Context Protocol servers | | `/model` | Switch AI model mid-session | | `/new` | Start a fresh session, reset model & autonomy to defaults | | `/plugins` | Manage plugins and marketplaces | | `/quit` | Exit droid (alias: `exit`, or press ) | | `/readiness-fix` | Fix failing agent readiness signals from latest report | | `/readiness-report` | Generate readiness report | | `/rename` | Rename current session | | `/review` | Start AI-powered code review workflow | | `/rewind-conversation` | Undo recent changes in the session | | `/sessions` | List and select previous sessions | | `/settings` | Configure application settings | | `/setup-incident-response` | Set up Slack auto-run for incident-response channel | | `/share` | Share session with organization | | `/skills` | Manage and invoke skills | | `/stats [period]` | Show usage statistics (supports relative periods, date ranges) | | `/status` | Show current droid status and configuration | | `/statusline` | Configure custom status line | | `/terminal-setup` | Configure terminal keybindings for | | `/themes` | Choose a color theme | | `/tree` | Browse and resume branches in the current session's fork tree | Slash commands stay usable while the agent is running: read-only and settings commands open right away over the live stream, commands that modify the conversation or session (such as `/new`, `/clear`, `/compress`) ask for confirmation before stopping the current run, and `/model`, `/fast`, and `/archive` apply only between turns. `/fork` also runs immediately: it copies the session in the background and keeps you in the original session, so there is nothing to interrupt. Archiving hides a session from your session lists without deleting any data. Restore an archived session anytime from the Archived tab in `/sessions` (where restores instead of archiving), or send a new message in the archived session to restore it automatically. For detailed information on slash commands, see the [interactive mode documentation](/droid-cli/quickstart#useful-slash-commands). ### Git worktrees Use `-w, --worktree [name]` to run a session inside a native [git worktree](https://git-scm.com/docs/git-worktree) so you can work on multiple branches of the same repository in parallel without file conflicts. This flag is available on both `droid` (interactive) and `droid exec`. By default, Droid creates worktrees under the [`worktreeDirectory` setting](/droid-cli/settings#infrastructure), which defaults to `~/.factory/worktrees`. Each worktree uses the path `/<8-character group>//` and has its own checkout and dedicated branch. **Branch naming:** - **`--worktree`** (no value): Creates/reuses a worktree on a branch named `-wt`. - **`--worktree `**: Uses `` as the branch. If the branch already exists, it is checked out in the worktree; otherwise it is created from `HEAD`. - If the target branch is already checked out in another worktree, the command fails with a clear error. **Examples:** ```bash # Interactive: derive branch from current branch (creates -wt) droid --worktree # Interactive: explicit branch name droid -w fix-auth-bug "start debugging the login flow" # Headless: isolate an automated task on its own branch droid exec --worktree refactor-tests --auto medium "migrate jest suites to vitest" # Run two parallel sessions on the same repo, each on its own branch droid --worktree feature-a & droid --worktree feature-b & ``` **Session lifecycle:** - **Interactive mode**: The worktree persists after the session ends so you can resume work, inspect changes, or push the branch. - **`droid exec` mode**: On exit, a clean worktree (no uncommitted changes) is removed automatically; a dirty worktree is preserved and its path is printed so you can review the work. - The underlying git branch is **never deleted** by Droid; only the worktree directory is removed during cleanup. When `--worktree` is active, Droid operates entirely inside the worktree directory. Run follow-up commands (tests, builds, installs) from that directory rather than the original repo root. Fresh worktrees may not have dependencies installed yet (for example `node_modules`); set them up if builds or tests fail with missing-module errors. ### Subcommand flags In addition to the global flags above, several `droid` subcommands accept their own flags. #### `droid update` flags | Flag | Description | Example | | :------------------------ | :------------------------------------------------- | :--------------------------- | | `-c, --check` | Check for updates without installing | `droid update --check` | | `-v, --version ` | Update to or roll back to a specific version | `droid update -v 0.174.0` | #### `droid search` flags | Flag | Description | Example | | :-------------------- | :----------------------------------------------------------------------- | :--------------------------------------- | | `--kind ` | Filter by entry kind: `message_text`, `document`, `tool_use`, `tool_result`, or `all` | `droid search "auth" --kind document` | | `--limit-sessions ` | Maximum number of sessions to return | `droid search "auth" --limit-sessions 5` | | `--limit-hits ` | Maximum number of matches per kind per session | `droid search "auth" --limit-hits 3` | | `--context-chars ` | Number of characters of context shown around each match | `droid search "auth" --context-chars 200`| | `--json` | Emit results as JSON | `droid search "auth" --json` | | `--reindex` | Drop the search cache and rebuild the local index | `droid search "auth" --reindex` | #### `droid mcp add` flags | Flag | Description | Example | | :------------------------- | :----------------------------------------------------------------------- | :--------------------------------------------------------------- | | `--type ` | Server transport: `http` (streamable HTTP), `sse` (legacy HTTP+SSE), or `stdio` (local) | `droid mcp add api https://api.example.com/mcp --type http` | | `--env ` | Set an environment variable for a stdio server (repeatable) | `droid mcp add gh "gh-mcp" --type stdio --env GH_TOKEN=$GH_TOKEN`| | `--header ` | Add an HTTP header for a remote server (repeatable) | `droid mcp add api https://api.example.com/mcp --type http --header "Authorization: Bearer $TOKEN"` | #### `droid plugin` flags | Flag | Description | Example | | :-------------------- | :----------------------------------------------------------------------- | :------------------------------------------------------- | | `-s, --scope ` | Installation scope: `user` (default) or `project` | `droid plugin install droid-control@factory-plugins --scope project` | #### `droid computer register` flags | Flag | Description | Example | | :----------- | :------------------------------------------------------- | :----------------------------------- | | `-y, --yes` | Skip interactive prompts and accept registration defaults | `droid computer register laptop -y` | ### MCP command reference The `/mcp` slash command opens an interactive manager UI for browsing and managing MCP servers. **Quick start:** Type `/mcp` and select **"Add from Registry"** to browse 40+ pre-configured servers (Linear, Sentry, Notion, Stripe, Vercel, and more). Select a server, authenticate if required, and you're ready to go. **CLI commands** for scripting and automation: ```bash droid mcp add --type http # Add HTTP (streamable) server droid mcp add --type sse # Add legacy HTTP+SSE server droid mcp add "" # Add stdio server droid mcp remove # Remove a server ``` See [MCP Configuration](/harness/mcp) for the full registry list, CLI options (`--env`, `--header`), configuration files, and how user vs project config layering works. ## Authentication 1. Generate an API key in the Factory API keys settings 2. Set the environment variable: ```bash macOS/Linux export FACTORY_API_KEY=fk-... ``` ```powershell Windows (PowerShell) $env:FACTORY_API_KEY="fk-..." ``` ```cmd Windows (CMD) set FACTORY_API_KEY=fk-... ``` **Persist the variable** in your shell profile (`~/.bashrc`, `~/.zshrc`, or PowerShell `$PROFILE`) for long-term use. Never commit API keys to source control. Use environment variables or secure secret management. ## Auto-updates Standalone Droid installations check for and install CLI updates automatically. Run `droid update` to trigger an update on demand. To keep a standalone installation pinned, disable its in-process updater: ```bash export FACTORY_DROID_AUTO_UPDATE_ENABLED=false ``` With this variable set, automatic updates are disabled and `droid update` does not install another version. Set it in the environment that launches every Droid process, including `droid daemon`. The npm distribution has auto-updates disabled at build time and does not require this variable. Running `droid update` from an npm installation reports that auto-update is unavailable. Setting `FACTORY_DROID_AUTO_UPDATE_ENABLED=true` explicitly overrides the npm build default. Enterprise administrators can disable CLI auto-updates for all organization members with the `disableAutoUpdate` organization setting. See the [release notes](/changelog/release-notes) and [organization controls](/enterprise/hierarchical-settings-and-org-control). Pin the CLI and disable auto-updates in reproducible environments. For local developer installations, leave auto-updates enabled to receive fixes and security updates. ## Exit codes | Code | Meaning | | :--- | :---------------------------- | | `0` | Success | | `1` | General runtime error | | `2` | Invalid CLI arguments/options | ## Common workflows ### Code review Interactive review inside a Droid session: ```text > /review ``` Non-interactive review through `droid exec`: ```bash # Analysis only droid exec "Review this PR for security issues" # With modifications droid exec --auto low "Review code and add missing type hints" ``` See the [Local Code Review documentation](/software-factory/code-review) for detailed guidance on review types, workflows, and best practices. ### Testing and debugging ```bash # Investigation droid exec "Analyze failing tests and explain root cause" # Fix and verify droid exec --auto medium "Fix failing tests and run test suite" ``` ### Refactoring ```bash # Planning droid exec "Create refactoring plan for auth module" # Execution droid exec --auto low --use-spec "Refactor auth module" ``` ### Parallel sessions on one repo ```bash # Work on two branches of the same repo at the same time, each in its own worktree droid --worktree feature-a & droid --worktree feature-b & # Fan out headless tasks across branches without clobbering each other's files droid exec -w migration-step-1 --auto medium "apply codemod A" & droid exec -w migration-step-2 --auto medium "apply codemod B" & wait ``` See [Git worktrees](#git-worktrees) for details. ### CI/CD integration ```yaml # GitHub Actions example - name: Run Droid Analysis env: FACTORY_API_KEY: ${{ secrets.FACTORY_API_KEY }} run: | droid exec --auto medium -f .github/prompts/deploy.md ``` Model IDs and reasoning options. Create your own shortcuts. Build specialized agents. External tool integration. # Droid Sessions API Create and drive Droid sessions: manage their lifecycle, settings, and messages. ## List sessions `GET /api/v0/sessions` Returns a paginated list of sessions for the authenticated user. This feature is enabled for selected organizations only. ```bash curl 'https://api.factory.ai/api/v0/sessions' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `limit` (`string`) - Query parameter. Maximum number of items to return (1-100) - `cursor` (`string`) - Query parameter. Cursor for pagination - `computerId` (`string`) - Query parameter. Computer ID to query directly **Response:** `200` - Response for status 200 ## Create a session `POST /api/v0/sessions` Creates a new session with the specified configuration. This feature is enabled for selected organizations only. ```bash curl -X POST 'https://api.factory.ai/api/v0/sessions' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Response:** `201` - Response for status 201 ## Get a session `GET /api/v0/sessions/{sessionId}` Returns detailed session information including settings and stats. This feature is enabled for selected organizations only. ```bash curl 'https://api.factory.ai/api/v0/sessions/{sessionId}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `sessionId` (`string`, required) - Path parameter. Session ID - `computerId` (`string`) - Query parameter. Computer ID to query directly **Response:** `200` - Response for status 200 ## Delete a session `DELETE /api/v0/sessions/{sessionId}` Soft-deletes a session (can be restored within retention period). This feature is enabled for selected organizations only. ```bash curl -X DELETE 'https://api.factory.ai/api/v0/sessions/{sessionId}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `sessionId` (`string`, required) - Path parameter. Session ID - `computerId` (`string`) - Query parameter. Computer ID to query directly **Response:** `204` - Response for status 204 ## Update a session `PATCH /api/v0/sessions/{sessionId}` Updates session configuration such as model and reasoning effort. This feature is enabled for selected organizations only. ```bash curl -X PATCH 'https://api.factory.ai/api/v0/sessions/{sessionId}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Parameters** - `sessionId` (`string`, required) - Path parameter. Session ID **Response:** `200` - Response for status 200 ## Interrupt a session `POST /api/v0/sessions/{sessionId}/interrupt` Interrupts a running agent loop. Idempotent if already idle. This feature is enabled for selected organizations only. ```bash curl -X POST 'https://api.factory.ai/api/v0/sessions/{sessionId}/interrupt' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `sessionId` (`string`, required) - Path parameter **Response:** `200` - Response for status 200 ## Get session messages `GET /api/v0/sessions/{sessionId}/messages` Returns paginated message history with optional role filtering. This feature is enabled for selected organizations only. ```bash curl 'https://api.factory.ai/api/v0/sessions/{sessionId}/messages' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `sessionId` (`string`, required) - Path parameter - `limit` (`string`) - Query parameter. Maximum number of items to return (1-100) - `cursor` (`string`) - Query parameter. Cursor for pagination - `computerId` (`string`) - Query parameter. Computer ID to query directly - `role` (`string`) - Query parameter. Filter messages by role Allowed values: user, assistant, tool. **Response:** `200` - Response for status 200 ## Add a message to a session `POST /api/v0/sessions/{sessionId}/messages` Adds a user message and optionally waits for agent completion. Supports text, images, and file attachments. This feature is enabled for selected organizations only. ```bash curl -X POST 'https://api.factory.ai/api/v0/sessions/{sessionId}/messages' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Parameters** - `sessionId` (`string`, required) - Path parameter **Response:** `200` - Response for status 200 ## Get a message by ID `GET /api/v0/sessions/{sessionId}/messages/{messageId}` Returns a single message from the session by its ID. This feature is enabled for selected organizations only. ```bash curl 'https://api.factory.ai/api/v0/sessions/{sessionId}/messages/{messageId}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `sessionId` (`string`, required) - Path parameter. Session ID - `messageId` (`string`, required) - Path parameter. Message ID - `computerId` (`string`) - Query parameter. Computer ID to query directly **Response:** `200` - Response for status 200 # Droid Exec (Headless) Non-interactive execution mode for CI/CD pipelines and automation scripts. Droid Exec is Factory's headless execution mode designed for automation workflows. Unlike the interactive CLI, `droid exec` runs as a one-shot command that completes a task and exits, making it ideal for CI/CD pipelines, shell scripts, and batch processing. Explore the [Droid Exec cookbooks](/software-factory/code-review-ci). ## Summary and goals Droid Exec is a one-shot task runner: it produces readable logs (and structured artifacts when requested), keeps mutations and command execution opt-in so runs are secure by default, fails fast on permission violations with clear errors, and composes simply for batch and parallel work. Single run execution that writes to stdout/stderr for CI/CD integration Read-only by default with explicit opt-in for mutations via autonomy levels Designed for shell scripting, parallel execution, and pipeline integration Structured output formats and artifacts for automated processing ## Execution model Each run is a single non-interactive pass that writes to stdout/stderr. The default is spec-mode, where the agent is only allowed to execute read-only operations; add `--auto` to enable edits and commands, with risk tiers gating what can run. CLI help (excerpt): ```text Usage: droid exec [options] [prompt] Execute a single command (non-interactive mode) Arguments: prompt The prompt to execute Options: -o, --output-format Output format (default: "text") --input-format Input format: stream-json for multi-turn sessions; stream-jsonrpc is controlled via JSON-RPC requests -f, --file Read prompt from file --auto Autonomy level: low|medium|high --skip-permissions-unsafe Skip ALL permission checks - allows all permissions (unsafe) -s, --session-id Existing session to continue (requires a prompt) --fork Fork an existing session and continue from it -m, --model Model ID to use -r, --reasoning-effort Reasoning effort (defaults per model) --spec-model Model ID to use for spec mode --spec-reasoning-effort Reasoning effort for spec mode --use-spec Start in spec mode --restrict-tools Restrict the run to only the specified tools (comma or space separated list) --additional-tools Enable additional tools beyond the defaults (comma or space separated list) --disabled-tools Disable specific tools (comma or space separated list) --disable-builtin-skills Disable Factory-provided builtin skills for this session --list-tools List available tools for the selected model and exit --cwd Working directory path -w, --worktree [name] Run in a git worktree --worktree-dir Directory for worktree creation --tag Session tag (name or JSON, repeatable) --log-group-id Log group ID for filtering logs --append-system-prompt Append custom text to end of system prompt --append-system-prompt-file Append file contents to end of system prompt --mission Run in mission mode (multi-agent orchestration) --worker-model Model for mission workers --worker-reasoning-effort Reasoning effort for mission workers --validator-model Model for mission validators --validator-reasoning-effort Reasoning effort for mission validators -h, --help display help for command ``` Use any [available model ID](/models) with `--model` or `--spec-model`. For custom models, see [Custom Models (BYOK)](/model-independence/byok). For multi-turn sessions, use `--input-format stream-jsonrpc` with a matching `--output-format stream-jsonrpc` and drive the session over [raw JSON-RPC](#build-custom-flows-on-raw-json-rpc). The older `stream-json` input mode still parses but is deprecated and prints a warning. ## Installation **macOS/Linux:** `curl -fsSL https://app.factory.ai/cli | sh` **Homebrew:** `brew install --cask droid` **Windows:** `irm https://app.factory.ai/cli/windows | iex` **npm:** `npm install -g droid` Generate your API key from the Factory Settings Page Export your API key as an environment variable: ```bash export FACTORY_API_KEY=fk-... ``` ## Quickstart ### Direct prompt ```bash droid exec "analyze code quality" droid exec "fix the bug in src/main.js" --auto low ``` ### From file ```bash droid exec -f prompt.md ``` ### Pipe ```bash echo "summarize repo structure" | droid exec ``` ### Session continuation ```bash droid exec --session-id "continue with next steps" ``` ## Autonomy Levels Droid exec uses tiered autonomy to control what operations can run without manual confirmation. It starts read-only by default; use explicit flags when an automation needs to modify files, install dependencies, or touch external systems. | Mode | Allows | Blocks or limits | Best for | | :--- | :----- | :--------------- | :------- | | Default, no flags | Read-only file inspection, directory listing, process and environment inspection, and git read operations such as `git status`, `git log`, and `git diff`. | File edits, package installs, git writes, system changes, and deployments. | Analysis, planning, architecture review, security review, and reports. | | `--auto low` | Low-risk project file creation and edits. | System modifications, package installs, remote writes, and production changes. | Documentation updates, formatting, adding comments, and simple local edits. | | `--auto medium` | Low-risk actions plus common local development operations such as trusted package installs, builds, tests, `git commit`, `git checkout`, and `git pull`. | `git push`, sudo commands, production changes, and sensitive or irreversible operations. | Local development, dependency updates, test fixes, and CI preparation. | | `--auto high` | High-risk automation including remote writes, deployments, database migrations, and commands that may access sensitive data, subject to Droid's remaining hard safety checks. | Extreme destructive operations such as `sudo rm -rf /` and system-wide unsafe changes. | CI/CD pipelines, staging deployments, and fully automated workflows in controlled environments. | | `--skip-permissions-unsafe` | All operations without confirmation. | Cannot be combined with `--auto` flags. Use only where the environment is disposable. | Throwaway containers, ephemeral CI runners, and temporary VMs. | Examples: ```bash # Analyze and plan without making changes droid exec "Analyze the authentication system and list the files that would need changes for an OAuth2 migration." # Allow simple file edits droid exec --auto low "add JSDoc comments to all exported functions" # Allow normal local development operations droid exec --auto medium "install deps, run tests, and fix failing checks" # Allow deployment-oriented automation in a controlled environment droid exec --auto high "fix the bug, run tests, commit, and push the branch" ``` `--skip-permissions-unsafe` bypasses all permission checks. Only use it in completely isolated environments such as Docker containers, throwaway VMs, or ephemeral CI runners that are destroyed after the job. ```bash # In a disposable Docker container for CI testing docker run --rm -v $(pwd):/workspace alpine:latest sh -c " apk add curl bash && curl -fsSL https://app.factory.ai/cli | sh && droid exec --skip-permissions-unsafe 'Install dependencies, modify system configs, run privileged integration tests, and clean up test databases' " ``` ### Fail-fast behavior If a requested action exceeds the current autonomy level, droid exec stops immediately with a clear error message, returns a non-zero exit code, and performs no partial changes. This ensures predictable behavior in automation scripts and CI/CD pipelines. ## Output formats and artifacts Droid exec supports three output formats for different use cases: ### text (default) Human-readable output for direct consumption or logs: ```bash $ droid exec --auto low "create a python file that prints 'hello world'" Perfect! I've created a Python file named `hello_world.py` in your home directory that prints 'hello world' when executed. ``` ### json Structured JSON output for parsing in scripts and automation: ```bash $ droid exec "summarize this repository" --output-format json { "type": "result", "subtype": "success", "is_error": false, "duration_ms": 5657, "num_turns": 1, "result": "This is a Factory documentation repository containing guides for CLI tools, web platform features, and onboarding procedures...", "session_id": "8af22e0a-d222-42c6-8c7e-7a059e391b0b" } ``` Use JSON format when you need to parse the result in a script, check success/failure programmatically, extract session IDs for continuation, or process results in a pipeline. ### Build custom flows on raw JSON-RPC For custom integrations, you can run Droid as a long-lived subprocess and drive the full JSON-RPC control surface over stdin/stdout: ```bash droid exec --input-format stream-jsonrpc --output-format stream-jsonrpc --auto low ``` This is the lowest-level integration path for building your own interaction model around Droid. Each stdin line is one JSON-RPC request; each stdout line is a JSON-RPC response, server request, or notification. Your client spawns `droid exec` with the project `cwd` and desired flags, writes newline-delimited requests with unique IDs, reads stdout line-by-line to match responses by `id`, and implements timeouts, process cleanup, and session persistence. The core methods and events: Start a new session; a client calls this (or `droid.load_session` to resume an existing session) first. Send a turn to the session. Event stream your client handles: assistant text deltas, tool events, token usage, errors, and turn completion. Server-to-client request your client must answer; `droid.ask_user` arrives the same way. Further session methods interrupt work, update settings, manage MCP servers/tools, inspect context, fork sessions, or compact history. On top of raw stdin/stdout you can build web, desktop, or IDE agent experiences with your own UX and controls; chat or copiloting surfaces that route user actions into Droid turns; CI and workflow runners that execute Droid tasks and surface progress in build logs; orchestrators that queue work, resume sessions, fork conversations, and persist results; policy layers that approve, deny, transform, or audit tool permission requests; and bridges from Droid events into your own protocol, message bus, telemetry, or storage layer. To disable Factory-provided builtin skills, set `disableBuiltinSkills` when initializing or loading the session: ```json {"jsonrpc":"2.0","id":"1","method":"droid.initialize_session","params":{"machineId":"my-machine","cwd":"/workspace","disableBuiltinSkills":true}} {"jsonrpc":"2.0","id":"2","method":"droid.load_session","params":{"sessionId":"session-123","disableBuiltinSkills":true}} ``` In stream JSON-RPC mode, use the request field instead of the `--disable-builtin-skills` CLI flag. The TypeScript SDK exposes the same option through `createSession({ disableBuiltinSkills: true })` and `resumeSession(sessionId, { disableBuiltinSkills: true })`. For protocol reference and implementation patterns, see the low-level client and process transport in the [TypeScript SDK](/sdk/typescript). Prefer an SDK when possible: - **TypeScript**: [`@factory/droid-sdk`](/sdk/typescript) for Node.js apps, streaming, multi-turn sessions, structured output, permissions, tool controls, and SDK-backed MCP tools - **Python**: [`droid-sdk`](/sdk/python) for asyncio apps, streaming, direct client control, notifications, permissions, and typed event handling For automated pipelines, you can also direct the agent to write specific artifacts: ```bash droid exec --auto low "Analyze dependencies and write to deps.json" droid exec --auto low "Generate metrics report in CSV format to metrics.csv" ``` ## Working directory Use `--cwd` to scope execution: ```bash droid exec --cwd /home/runner/work/repo "Map internal packages and dump graphviz DOT to deps.dot" ``` Use `-w, --worktree [name]` to run the task inside an isolated [git worktree](/droid-cli/cli-reference#git-worktrees) on its own branch. This is useful for fanning out parallel `droid exec` jobs against the same repo without file conflicts: ```bash droid exec --worktree codemod-a --auto medium "apply codemod A" & droid exec --worktree codemod-b --auto medium "apply codemod B" & wait ``` Clean worktrees are auto-removed on exit; dirty ones are preserved so you can review and push the work. ## Models and reasoning effort Choose a model with `-m` and adjust reasoning with `-r`. See the [Factory-Managed Inference](/models) for available models. ```bash droid exec -m claude-sonnet-4-5-20250929 -r medium -f plan.md ``` Use `--use-spec` to start in specification mode, where the agent plans before executing: ```bash droid exec --use-spec --auto low "refactor the auth module" ``` You can also use a different model for the spec phase: ```bash droid exec --use-spec --spec-model claude-haiku-4-5-20251001 --auto medium "implement feature X" ``` ### Tool controls List available tools for a model: ```bash droid exec --list-tools droid exec --model gpt-5.3-codex --list-tools --output-format json ``` Enable or disable specific tools: ```bash # Restrict to only specific tools droid exec --auto low --restrict-tools ApplyPatch "refactor files" # Enable additional tools droid exec --additional-tools ApplyPatch "refactor files" # Disable specific tools droid exec --auto medium --disabled-tools execute-cli "run edits only" ``` ### Skill controls Use `--disable-builtin-skills` when an interactive or non-interactive session should expose only skills supplied outside the Droid CLI: ```bash # Start the interactive TUI without builtin skills droid --disable-builtin-skills # Run one non-interactive session without builtin skills droid exec --disable-builtin-skills "analyze this repository" ``` This removes Factory-provided builtin skills while preserving user, project, plugin, automation, mission, runtime, and dynamically supplied skills. For an interactive launch, the restriction remains active when you create, resume, or switch sessions. It is also inherited by child sessions, including subagents and mission workers. The flag is not supported in Agent Client Protocol (ACP) mode. Use it with standard `droid` or `droid exec`, or set `disableBuiltinSkills` on `droid.initialize_session` and `droid.load_session` requests in stream JSON-RPC mode. ### Custom models You can configure custom models to use with droid exec by adding them to your `~/.factory/settings.json` file: ```json { "customModels": [ { "model": "my-codex-model", "displayName": "My Custom Model", "baseUrl": "https://api.openai.com/v1", "apiKey": "${OPENAI_API_KEY}", "provider": "openai" } ] } ``` To use a custom model, use the `custom:` prefix followed by the display name (with spaces replaced by dashes) and the index: ```bash droid exec --model "custom:My-Custom-Model-0" "analyze this codebase" ``` If you have multiple custom models configured: ```json { "customModels": [ { "model": "moonshotai/kimi-k2-instruct-0905", "displayName": "Kimi K2 [Groq]", "baseUrl": "https://api.groq.com/openai/v1", "apiKey": "${GROQ_API_KEY}", "provider": "generic-chat-completion-api", "maxOutputTokens": 16384 }, { "model": "openai/gpt-oss-20b", "displayName": "GPT-OSS-20B [OpenRouter]", "baseUrl": "https://openrouter.ai/api/v1", "apiKey": "YOUR_OPENROUTER_KEY", "provider": "generic-chat-completion-api", "maxOutputTokens": 32000 } ] } ``` You would reference them as: - `--model "custom:Kimi-K2-[Groq]-0"` - `--model "custom:GPT-OSS-20B-[OpenRouter]-1"` The index corresponds to the position in the `customModels` array (0-based). Reasoning effort (`-r` / `--reasoning-effort`) does not apply to custom models; it is controlled by the model and provider you configure. ## Sessions, tagging, and logs Forking lets you branch off an existing session without disturbing the original; the new run starts from the forked session's history and is assigned a fresh session ID. ```bash # Continue a session in-place droid exec --session-id "next steps" # Branch off a session into a new run droid exec --fork --auto low "try an alternative refactor" ``` Use `--tag` to attach searchable labels to a run. The flag is repeatable and accepts either a plain name or a JSON object for structured metadata. Pair it with `--log-group-id` to bucket logs from related runs together for easier filtering and aggregation downstream. ```bash droid exec \ --tag release-v2 \ --tag '{"name":"team","value":"platform"}' \ --log-group-id ci-nightly \ --auto low "run nightly checks and write report.md" ``` ## Customizing the system prompt Use `--append-system-prompt` to append additional text to the end of the system prompt for a single run, or `--append-system-prompt-file` to append the contents of a file. Both flags can be combined and are useful for injecting project-specific guidance, style guides, or invariants without modifying global settings. ```bash droid exec \ --append-system-prompt "Always prefer functional React components." \ --auto low "review src/components for class-based components" droid exec \ --append-system-prompt-file ./docs/style-guide.md \ --auto low "lint prose in README.md against the style guide" ``` ## Mission Mode Mission Mode runs `droid exec` as a multi-agent orchestrator that plans work, delegates to worker agents, and validates results. Enable it with `--mission` and optionally select dedicated models and reasoning effort levels for the worker and validator roles. Mission Mode requires `--auto high` or `--skip-permissions-unsafe` because the orchestrator needs high autonomy to manage workers. ```bash droid exec --mission \ --worker-model claude-sonnet-4-5-20250929 \ --worker-reasoning-effort medium \ --validator-model claude-sonnet-4-5-20250929 \ --validator-reasoning-effort high \ --auto high \ "ship the new billing webhook end-to-end" ``` ### Mission Mode flags Run in Mission Mode (multi-agent orchestration). Model ID used by mission worker agents. Reasoning effort for mission workers. Model ID used by mission validator agents. Reasoning effort for mission validators. The top-level `-m, --model` and `-r, --reasoning-effort` flags still apply to the orchestrator itself; the worker and validator overrides only affect the agents the orchestrator spawns. ## Batch and parallel patterns Shell loops (bounded concurrency): ```bash # Process files in parallel (GNU xargs -P) find src -name "*.ts" -print0 | xargs -0 -P 4 -I {} \ droid exec --auto low "Refactor file: {} to use modern TS patterns" ``` Background job parallelization: ```bash # Process multiple directories in parallel with job control for path in packages/ui packages/models apps/factory-app; do ( cd "$path" && droid exec --auto low "Run targeted analysis and write report.md" ) & done wait # Wait for all background jobs to complete ``` Chunked inputs: ```bash # Split large file lists into manageable chunks git diff --name-only origin/main...HEAD | split -l 50 - /tmp/files_ for f in /tmp/files_*; do list=$(tr '\n' ' ' < "$f") droid exec --auto low "Review changed files: $list and write to review.json" done rm /tmp/files_* # Clean up temporary files ``` Workflow Automation (CI/CD): ```yaml # Dead code detection and cleanup suggestions name: Code Cleanup Analysis on: schedule: - cron: '0 1 * * 0' # Weekly on Sundays workflow_dispatch: jobs: cleanup-analysis: strategy: matrix: module: ['src/components', 'src/services', 'src/utils', 'src/hooks'] steps: - uses: actions/checkout@v4 - run: droid exec --cwd "${{ matrix.module }}" --auto low "Identify unused exports, dead code, and deprecated patterns. Generate cleanup recommendations in cleanup-report.md" ``` Pull request review, security review, and PR description fill already ship as a packaged action, [`Factory-AI/droid-action`](https://github.com/Factory-AI/droid-action), so you do not need to script those flows around `droid exec`. See [Automated Code Review](/software-factory/code-review-ci) to set it up. ## Unique usage examples License header enforcer: ```bash git ls-files "*.ts" | xargs -I {} \ droid exec --auto low "Ensure {} begins with the Apache-2.0 header; add it if missing" ``` API contract drift check (read-only): ```bash droid exec "Compare openapi.yaml operations to our TypeScript client methods and write drift.md with any mismatches" ``` Security sweep: ```bash droid exec --auto low "Run a quick audit for sync child_process usage and propose fixes; write findings to sec-audit.csv" ``` ## Exit behavior `droid exec` exits `0` on success and non-zero on failure (permission violation, tool error, unmet objective). Treat non-zero as failed in CI. ## Best practices - Favor `--auto low`; keep mutations minimal and commit/push in scripted steps. - Avoid `--skip-permissions-unsafe` unless fully sandboxed. - Ask the agent to emit artifacts your pipeline can verify. - Use `--cwd` to constrain scope in monorepos. Orchestrate multi-agent missions from the command line. Wire Droid Exec into GitHub Actions for automated PR review. # Droid Python SDK Use the Factory Droid SDK for Python to build custom agents, workflows, and product integrations. The Droid SDK lets you run the same agent harness that powers Factory's CLI, desktop application, and web platform from your own code. It manages conversation context, tool execution, permissions, streaming, model selection, and multi-step execution. Your application decides when Droid runs, which tools and models it can use, and what happens to the result. ## Overview The Python SDK runs a local `droid` subprocess and provides one-shot runs, persistent sessions, streaming, and model and session discovery. ### Choose an API | Goal | API | | --- | --- | | Run one prompt and return its result | `await run(...)` | | Keep context across prompts | `Session(...)` | | Continue a saved session | `Session.resume(...)` | | Discover selectable models | `await list_models(...)` | | Find saved sessions | `await list_sessions(...)` | Use `Session` when prompts need shared history or session operations such as interruption, compaction, or rewind. There is no `session.run()`; every session turn goes through `stream()`. ### What you can build Invoke Droid when a pull request opens, a support ticket is escalated, an incident is created, or scheduled repository maintenance is due. The agent does not need a chat interface. ## Install and authenticate Install the `droid-sdk` package and import it as `droid_sdk`: ```bash pip install droid-sdk ``` The package provides an async API for running Droid locally. Requirements: - Python 3.10 or later - `droid` on `PATH` - an authenticated Droid CLI session or a Factory API key The SDK uses the Droid CLI's local authentication by default. It reads `FACTORY_API_KEY` when present; pass `api_key=` only when the key comes from application configuration. The SDK never places keys in command arguments, logs, or exception messages. ## Quick start ### Run one prompt ```python import asyncio from droid_sdk import run async def main() -> None: result = await run("Summarize this repository.") if result.success: print(result.text) else: print(result.subtype) asyncio.run(main()) ``` `run()` starts a session, runs one turn, and closes everything it created. The saved session remains resumable. See [Result types](#result-types) for every terminal outcome. ### Continue a conversation ```python import asyncio from pathlib import Path from droid_sdk import Session async def run_turn(session: Session, prompt: str) -> str: async with session.stream(prompt) as stream: async for _ in stream: pass result = stream.result if not result.success: raise RuntimeError(result.subtype) return result.text async def main() -> None: async with Session(cwd=Path.cwd()) as session: print(await run_turn(session, "What does this project do?")) print(await run_turn(session, "What should I test first?")) asyncio.run(main()) ``` The second turn includes context from the first. Exiting the session closes the subprocess and session-owned resources. See [Default stream types](#default-stream-types) for every yielded value. ## Core concepts ### Session A `Session` owns conversation history, settings, a working directory, and one Droid connection. It accepts one active turn at a time. ### Turn A turn begins with `session.stream(prompt)` and ends when the stream yields a `RunResult`. Create separate sessions for parallel work. ### Result A turn's outcome is a value, not an exception: `RunSuccess`, `RunInterrupted`, or `RunFailure`. Interrupted turns, execution failures, and structured-output failures do not raise. Failures in the machinery around a turn (setup, connection, process, protocol, timeout, and cancellation) raise exceptions. ### Ownership and cleanup | Object | Owner | | --- | --- | | Resources created by `run()` | `run()` | | `Session` used with `async with` | The context manager | | Manually opened `Session` | The caller | | In-process MCP servers | The session | ### Typing The package includes `py.typed`, and the public API type-checks under strict Pyright and mypy. Public unions narrow with `isinstance()`. High-level value models are immutable data classes. | Call | Static return type | | --- | --- | | `run(..., output=None)` | `RunResult[None]` | | `run(..., output=Model)` | `RunResult[Model]` | | `run(..., output=JsonSchema(...))` | `RunResult[JsonObject]` | | `session.stream(...)` | `RunStream[T, StreamMessage[T]]` | | Partial `session.stream(...)` | `RunStream[T, StreamEvent[T]]` | | `stream.result` | `RunResult[T]` | ## Sessions Use a session when turns share history. ### Create and use a session `Session` is lazy. It creates the subprocess and session on entry: ```python async with Session( cwd=Path.cwd(), model="auto", ) as session: async with session.stream("Review the project.") as stream: async for message in stream: handle(message) ``` `model="auto"` selects the [Factory Router](#use-the-factory-router). Configure behavior and attach handlers with `SessionConfig` and `InteractionHandlers`: ```python from droid_sdk import ( Autonomy, InteractionHandlers, PermissionRequest, PermissionResponse, SessionConfig, ToolConfirmationOutcome, ) def approve(request: PermissionRequest) -> PermissionResponse: return request.respond(ToolConfirmationOutcome.PROCEED_ONCE) config = SessionConfig( autonomy=Autonomy.LOW, disabled_tools={"Execute"}, # any iterable of tool IDs works here disable_builtin_skills=True, ) async with Session( config=config, interactions=InteractionHandlers(on_permission=approve), ) as session: ... ``` Handlers are plain callables that inspect the request and choose an offered outcome; see [Permissions and user input](#permissions-and-user-input) for the full contract. `SessionConfig` also accepts mode-specific models, MCP servers, tags, source attribution, `machine_id`, automatic permission rejection, and native-tool overrides. ### Customize the system prompt Set `system_prompt` when creating a session. A string replaces Droid's default behavioral prompt: ```python from droid_sdk import Session, SessionConfig config = SessionConfig( system_prompt="Act as a focused dependency-analysis agent.", ) async with Session(config=config) as session: ... ``` To keep Droid's default prompt and append instructions: ```python config = SessionConfig( system_prompt={ "type": "preset", "preset": "droid", "append": "Prioritize security findings and cite relevant files.", } ) ``` System prompts are set at session creation and preserved when the session is resumed or forked. Older Droid versions that do not support custom system prompts raise `SessionError`. ### Resume a saved session ```python async with Session.resume( session_id, interactions=InteractionHandlers(on_question=answer), disabled_tools={"Execute"}, ) as session: async with session.stream("Continue from the previous prompt.") as stream: async for _ in stream: pass ``` Get a session ID from `session.id`, `result.session_id`, or [`list_sessions()`](#list-saved-sessions). Resume restores conversation history, working directory, title, and settings; runtime concerns such as handlers and MCP servers must be attached again (see the table below). It does not accept a new working directory or model. ### Read session state ```python print(session.id) print(session.cwd) print(session.settings.model) print(session.settings.mode) ``` These properties are read-only. `settings` is an immutable snapshot, replaced when Droid reports a settings change. ### Update session settings ```python await session.update_settings( model="model-id", reasoning_effort=ReasoningEffort.HIGH, autonomy=Autonomy.MEDIUM, disabled_tools={"Execute"}, ) ``` Only supplied fields change. Set nullable Spec-model fields to `None` to clear them. `update_settings()` accepts: | Field | Type | | --- | --- | | `model` | `str \| None` | | `reasoning_effort` | `ReasoningEffort \| None` | | `mode` | `Mode \| None` | | `autonomy` | `Autonomy \| None` | | `spec_model` | `str \| None` | | `spec_reasoning_effort` | `ReasoningEffort \| None` | | `tags` | `Sequence[SessionTag] \| None` | | `compaction_token_limit` | `int \| None` | | `compaction_threshold_check_enabled` | `bool \| None` | | `additional_tools` | `Iterable[str] \| None` | | `enabled_tools` | `Iterable[str] \| None` | | `disabled_tools` | `Iterable[str] \| None` | | `restrict_tools` | `Iterable[str] \| None` | It returns `UpdateSettingsResult`. The result currently has no fields. ### Rename a session ```python await session.rename("Authentication review") ``` `rename(title: str)` returns `None`. ### List saved sessions ```python from droid_sdk import list_sessions for saved in await list_sessions(limit=10): print(saved.id, saved.title, saved.modified_at) ``` Pass `all_workspaces=True` to list sessions across working directories. `list_sessions()` reads local files without starting Droid. Results are newest first. Timestamps are timezone-aware. ### What persists | Restored | Attach again | | --- | --- | | Conversation history | Interaction handlers | | Working directory | Session-scoped MCP servers | | Title | Observability sinks | | Session settings | | ### Close sessions safely ```python session = Session() await session.open() try: async with session.stream("Review the project.") as stream: async for _ in stream: pass finally: await session.close() ``` `open()` and `close()` are idempotent, but a closed session cannot be reopened. Calling an active method before `open()` raises `SessionNotOpenError`. Concurrent `open()` calls share one startup attempt. If one waiter is cancelled, startup continues for the others. Cancelling the final waiter cancels startup and completes resource cleanup before the session becomes retryable. `close()` racing startup waits for startup cleanup and leaves the session closed. ### Subscribe to raw notifications ```python unsubscribe = session.on_notification(handle_notification, type="custom_type") try: ... finally: unsubscribe() ``` High-level streams ignore unknown notifications. `on_notification()` exposes them without creating a second client. ## Streaming and results Use top-level `run()` when only the result matters. Session turns always use `stream()`. `session.stream()` accepts: | Argument | Type | | --- | --- | | `prompt` | `str` | | `images` | `Sequence[Image]` | | `files` | `Sequence[Document]` | | `output` | `type[BaseModel] \| JsonSchema \| None` | | `timeout` | `float \| None` | | `include_partial_messages` | `bool` | ### Complete messages ```python from droid_sdk import AssistantMessage async with session.stream("Run the tests.") as stream: async for message in stream: if isinstance(message, AssistantMessage): print(message.text) ``` ### Default stream types ```python StreamMessage[T] = ( UserMessage | AssistantMessage | ToolCall | ToolResult | HookExecution | ErrorEvent | RunResult[T] ) ``` `session.stream(prompt)` yields: | Type | Fields | | --- | --- | | `UserMessage` | Conversation-message fields | | `AssistantMessage` | Conversation-message fields | | `ToolCall` | `name`, `tool_use_id`, `input` | | `ToolResult` | `tool_use_id`, `tool_name`, `content`, `is_error` | | `HookExecution` | `hook_id`, `event_name`, `matcher`, `tool_call_id`, `command`, `timeout`, `status`, `exit_code`, `stdout`, `stderr`, `suppress_output` | | `ErrorEvent` | `message`, `error_type`, `timestamp` | | `RunResult[T]` | Terminal result described below | Values appear in delivery order. `RunResult[T]` is always last. ### Partial events ```python from droid_sdk import Session, TextDelta async with Session() as session: async with session.stream( "Explain the failing test.", include_partial_messages=True, ) as stream: async for event in stream: if isinstance(event, TextDelta): print(event.text, end="", flush=True) print() if not stream.result.success: print(f"Turn ended: {stream.result.subtype}") ``` The default stream yields complete messages. Set `include_partial_messages=True` to add deltas and operational events. Both modes yield the terminal result and cache it in `stream.result`. See [Partial stream types](#partial-stream-types) for the complete event union. ```python from droid_sdk import RunResult, TextDelta async with session.stream( "Explain the failing test.", include_partial_messages=True, ) as stream: async for event in stream: if isinstance(event, TextDelta): print(event.text, end="", flush=True) elif isinstance(event, RunResult) and not event.success: print(f"\nTurn ended: {event.subtype}") ``` ### Partial stream types ```python StreamEvent[T] = ( StreamMessage[T] | TextDelta | TextComplete | ThinkingDelta | ThinkingComplete | ToolCallDelta | ToolProgress | TokenUsageUpdate | WorkingStateChanged | PermissionResolved | SettingsUpdated | SessionTitleUpdated | SessionWorkingDirectoryChanged | McpStatusChanged | McpAuthRequired | McpAuthCompleted ) ``` With `include_partial_messages=True`, the stream yields every `StreamMessage[T]` plus: | Type | Fields | | --- | --- | | `TextDelta` | `message_id`, `block_index`, `text` | | `TextComplete` | `message_id`, `block_index` | | `ThinkingDelta` | `message_id`, `block_index`, `text` | | `ThinkingComplete` | `message_id`, `block_index`, `duration` | | `ToolCallDelta` | `tool_use` | | `ToolProgress` | `tool_use_id`, `tool_name`, `content`, `update` | | `TokenUsageUpdate` | Token-usage fields | | `WorkingStateChanged` | `state` | | `PermissionResolved` | `request_id`, `tool_use_ids`, `selected_option` | | `SettingsUpdated` | `settings` | | `SessionTitleUpdated` | `title` | | `SessionWorkingDirectoryChanged` | `cwd` | | `McpStatusChanged` | `servers`, `summary` | | `McpAuthRequired` | `server_name`, `auth_url`, `message`, `state` | | `McpAuthCompleted` | `server_name`, `outcome`, `message` | `ToolProgress.update` is a `ToolProgressUpdate`: | Type | Fields | | --- | --- | | `ToolProgressUpdate` | `type`, `tool_name`, `status`, `details`, `text`, `error`, `timestamp`, `parameters`, `value_snippet`, `terminal_id`, `full_output`, `subagent_session_id` | `ToolProgressUpdate.type` is `"tool_call"`, `"tool_result"`, `"error"`, `"status"`, or `"message"`. Unknown high-level events are ignored. Use `session.on_notification()` when the application needs raw notifications. ### Message model ```python from droid_sdk import AssistantMessage, ConversationMessage, UserMessage completed: list[ConversationMessage] = [] async with session.stream("Explain the failing test.") as stream: async for event in stream: if isinstance(event, (UserMessage, AssistantMessage)): completed.append(event) save(event) ``` `ConversationMessage` defines the fields shared by user and assistant messages: | Field | Type | Meaning | | --- | --- | --- | | `id` | `str` | Stable message ID | | `content` | `tuple[ContentBlock, ...]` | Ordered canonical content | | `text` | `str` | Visible text blocks joined in order | | `parent_id` | `str \| None` | Parent message, when present | | `created_at` | `datetime` | Timezone-aware creation time | | `updated_at` | `datetime` | Timezone-aware update time | `ContentBlock` is a typed union for text, thinking, tool use, tool result, redacted thinking, images, and documents. Narrow blocks with `isinstance()`. ```python ContentBlock = ( TextBlock | ThinkingBlock | RedactedThinkingBlock | ToolUseBlock | ToolResultBlock | ImageBlock | DocumentBlock ) ``` | Type | Fields | | --- | --- | | `TextBlock` | `id`, `text` | | `ThinkingBlock` | `id`, `thinking`, `signature`, `signature_provider`, `duration` | | `RedactedThinkingBlock` | `id`, `data` | | `ToolUseBlock` | `id`, `name`, `input`, `thought_signature` | | `ToolResultBlock` | `id`, `tool_use_id`, `content`, `is_error` | | `ImageBlock` | `id`, `source`, `generated` | | `DocumentBlock` | `id`, `source` | `Message` is `StreamMessage[T]` without `RunResult[T]`. Partial events are not messages. `TextDelta.message_id` and `block_index` identify the assistant content block that eventually appears in `AssistantMessage.content`. `ToolCall.tool_use_id` matches the corresponding `ToolUseBlock.id`. A complete message can repeat text already delivered through `TextDelta`. Use deltas for live rendering and complete messages for persistence. Do not concatenate both. `stream.result.messages` is the ordered tuple of complete `Message` values emitted before the result. It excludes the result and partial events. ### Handle the result The terminal `RunResult` is yielded by the iterator and cached on the stream: ```python from droid_sdk import RunResult async with session.stream("Run the tests.") as stream: async for message in stream: if isinstance(message, RunResult): print(message.subtype) result = stream.result if result.success: print(result.text) ``` Reading `stream.result` before completion raises `StreamIncompleteError`. ### Result types `RunResult[T]` is a union of three terminal states: ```python RunResult[T] = RunSuccess[T] | RunInterrupted[T] | RunFailure[T] ``` | Type | Subtype | `success` | `interrupted` | | --- | --- | --- | --- | | `RunSuccess[T]` | `success` | `True` | `False` | | `RunInterrupted[T]` | `interrupted` | `False` | `True` | | `RunFailure[T]` | `error_during_execution` | `False` | `False` | | `RunFailure[T]` | `error_structured_output` | `False` | `False` | | Field | Type | Meaning | | --- | --- | --- | | `subtype` | `str` | Terminal state | | `text` | `str` | Final assistant text; reconstructed from deltas if no complete message arrived | | `messages` | `tuple[Message, ...]` | Complete messages from the turn | | `usage` | `Usage \| None` | Per-turn token and credit usage | | `duration` | `timedelta` | Wall-clock duration | | `turn_count` | `int` | SDK turn count, currently `1` | | `session_id` | `str` | Session that ran the turn | | `output` | `T \| None` | Locally adapted structured output | | `structured_output` | `FrozenJsonObject \| None` | Immutable raw structured output | | `output_validation_error` | `ValidationError \| None` | Pydantic failure | | `structured_output_error` | `StructuredOutputError \| None` | Droid failure | | `error` | `ErrorEvent \| None` | Terminal execution error | All result variants retain partial output. `RunFailure` also exposes `error` and the server's `structured_output_error` when available. A successful turn may have no structured output. `RunResult` does not define truthiness. Check `success` or `subtype`. ### Token and context usage ```python if result.usage: print(result.usage.input_tokens) print(result.usage.output_tokens) print(result.usage.cache_read_tokens) print(result.usage.cache_creation_tokens) print(result.usage.thinking_tokens) print(result.usage.factory_credits) ``` `TokenUsageUpdate` contains cumulative committed session usage. Context occupancy is separate: ```python context = await session.context() print(context.used, context.remaining, context.limit, context.accuracy) ``` `ContextUsage.updated_at` records when Droid measured the value. | Type | Fields | | --- | --- | | `Usage` | `input_tokens`, `output_tokens`, `cache_creation_tokens`, `cache_read_tokens`, `thinking_tokens`, `factory_credits` | | `ContextUsage` | `used`, `remaining`, `limit`, `accuracy`, `updated_at` | ### Errors ```python async with session.stream("Run the tests.") as stream: async for _ in stream: pass result = stream.result if result.subtype == "interrupted": print("The turn was interrupted.") elif result.subtype in ("error_during_execution", "error_structured_output"): print(result.error.message if result.error else result.subtype) ``` Setup, connection, process, protocol, and timeout failures raise `DroidError` subclasses. Python programming errors use normal built-in exceptions. `asyncio.CancelledError` is never wrapped. ### Stream concurrency A session may have one active stream. Starting another raises `SessionBusyError`. Different sessions may run concurrently. ### Timeout ```python async with session.stream("Perform a long review.", timeout=60) as stream: async for event in stream: handle(event) ``` On expiry, the SDK interrupts the turn and raises `RunTimeoutError`. ### Interrupt a turn Another task may stop active work: ```python await session.interrupt() ``` The active stream yields an interrupted result if its consumer remains attached. The session stays open. ### Cancellation and interruption Cancelling the task running a turn sends a best-effort interrupt, releases the subscription, then re-raises `asyncio.CancelledError`. Use the stream context manager when iteration may stop early: ```python async with session.stream("Inspect every test.") as stream: async for event in stream: if should_stop(event): break ``` Context exit interrupts unfinished work. A bare async iterator cannot guarantee immediate cleanup on `break`. `await stream.aclose()` explicitly interrupts and detaches an unfinished stream; it is idempotent. ## Models Model IDs are strings because availability depends on account and organization policy. Omit `model` to use the configured default, which lives in `~/.factory/settings.json` under `sessionDefaultSettings`. An unknown model ID is rejected by the backend: the turn returns `RunFailure(subtype="error_during_execution")` with the rejection message in `error.message`. ### Discover available models Use `list_models()` to inspect the models available to the current account and project before opening a session. ```python from droid_sdk import list_models models = await list_models() for model in models: print(model.id, model.default_reasoning_effort) ``` `list_models()` starts a one-shot Droid process, requests the model catalog without creating a session, and closes the process before returning. Disabled models are hidden by default. Include them when building a model picker or explaining why a model cannot be selected: ```python from pathlib import Path from droid_sdk import list_models models = await list_models( include_disabled=True, cwd=Path.cwd(), ) for model in models: status = ( f"disabled: {model.disabled_reason}" if model.disabled else "available" ) print(model.display_name, status) ``` The catalog includes built-in and configured custom models after applying feature flags, region availability, and organization policy. Passing `cwd` applies project-specific settings. Custom endpoint and credential configuration is never returned. The keyword-only arguments are: | Argument | Type | Purpose | | --- | --- | --- | | `include_disabled` | `bool` | Include disabled models and their `disabled_reason` | | `cwd` | `str \| Path \| None` | Apply project settings from this directory | | `runtime` | `Runtime \| None` | Configure the Droid process and observability | | `api_key` | `str \| None` | Override `FACTORY_API_KEY` for this process | Each immutable `ModelInfo` reports the model ID, display name, provider, supported and default reasoning efforts, custom-model status, disabled status, and optional capabilities. If the installed Droid version does not support model discovery, `list_models()` raises `DroidProtocolError` with instructions to update Droid. ### Select a model ```python from droid_sdk import ReasoningEffort, run result = await run( "Review this repository.", model="model-id", reasoning_effort=ReasoningEffort.HIGH, ) ``` Omit `reasoning_effort` to use the model default. Change the model for later turns: ```python await session.update_settings( model="model-id", reasoning_effort=ReasoningEffort.HIGH, ) ``` See [Update session settings](#update-session-settings) for the full `update_settings()` contract. ### Use the Factory Router The model ID `auto` selects the [Factory Router](/model-independence/factory-router), which routes each prompt to the model with the best balance of quality, latency, and cost. Some product surfaces label it Auto Model; the model ID is `auto` everywhere. ```python from droid_sdk import run result = await run("Review this repository.", model="auto") ``` Pin a session to the router the same way: ```python async with Session(model="auto") as session: ... ``` Move a live session onto the router: ```python await session.update_settings(model="auto") ``` Omit `reasoning_effort`; the router chooses the effort along with the model and ignores a supplied value. `session.settings.model` reports `auto`; the underlying model can differ per response. The model ID `auto` is unrelated to `Mode.AUTO`, the default interaction mode, and to `Autonomy`, the permission level. #### See which model handled a response Assistant messages in the wire `create_message` notification carry the underlying model in `modelId`; `routerId` is `"auto"` when the router made the choice. High-level messages omit these fields. Subscribe with [`on_notification()`](#subscribe-to-raw-notifications): ```python from collections.abc import Mapping def report_routing(notification: Mapping[str, object]) -> None: message = notification.get("message") if isinstance(message, Mapping) and message.get("role") == "assistant": print(message.get("modelId"), message.get("routerId")) unsubscribe = session.on_notification(report_routing, type="create_message") ``` Wire payloads use camelCase keys and evolve server-side; treat absent keys as normal. ### Configure mode-specific models ```python from droid_sdk import Mode, ReasoningEffort, Session, SessionConfig config = SessionConfig( mode=Mode.SPEC, spec_model="model-id", spec_reasoning_effort=ReasoningEffort.HIGH, ) async with Session(model="model-id", config=config) as session: async with session.stream("Draft an implementation plan.") as stream: async for _ in stream: pass ``` The primary model handles Auto turns. `spec_model` handles Spec turns. Switch modes on a live session with `enter_spec()` and `leave_spec()`; see [Spec mode](#spec-mode). ### Model configuration contract | Field | Type | Used by | | --- | --- | --- | | `model` | `str \| None` | `run()`, `Session`, `update_settings()` | | `reasoning_effort` | `ReasoningEffort \| None` | Same | | `spec_model` | `str \| None` | `SessionConfig`, `update_settings()` | | `spec_reasoning_effort` | `ReasoningEffort \| None` | Same | `None` uses the configured default. Setting a Spec field to `None` through `update_settings()` clears it. ## Inputs and outputs ### Images and documents Attach images and files to a turn with the `images` and `files` options: ```python from droid_sdk import Document, Image, run result = await run( "Compare these files.", images=[ Image.from_path("screenshot.png"), Image.from_bytes(image_bytes, media_type="image/png"), ], files=[ Document.from_path("report.pdf"), Document.from_text(source, name="auth.py"), ], ) ``` The constructors read local data and encode it for the turn. Supported image types are PNG, JPEG, GIF, and WebP; image URLs are unsupported. Invalid local input raises `InvalidAttachmentError` before the turn starts. Attachments are limited to `MAX_ATTACHMENT_BYTES` (5 MiB), and PDFs to `MAX_PDF_ATTACHMENT_BYTES` (3 MiB). ### Input schemas | Type | Fields | | --- | --- | | `Image` | `source: Base64ImageSource` | | `Base64ImageSource` | `data`, `media_type` | | `Document` | `source: TextDocumentSource \| PdfDocumentSource` | | `TextDocumentSource` | `data`, `name`, `mime` | | `PdfDocumentSource` | `data`, `parsed_data`, `name`, `path` | | Constructor | Returns | | --- | --- | | `Image.from_path(path)` | `Image` | | `Image.from_bytes(data, media_type=...)` | `Image` | | `Document.from_path(path)` | `Document` | | `Document.from_text(text, name=..., mime=...)` | `Document` | | `Document.from_bytes(data, name=...)` | `Document` | `images` accepts `Sequence[Image]`. `files` accepts `Sequence[Document]`. Wire payloads use canonical `mediaType` fields. The optional `mime` hint on text documents is forwarded with the document. ### Structured output #### Return a Pydantic model ```python from typing import Literal from pydantic import BaseModel from droid_sdk import RunSuccess, run class Finding(BaseModel): severity: Literal["low", "medium", "high"] message: str class Review(BaseModel): summary: str findings: list[Finding] result = await run( "Review the authentication code.", output=Review, ) if isinstance(result, RunSuccess) and result.output is not None: print(result.output.summary) # success guarantees output when requested elif result.output_validation_error is not None: print(result.output_validation_error) ``` `output` accepts a `BaseModel` subclass, not an arbitrary class. Unsupported types raise `TypeError` before the turn starts. With `output=Review`, the return type is `RunResult[Review]`. When structured output arrives, the SDK validates it against the model. If output was requested, `RunSuccess` guarantees `result.output` is set. Missing or invalid output turns the result into `RunFailure` with subtype `error_structured_output`; the failure keeps the text, messages, usage, raw `structured_output`, and `output_validation_error` for inspection. #### Use raw JSON Schema ```python from droid_sdk import JsonObject, JsonSchema, RunResult, run schema = JsonSchema( { "type": "object", "properties": {"summary": {"type": "string"}}, "required": ["summary"], } ) result: RunResult[JsonObject] = await run( "Summarize the repository.", output=schema, ) if result.output is not None: print(result.output["summary"]) ``` `JsonSchema` accepts an object-shaped schema. Droid reports unsupported or invalid schemas through the normal result subtype. ### Output contract | `output` argument | Return type | | --- | --- | | Omitted or `None` | `RunResult[None]` | | `type[BaseModel]` | `RunResult[Model]` | | `JsonSchema` | `RunResult[JsonObject]` | `JsonSchema.schema` is `FrozenJsonObject`. Schema mappings must contain only JSON-compatible values; the SDK validates this and freezes them recursively. Raw schema output remains `JsonObject`. ```python JsonValue = bool | int | float | str | list["JsonValue"] | dict[str, "JsonValue"] | None JsonObject = dict[str, JsonValue] FrozenJsonValue = ( bool | int | float | str | tuple["FrozenJsonValue", ...] | Mapping[str, "FrozenJsonValue"] | None ) FrozenJsonObject = Mapping[str, FrozenJsonValue] ``` Both local and Droid-side output failures produce `RunFailure` with subtype `error_structured_output`. To tell them apart, check `structured_output_error.code`: local failures use `local_validation_failed` (with the Pydantic error in `output_validation_error`) or `local_output_missing`; Droid-side failures use Droid's own codes. ## Permissions and user input ### Autonomy Autonomy controls which actions require approval in Auto mode: | Level | Behavior | | --- | --- | | `Autonomy.OFF` | Ask before every action | | `Autonomy.LOW` | Allow edits and read-only commands | | `Autonomy.MEDIUM` | Allow reversible commands | | `Autonomy.HIGH` | Allow commands without approval | Configure it through `SessionConfig`: ```python config = SessionConfig(autonomy=Autonomy.LOW) ``` Autonomy does not select Auto or Spec mode. `Mode` controls that. ### Permission handler ```python from droid_sdk import ( InteractionHandlers, PermissionRequest, PermissionResponse, Session, ToolConfirmationOutcome, ) from droid_sdk.permissions import CreateFile def approve(request: PermissionRequest) -> PermissionResponse: if request.actions and all( isinstance(action, CreateFile) for action in request.actions ): return request.respond(ToolConfirmationOutcome.PROCEED_ONCE) return request.respond(ToolConfirmationOutcome.CANCEL) session = Session( interactions=InteractionHandlers(on_permission=approve), ) ``` Droid decides which outcomes are on offer. `respond()` accepts only an outcome listed in `request.options`, plus optional `comment` or `edited_spec_content`. An absent handler, an invalid response, or a handler exception cancels the request. Handler failures appear as `ErrorEvent` values; they do not raise through the active stream. ### AskUser handler ```python from droid_sdk import QuestionRequest, QuestionResponse def answer(request: QuestionRequest) -> QuestionResponse: answers = [ question.answer(question.options[0] if question.options else "none") for question in request.questions ] return request.submit(answers) ``` Cancel the questionnaire: ```python return request.cancel() ``` Answers are always strings on the wire. Without a handler, or when a handler fails or returns an invalid shape, the questionnaire is cancelled and the stream receives an `ErrorEvent`. `question.answer(value)` answers one value. `question.answer_multiple(values)` joins multi-select values with `", "` to produce the wire-compatible string. Handlers run inside the active turn. Use async handlers for I/O, and do not start another turn on the same session from a handler. ### Interaction contracts ```python PermissionHandler = Callable[ [PermissionRequest], PermissionResponse | Awaitable[PermissionResponse], ] QuestionHandler = Callable[ [QuestionRequest], QuestionResponse | Awaitable[QuestionResponse], ] ``` | Type | Fields | | --- | --- | | `InteractionHandlers` | `on_permission`, `on_question` | | `PermissionRequest` | `actions`, `options`, `associated_session_ids`, `plan` | | `PermissionOption` | `label`, `value` | | `Plan` | `text`, `title` | | `PermissionResponse` | `selected_option`, `comment`, `edited_spec_content` | | `QuestionRequest` | `tool_call_id`, `questions` | | `Question` | `index`, `topic`, `question`, `options`, `multi_select` | | `QuestionAnswer` | `index`, `question`, `answer` | | `QuestionResponse` | `cancelled`, `answers` | `PermissionAction` is a discriminated union: ```python PermissionAction = ( EditAction | ExecuteAction | CreateFile | AskUserAction | ExitSpecModeAction | ApplyPatchAction | McpToolAction | SandboxViolationAction | DroidShieldViolationAction ) ``` Every action includes its `tool_use`, confirmation type, and typed details: | Type | Detail fields | | --- | --- | | `EditAction` | `file_path`, `file_name`, `old_content`, `new_content` | | `ExecuteAction` | `full_command`, `command`, `extracted_commands`, `impact_level`, `risk_level_reason` | | `CreateFile` | `file_path`, `file_name`, `content` | | `AskUserAction` | `questionnaire`, `questions`, `parse_error` | | `ExitSpecModeAction` | `plan`, `title` | | `ApplyPatchAction` | `file_path`, `file_name`, `patch_content`, `old_content`, `new_content`, `files` | | `ApplyPatchFile` | `file_path`, `file_name`, `operation`, `move_to`, `old_content`, `new_content` | | `McpToolAction` | `tool_name`, `server_name`, `actual_tool_name`, `impact_level` | | `SandboxViolationAction` | `violating_tool_name`, `target`, `operation`, `violation_type`, `reason`, `violation_reason`, `is_org_deny` | | `DroidShieldViolationAction` | `command`, `reason` | `AskUserParseError(message: str, line: int | None = None)` records malformed questionnaire details. `ApplyPatchFile` requires `file_path`, `file_name`, and `operation` (`"create"`, `"update"`, or `"delete"`); `move_to`, `old_content`, and `new_content` default to `None`. `ToolConfirmationOutcome` defines: | Outcome | Meaning | | --- | --- | | `PROCEED_ONCE` | Approve this request | | `PROCEED_ALWAYS` | Persist the offered rule | | `PROCEED_ALWAYS_EXACT_PATH` | Persist the exact file path | | `PROCEED_AUTO_RUN` | Continue with automatic approvals | | `PROCEED_AUTO_RUN_LOW` | Continue at low autonomy | | `PROCEED_AUTO_RUN_MEDIUM` | Continue at medium autonomy | | `PROCEED_AUTO_RUN_HIGH` | Continue at high autonomy | | `PROCEED_NEW_SESSION` | Continue in a new session | | `PROCEED_NEW_SESSION_LOW` | New session at low autonomy | | `PROCEED_NEW_SESSION_MEDIUM` | New session at medium autonomy | | `PROCEED_NEW_SESSION_HIGH` | New session at high autonomy | | `PROCEED_EDIT` | Submit edited plan content | | `PROCEED_ALWAYS_TOOLS` | Persist approval for MCP tools | | `PROCEED_ALWAYS_SERVER` | Persist approval for an MCP server | | `CANCEL` | Reject the request | Only outcomes present in `PermissionRequest.options` are valid. ### Tool controls ```python config = SessionConfig( additional_tools={"CustomTool"}, enabled_tools={"Read"}, disabled_tools={"Execute", "Edit"}, restrict_tools={"Read", "Grep"}, ) ``` Each of the four sets has a distinct role: - `additional_tools` adds IDs to the catalog. - `enabled_tools` enables otherwise available tools. - `disabled_tools` removes tools. - `restrict_tools` is a restrictive allowlist and never elevates permission. ```python for tool in await session.list_tools(model="model-id", mode=Mode.AUTO): print(tool.id, tool.allowed) ``` `update_settings()` accepts the same override fields for later turns. The four override parameters accept any iterable of tool IDs, such as a set, list, or tuple. Passing a bare string raises `TypeError`. `list_tools()` returns `list[ToolInfo]`. | Type | Fields | | --- | --- | | `ToolInfo` | `id`, `display_name`, `description`, `category`, `default_allowed`, `allowed` | | `ListToolsOptions` | `model`, `mode`, `autonomy`, `spec_model`, `additional_tools`, `enabled_tools`, `disabled_tools`, `restrict_tools`, `skip_permissions_unsafe` | ## Extensions | Extension | Use it for | | --- | --- | | Skills | Reusable instructions and supporting files | | In-process MCP tools | Python functions exposed to Droid | | External MCP servers | Tools from another process or service | | Hooks | Commands run at Droid lifecycle points | ### Skills ```python skills = await session.list_skills() for skill in skills.skills: print(skill.name, skill.enabled) ``` `skills.project_available` reports whether a project skill scope is available; it is `None` when Droid does not report it. Enable or disable a skill: ```python await session.enable_skill("review", scope="project") await session.disable_skill("legacy", scope="user") ``` Skills may come from project, personal, built-in, or automation settings. #### Skill schemas | Type | Fields | | --- | --- | | `SkillsResult` | `skills`, `project_available` | | `SkillInfo` | `name`, `description`, `location`, `file_path`, `enabled`, `user_invocable`, `version`, `content`, `resources`, `disabled_by` | | `SkillResource` | `name`, `path`, `type` | | `SkillMutationResult` | `success` | `scope` is `"user"` or `"project"`. `SkillResource.type` is `"reference"` or `"asset"`. `SkillInfo.location` is `"project"`, `"personal"`, `"builtin"`, or `"automation"`. ### In-process MCP tools In-process servers require the `mcp` extra (`pip install "droid-sdk[mcp]"`). Use `@tool` to expose an annotated Python function: ```python from droid_sdk.mcp import create_sdk_mcp_server, tool @tool("lookup_owner", "Return the owner of a repository file.") async def lookup_owner(path: str) -> str: return f"Owner for {path}: platform-team" server = create_sdk_mcp_server( name="review-tools", tools=[lookup_owner], version="1.0.0", ) config = SessionConfig(mcp_servers=[server]) ``` The decorator derives JSON Schema from type annotations. Invalid input is returned to Droid as a tool error and is not passed to the function. Tool functions may be synchronous or asynchronous. They may return text or a typed `ToolResponse`. The SDK starts an authenticated loopback server and closes it with the session. Attach in-process servers again when resuming. #### In-process MCP schemas | Type | Fields | | --- | --- | | `DroidTool` | `name`, `description`, `input_schema`, `handler`, `output_schema` | | `SdkMcpServer` | `name`, `version`, `tools` | | `ToolResponse` | `content`, `is_error`, `structured_content` | `tool(name, description)` is a decorator; `tool(name, description, function)` is the equivalent direct call. A return annotation that describes an object is validated at call time, returned as structured content, and advertised through the MCP `outputSchema` field. `create_sdk_mcp_server(name, tools, version="1.0.0")` returns `SdkMcpServer`. `SdkMcpServer.config` is the active `HttpMcpServerConfig` or `None`; `await server.start()` returns that config, and `await server.close()` stops the server. Sessions normally own these calls. ### External MCP servers External server configs do not need the `mcp` extra: ```python from droid_sdk import HttpMcpServerConfig, StdioMcpServerConfig config = SessionConfig( mcp_servers=[ HttpMcpServerConfig(name="docs", url="https://example.com/mcp"), StdioMcpServerConfig( name="search", command="python", args=["-m", "search_server"], ), ], ) ``` HTTP, SSE, and stdio transports are supported. HTTP and SSE configurations support headers and OAuth settings. Inspect connected servers and tools: ```python servers = await session.list_mcp_servers() tools = await session.list_mcp_tools() print(servers.summary) for server in servers.servers: print(server.name, server.status) ``` Session methods can add, remove, enable, disable, and authenticate external servers. Configuration mutations affect the user's Droid settings. Servers passed through `SessionConfig` are session-scoped. | Session MCP method | Signature | | --- | --- | | `list_mcp_servers` | `() -> McpServersResult` | | `list_mcp_tools` | `() -> list[McpToolInfo]` | | `add_mcp_server` | `(config) -> McpMutationResult` | | `remove_mcp_server` | `(name: str) -> McpMutationResult` | | `enable_mcp_server` / `disable_mcp_server` | `(name: str) -> McpMutationResult` | | `enable_mcp_tool` / `disable_mcp_tool` | `(server_name: str, tool_name: str) -> McpMutationResult` | | `authenticate_mcp_server` | `(name: str) -> McpMutationResult` | | `cancel_mcp_auth` / `clear_mcp_auth` | `(name: str) -> McpMutationResult` | | `submit_mcp_auth_code` | `(name: str, *, code: str, state: str) -> McpMutationResult` | | `submit_mcp_auth_error` | `(name: str, *, error: str, state: str, error_description: str \| None = None) -> McpMutationResult` | #### External MCP schemas ```python McpServerConfig = ( StdioMcpServerConfig | HttpMcpServerConfig | SseMcpServerConfig | SdkMcpServer ) ``` | Type | Fields | | --- | --- | | `StdioMcpServerConfig` | `name`, `command`, `args`, `env` | | `HttpMcpServerConfig` | `name`, `url`, `headers`, `oauth` | | `SseMcpServerConfig` | `name`, `url`, `headers`, `oauth` | | `HttpHeader` | `name`, `value` | | `McpOAuthOptions` | `scopes`, `resource`, `authorization_server_issuer`, `client_metadata_url`, `client_id`, `client_secret`, `callback_port`, `token_endpoint_auth_method` | | `McpServersResult` | `servers`, `summary` | | `McpServerStatusInfo` | `name`, `status`, `source`, `is_managed`, `error`, `tool_count`, `server_type`, `has_auth_tokens`, `requires_auth`, `pending_auth_url`, `pending_auth_message`, `pending_auth_state` | | `McpStatusSummary` | `total`, `connected`, `connecting`, `failed`, `disabled`, `config_error` | | `McpConfigError` | `path`, `message` | | `McpToolInfo` | `server_name`, `name`, `description`, `is_enabled`, `is_read_only`, `input_schema` | | `McpToolInputSchema` | `type`, `properties`, `required` | | `McpMutationResult` | `success` | `oauth` accepts `McpOAuthOptions` or `False`. Remove, enable, and disable mutate user-scoped MCP configuration only. These methods do not accept a project-scope argument. ### Hooks There is no Python API for defining hooks; configure them in `.factory/hooks.json`. Hook execution appears as `HookExecution` values in the run stream. `HookExecution.status` is `"started"`, `"completed"`, or `"error"`. ## Session lifecycle Fork, compact, and rewind create successor sessions. The successor is an opened `Session` that takes over the existing Droid connection; `fork()` returns it directly, while `compact()` and `rewind()` return it on their outcome objects. ### Fork Fork copies the conversation into a new session and continues there. ```python fork = await session.fork( title="Alternative approach", tags=[SessionTag(name="experiment")], ) async with fork: async with fork.stream("Try the other strategy.") as stream: async for _ in stream: pass ``` ### Compact Compaction summarizes older conversation history to free context-window space. ```python outcome = await session.compact(instructions="Keep decisions and unresolved failures.") async with outcome.session as compacted: print(outcome.removed_count) ``` ### Rewind Rewind returns the conversation to an earlier message and can restore or delete files changed since. ```python info = await session.rewind_info(message_id) outcome = await session.rewind( message_id, restore=info.available_files, delete=info.created_files, title="Before the failed change", ) async with outcome.session as rewound: print(outcome.restored_count, outcome.deleted_count) print(outcome.failed_restore_count, outcome.failed_delete_count) ``` `info.evicted_files` explains files that cannot be restored. ### Lifecycle contracts | Method | Arguments | Returns | | --- | --- | --- | | `fork()` | `title`, `tags` | Opened `Session` | | `compact()` | `instructions` | `CompactOutcome` | | `rewind_info()` | `message_id` | `RewindInfo` | | `rewind()` | `message_id`, `restore`, `delete`, `title` | `RewindOutcome` | | Type | Fields | | --- | --- | | `CompactOutcome` | `session`, `removed_count` | | `RewindInfo` | `available_files`, `created_files`, `evicted_files` | | `RewindFileSnapshot` | `file_path`, `content_hash`, `size` | | `RewindFileCreation` | `file_path` | | `RewindEvictedFile` | `file_path`, `reason` | | `RewindOutcome` | `session`, `restored_count`, `deleted_count`, `failed_restore_count`, `failed_delete_count` | ### Successor ownership After a successful replacement: - the returned successor owns the connection and runtime resources - the source session becomes retired - `id`, `cwd`, and `settings` remain readable on the source - active methods on the source raise `SessionReplacedError` - closing the source is a no-op Replacing a session with an active turn raises `SessionBusyError`. Only one replacement may run at a time. `open()` and another replacement raise `SessionBusyError` while replacement is active. `close()` racing a replacement waits; on success it closes the successor, and on rollback it closes the source. Cancelling a replacement before handoff restores the source to open state. If cancellation arrives after a successor is created, the SDK reloads the source with its attached policies and retires the detached successor. If source restoration fails, the SDK closes the connection rather than leaving an ambiguous owner. ## Spec mode Spec mode lets Droid inspect a codebase and propose a plan without changing files. ### Enter Spec Mode ```python await session.enter_spec( model="model-id", reasoning_effort=ReasoningEffort.HIGH, ) ``` Start directly in Spec mode with `SessionConfig(mode=Mode.SPEC)`. ### Leave without approving a plan ```python await session.leave_spec() ``` This changes the mode only. It does not approve the plan or start implementation. ### Approve a plan Plan approval arrives through the permission handler: ```python def approve(request: PermissionRequest) -> PermissionResponse: if request.plan: print(request.plan.text) return request.respond(ToolConfirmationOutcome.PROCEED_ONCE) return request.respond(ToolConfirmationOutcome.CANCEL) ``` Return any offered `PROCEED_NEW_SESSION*` outcome to hand implementation to a new session. Return `PROCEED_EDIT` with `edited_spec_content` to revise the plan. `CANCEL` ends the turn with an interrupted result. ### Mode contract | API | Arguments | Returns | | --- | --- | --- | | `SessionConfig(mode=...)` | `Mode.AUTO` or `Mode.SPEC` | `SessionConfig` | | `enter_spec()` | `model`, `reasoning_effort` | `UpdateSettingsResult` | | `leave_spec()` | None | `UpdateSettingsResult` | Entering or leaving Spec mode only changes settings. Plan approval is an `ExitSpecModeAction` permission request, and only an offered `ToolConfirmationOutcome` is valid. ## Observability and advanced APIs ### Observability ```python from droid_sdk import Runtime, run from droid_sdk.observability import LogEvent, Observability class PrintLogger: def log(self, event: LogEvent) -> None: print(event.level, event.name, event.message) runtime = Runtime(observability=Observability(logger=PrintLogger())) result = await run("Check repository status.", runtime=runtime) ``` Logger, metric, and trace-context sinks are synchronous and best-effort. Sink failures never fail a Droid operation. Events exclude prompts, messages, thinking, tool inputs and results, file contents, raw process output, stack traces, and credentials. #### Observability schemas | Type | Fields or methods | | --- | --- | | `Observability` | `logger`, `metrics`, `tracing` | | `Logger` | `log(event: LogEvent) -> None` | | `LogEvent` | `level`, `name`, `message`, `attributes`, `error` | | `SerializedError` | `name`, `message`, `code` | | `MetricSink` | `record(event: MetricEvent) -> None` | | `MetricEvent` | `name`, `kind`, `value`, `unit`, `attributes` | | `TraceContextProvider` | `inject(carrier: TraceContext) -> None` | | `TraceContext` | `traceparent`, `tracestate` | `attributes` is `Mapping[str, str | int | float | bool | None]`. Log levels are `"debug"`, `"info"`, `"warn"`, and `"error"`. Metric kinds are `"counter"` and `"histogram"`. `TraceContext` is deliberately mutable so tracing providers can inject values into it. ### Custom runtime `Runtime` holds process configuration: ```python runtime = Runtime( executable=Path("/opt/factory/bin/droid"), args=["--flag"], env={"EXAMPLE": "value"}, ) ``` Environment entries extend the current process environment when the SDK starts Droid. #### Runtime schema | Field | Type | | --- | --- | | `executable` | `str \| Path \| None` | | `args` | `Sequence[str]` | | `env` | `Mapping[str, str]` | | `observability` | `Observability \| None` | ## Examples Run commands from the repository root: | Example | Command | Expected result | | --- | --- | --- | | Attachments | `uv run python examples/attachments.py` | Live image, text, and PDF turn | | Custom system prompt | `uv run python examples/custom_system_prompt.py` | Live turn with instructions appended to Droid's prompt | | Factory Router | `uv run python examples/factory_router.py` | Live routed turns with per-response model IDs | | Interaction helpers | `uv run python examples/interaction_helpers.py` | Offline typed responses | | Interactions | `uv run python examples/interactions.py` | Live permission/question turn | | Interactive session | `uv run python examples/interactive_session.py` | Two live turns sharing history | | Model discovery | `uv run python examples/list_models.py` | Available and disabled model metadata | | Saved sessions | `uv run python examples/list_saved_sessions.py` | Local saved-session count | | Observability | `uv run python examples/observability.py` | Offline isolated sink counts | | One-shot run | `uv run python examples/one_shot.py` | Live one-turn result | | Resume | `uv run python examples/resume_session.py --session-id ID` | Live resumed turn | | SDK MCP | `uv run python examples/sdk_mcp.py` | Live in-process MCP result | | Session operations | `uv run python examples/session_operations.py` | Live settings, discovery, fork, and compact | | Stream events | `uv run python examples/stream_events.py` | Live tour of every stream event type | | Structured output | `uv run python examples/structured_output_model.py` | Live validated model output | Live examples require an authenticated local Droid CLI or `FACTORY_API_KEY`; every model call uses a finite timeout. ## API reference ### Root constants and supporting types | Export | Contract | | --- | --- | | `__version__` | Installed package version as `str` | | `MAX_ATTACHMENT_BYTES` | General attachment limit, `5 * 1024 * 1024` | | `MAX_PDF_ATTACHMENT_BYTES` | PDF limit, `3 * 1024 * 1024` | | `ImageMediaType` | Supported image MIME literal union | | `JsonValue`, `JsonObject` | Mutable JSON input/output aliases | | `FrozenJsonValue`, `FrozenJsonObject` | Recursively immutable JSON aliases | | `ApplyPatchFile` | Per-file patch details documented above | | `AskUserParseError` | `message`, `line` | | `ModelInfo` | Immutable model-discovery metadata | | `SandboxSettings` | `enabled`, `mode` | | `SessionSettingsUpdate` | Partial settings notification documented above | | `SystemPromptConfig` | `str \| SystemPromptPreset` | | `SystemPromptPreset` | Typed dictionary with required `type="preset"`, `preset="droid"`, and non-empty `append` | ### Top-level functions | API | Returns | Purpose | | --- | --- | --- | | `run(prompt, **options)` | `RunResult[T]` | Run one turn | | `list_models(**options)` | `list[ModelInfo]` | List selectable models | | `list_sessions(**filters)` | `list[SavedSession]` | List saved sessions | | `run()` argument | Type | | --- | --- | | `prompt` | `str` | | `cwd` | `str \| Path \| None` | | `model` | `str \| None` | | `reasoning_effort` | `ReasoningEffort \| None` | | `images` | `Sequence[Image]` | | `files` | `Sequence[Document]` | | `output` | `type[BaseModel] \| JsonSchema \| None` | | `timeout` | `float \| None` | | `config` | `SessionConfig \| None` | | `interactions` | `InteractionHandlers \| None` | | `runtime` | `Runtime \| None` | | `api_key` | `str \| None` | ### `Session` | Member | Returns | | --- | --- | | `Session(...)` | Lazy `Session` | | `Session.resume(id, ...)` | Lazy `Session` | | `open()` | `None` | | `close()` | `None` | | `id` | `str` | | `cwd` | `Path \| None` | | `settings` | `SessionSettings` | | `stream()` | `RunStream[T, E]` | | `interrupt()` | `None` | | `update_settings()` | `UpdateSettingsResult` | | `rename()` | `None` | | `on_notification()` | `Callable[[], None]` | | `list_tools()` | `list[ToolInfo]` | | `list_skills()` | `SkillsResult` | | `enable_skill()` / `disable_skill()` | `SkillMutationResult` | | `list_mcp_servers()` | `McpServersResult` | | `list_mcp_tools()` | `list[McpToolInfo]` | | MCP mutation methods | `McpMutationResult` | | `context()` | `ContextUsage` | | `fork()` | Opened `Session` | | `compact()` | `CompactOutcome` | | `rewind_info()` | `RewindInfo` | | `rewind()` | `RewindOutcome` | | `enter_spec()` / `leave_spec()` | `UpdateSettingsResult` | ### `RunStream` | Member | Returns | | --- | --- | | Async iteration | `StreamMessage[T]` or `StreamEvent[T]` values | | `result` | Cached `RunResult[T]` | | `aclose()` | `None` | ### Main enums | Enum | Values | | --- | --- | | `Mode` | `AUTO`, `SPEC` | | `Autonomy` | `OFF`, `LOW`, `MEDIUM`, `HIGH` | | `ModelProvider` | `ANTHROPIC`, `OPENAI`, `GENERIC_CHAT_COMPLETION_API`, `FACTORY`, `GOOGLE`, `XAI`, `VOYAGE` | | `ReasoningEffort` | See supported values below | | `ToolCategory` | `READ`, `EDIT`, `EXECUTE`, `OTHER` | | `ToolConfirmationType` | `EDIT`, `EXECUTE`, `CREATE`, `ASK_USER`, `EXIT_SPEC_MODE`, `APPLY_PATCH`, `MCP_TOOL`, `SANDBOX_VIOLATION`, `DROID_SHIELD_VIOLATION` | | `ToolConfirmationOutcome` | Permission outcomes offered by Droid | | `WorkingState` | `IDLE`, `THINKING`, `STREAMING_ASSISTANT_MESSAGE`, `WAITING_FOR_TOOL_CONFIRMATION`, `EXECUTING_TOOL`, `COMPACTING_CONVERSATION` | | `ContextAccuracy` | `EXACT`, `ESTIMATED` | | `McpServerType` | `STDIO`, `HTTP`, `SSE` | | `McpServerStatus` | `CONNECTING`, `CONNECTED`, `DISCONNECTED`, `FAILED`, `DISABLED` | | `McpAuthOutcome` | `SUCCESS`, `CANCELLED`, `FAILED` | | `OAuthTokenEndpointAuthMethod` | `NONE`, `CLIENT_SECRET_BASIC`, `CLIENT_SECRET_POST` | | `SessionPlatform` | `SLACK`, `WEB`, `API`, `SESSIONS_API`, `JIRA`, `LINEAR`, `MICROSOFT_TEAMS`, `READINESS_REMEDIATION`, `READINESS_EVALUATION`, `AUTOMATION`, `WIKI_GENERATION`, `WIKI_CI_SETUP`, `TUI`, `DESKTOP`, `ACP`, `UNKNOWN` | | `SandboxOperation` | `READ`, `WRITE`, `NETWORK`, `TOOL` | | `SandboxViolationType` | `FILESYSTEM_READ`, `FILESYSTEM_WRITE`, `NETWORK`, `TOOL` | | `SandboxViolationReason` | `DENY_LIST`, `NOT_ALLOWED` | | `ErrorType` | `CONNECTION_ERROR`, `PROTOCOL_ERROR`, `SESSION_ERROR`, `TIMEOUT_ERROR`, `DROID_CLIENT_ERROR`, `PROCESS_EXIT_ERROR`, `ERROR` | `ReasoningEffort` defines `NONE`, `DYNAMIC`, `OFF`, `MINIMAL`, `LOW`, `MEDIUM`, `HIGH`, `EXTRA_HIGH`, and `MAX`. Model and tool IDs are plain strings, not enums. On the wire, these enums use the following values: ```text ToolCategory: read, edit, execute, other ToolConfirmationType: edit, exec, create, ask_user, exit_spec_mode, apply_patch, mcp_tool, sandbox_violation, droid_shield_violation OAuthTokenEndpointAuthMethod: none, client_secret_basic, client_secret_post SandboxOperation: read, write, network, tool SandboxViolationType: filesystem-read, filesystem-write, network, tool SandboxViolationReason: deny-list, not-allowed SessionPlatform: slack, web, api, sessions_api, jira, linear, microsoft-teams, readiness-remediation, readiness-evaluation, automation, wiki-generation, wiki-ci-setup, tui, desktop, acp, unknown ``` ### Exceptions All SDK-defined exceptions derive from `DroidError`. Python built-ins and `asyncio.CancelledError` do not. | Exception | Meaning | | --- | --- | | `RunTimeoutError` | The turn exceeded its deadline | | `StreamIncompleteError` | A stream result was read before completion | | `InvalidAttachmentError` | A local attachment was invalid | | `SessionNotOpenError` | An operation required an opened session | | `SessionBusyError` | The session already had an active turn or replacement in progress | | `SessionClosedError` | The session was closed | | `SessionReplacedError` | A successor retired the source session | | `SessionReplacementError` | Successor load or source restore failed | | `SessionNotFoundError` | A saved session ID was not found | | `InvalidWorkingDirectoryError` | A working directory is unavailable | | `DroidConnectionError` | The local connection failed | | `DroidProcessError` | The Droid process exited unexpectedly | | `DroidProtocolError` | Protocol negotiation or validation failed | Exception constructor metadata is public: | Exception | Additional constructor fields | | --- | --- | | `RunTimeoutError(message, ...)` | `request_id`, `method`, `timeout_duration` | | `SessionReplacedError(session_id, replacement_session_id)` | both session IDs | | `SessionReplacementError(session_id, replacement_session_id, ...)` | both IDs, `rollback_error`, `rollback_failed` | | `SessionNotFoundError(session_id)` | `session_id` | | `InvalidWorkingDirectoryError(cwd, message=None)` | `cwd` | | `DroidConnectionError(message, ...)` | `cwd`, `exec_path` | | `DroidProcessError(message, ...)` | `exit_code`, `signal` | | `DroidProtocolError(message, ...)` | `code`, `data` | ### Detailed session contracts #### Session construction contract `Session(...)` accepts: | Argument | Type | | --- | --- | | `cwd` | `str \| Path \| None` | | `model` | `str \| None` | | `reasoning_effort` | `ReasoningEffort \| None` | | `config` | `SessionConfig \| None` | | `interactions` | `InteractionHandlers \| None` | | `runtime` | `Runtime \| None` | | `api_key` | `str \| None` | `SessionConfig` fields: | Field | Type | | --- | --- | | `mode` | `Mode \| None` | | `autonomy` | `Autonomy \| None` | | `spec_model` | `str \| None` | | `spec_reasoning_effort` | `ReasoningEffort \| None` | | `mcp_servers` | `Sequence[McpServerConfig]` | | `machine_id` | `str \| None` | | `tags` | `Sequence[SessionTag]` | | `session_source` | `SessionSource \| None` | | `auto_reject_permission_requests` | `bool \| None` | | `disable_builtin_skills` | `bool \| None` | | `system_prompt` | `SystemPromptConfig \| None` | | `additional_tools` | `Iterable[str] \| None` | | `enabled_tools` | `Iterable[str] \| None` | | `disabled_tools` | `Iterable[str] \| None` | | `restrict_tools` | `Iterable[str] \| None` | `Session.resume(session_id, ...)` accepts only values that can be reattached: | Argument | Type | | --- | --- | | `session_id` | `str` | | `interactions` | `InteractionHandlers \| None` | | `mcp_servers` | `Sequence[McpServerConfig]` | | `runtime` | `Runtime \| None` | | `api_key` | `str \| None` | | `disabled_tools` | `Iterable[str] \| None` | | `auto_reject_permission_requests` | `bool \| None` | | `disable_builtin_skills` | `bool \| None` | | `session_source` | `SessionSource \| None` | #### Session state schemas | Type | Fields | | --- | --- | | `SessionSettings` | `model`, `reasoning_effort`, `mode`, `autonomy`, `spec_model`, `spec_reasoning_effort`, `tags`, `sandbox`, `system_prompt`, `additional_tools`, `enabled_tools`, `disabled_tools`, `restrict_tools` | | `SessionSettingsUpdate` | `model`, `reasoning_effort`, `mode`, `autonomy`, `spec_model`, `spec_reasoning_effort`, `tags`, `additional_tools`, `enabled_tools`, `disabled_tools`, `restrict_tools`, `compaction_threshold_check_enabled` | | `SandboxSettings` | `enabled: bool`, `mode: str \| None = None` | | `SessionTag` | `name`, `metadata` | | `SavedSession` | `id`, `title`, `owner`, `message_count`, `modified_at`, `created_at`, `cwd`, `is_favorite` | Every `SessionSettingsUpdate` field defaults to `None`. Its field types match the `update_settings()` table above. `list_sessions()` returns `list[SavedSession]`. Its filters are `cwd`, `all_workspaces`, and `limit`. `on_notification(callback, type=None)` passes `Mapping[str, object]` to the callback and returns an unsubscribe function. `SessionSource` requires `platform: SessionPlatform`; every other attribution field is optional and defaults to `None`. Required combinations are validated when converted to the wire protocol. ## Known limitations - The SDK runs local Droid subprocess sessions only. - The API is asyncio-only. - Hooks are configured through Droid files, not Python callbacks. - Image URLs are not supported by the local runtime. - One session can run one turn at a time. Build Node.js integrations with the Droid SDK. Run Droid non-interactively from the command line. # Droid TypeScript SDK Use the Factory Droid SDK for TypeScript to build custom agents, workflows, and product integrations. The Droid SDK lets you run the same agent harness that powers Factory's CLI, desktop application, and web platform from your own code. It manages conversation context, tool execution, permissions, streaming, model selection, and multi-step execution. Your application decides when Droid runs, which tools and models it can use, and what happens to the result. ## Overview The TypeScript SDK can start Droid as a Node.js subprocess or connect to an existing daemon. It provides one-shot runs, persistent sessions, streaming, model and session discovery, and in-process MCP tools. ### Choose an API | Goal | API | | --------------------------------------- | ----------------- | | Run one prompt and get the final result | `run()` | | Keep context across prompts | `createSession()` | | Continue a saved session | `resumeSession()` | | Discover selectable models | `listModels()` | | Find saved sessions | `listSessions()` | Start with `run()` for one prompt. Use a session when later prompts need the same conversation history. ### What you can build Invoke Droid when a pull request opens, a support ticket is escalated, an incident is created, or scheduled repository maintenance is due. The agent does not need a chat interface. ## Install and authenticate Install the public npm package: ```bash npm install @factory/droid-sdk ``` If your code imports Zod for structured output or SDK MCP tools, use Zod 3: ```bash npm install zod@^3.24.0 ``` Import Node.js APIs from `@factory/droid-sdk/node`. Browser and daemon clients are available from the root entrypoint. This guide covers the high-level public API; the exported TypeScript declarations define the complete compile-time API. The SDK requires Node.js 18 or later. By default, the SDK starts the `droid` CLI from `PATH`. Pass `execPath` to use a different `droid` executable. Set an API key: ```bash export FACTORY_API_KEY="your-key" ``` The SDK reads `FACTORY_API_KEY` from the process environment by default. You can also pass `apiKey` to `run()`, `createSession()`, or `resumeSession()`. ## Quick start ### Run one prompt ```typescript import { run } from '@factory/droid-sdk/node'; const result = await run('Summarize this repository.'); if (!result.success) { throw new Error(result.error?.message ?? `Run failed: ${result.subtype}`); } console.log(result.text); ``` `run()` creates a session, consumes its complete message stream, returns the terminal `DroidResult`, and closes the session. ### Continue a conversation ```typescript import { createSession, DroidMessageType } from '@factory/droid-sdk/node'; const session = await createSession(); try { for await (const message of session.stream('What does this project do?')) { if (message.type === DroidMessageType.Assistant) { console.log(message.text); } } for await (const message of session.stream('What should I test first?')) { if (message.type === DroidMessageType.Assistant) { console.log(message.text); } } } finally { await session.close(); } ``` In the above example, the second prompt uses context from the first turn. Remember to close sessions that your code creates. ## Core concepts ### Session A session holds conversation history, settings, and a working directory. Create one with `createSession()` or load a saved session with `resumeSession()`. ### Turn A turn starts when you send one prompt with `session.stream()`. A session can run one turn at a time. ### Stream message `session.stream()` returns an async iterable of complete messages. Narrow each message by its `type`. Messages can report user input, assistant output, tool activity, hooks, errors, and a terminal `DroidResult`. ### Result A normally completed turn ends with a `DroidResult`. `run()` consumes the stream and returns that same result type. See [Streaming and results](#streaming-and-results) for result handling, errors, token usage, and cancellation. ### Ownership and cleanup Cleanup depends on how the session was created. | API | Cleanup | | ----------------- | -------------------------------- | | `run()` | Closes its session automatically | | `createSession()` | Call `await session.close()` | | `resumeSession()` | Call `await session.close()` | Use `finally` blocks so cleanup also runs after an error. ## Sessions Use a session when prompts need shared conversation history. A session owns its working directory, settings, and one active turn at a time. ### Create and use a session ```typescript import { createSession, DroidMessageType } from '@factory/droid-sdk/node'; const session = await createSession({ cwd: process.cwd(), }); async function send(prompt: string) { for await (const message of session.stream(prompt)) { if (message.type === DroidMessageType.Assistant) { console.log(message.text); } } } try { await send('What does this project do?'); await send('What should I test first?'); } finally { await session.close(); } ``` The second turn uses context from the first. `cwd` defaults to `process.cwd()`. Model and reasoning defaults come from Droid settings. ### Customize the system prompt Set `systemPrompt` when creating a session. A string replaces Droid's default behavioral prompt: ```typescript const session = await createSession({ systemPrompt: 'Act as a focused dependency-analysis agent.', }); ``` To keep Droid's default prompt and append instructions: ```typescript const session = await createSession({ systemPrompt: { type: 'preset', preset: 'droid', append: 'Prioritize security findings and cite relevant files.', }, }); ``` `systemPrompt` is available with `run()`, `createSession()`, and `droid.sessions.create()`. It is set at session creation and preserved when the session is resumed or forked. ### Resume a saved session ```typescript import { DroidMessageType, resumeSession } from '@factory/droid-sdk/node'; async function continueSession(sessionId: string) { const session = await resumeSession(sessionId); try { for await (const message of session.stream( 'Continue from the last conversation.' )) { if (message.type === DroidMessageType.Assistant) { console.log(message.text); } } } finally { await session.close(); } } ``` `resumeSession()` restores the saved conversation, working directory, and session settings. It does not accept `cwd`, `modelId`, or reasoning options. Resume options can provide new permission and AskUser handlers, disabled tools, MCP servers, an abort signal, and observability sinks. ### Read session state Each handle exposes its ID, current working directory, and current settings. ```typescript console.log(session.id); console.log(session.cwd); console.log(session.settings.modelId); console.log(session.settings.reasoningEffort); console.log(session.settings.interactionMode); ``` `settings` is read-only. The SDK updates `settings` and `cwd` when Droid reports a change. ### Update session settings Use `updateSettings()` for changes that should apply to later turns in the current session. ```typescript import { AutonomyLevel, ReasoningEffort } from '@factory/droid-sdk/node'; await session.updateSettings({ modelId: 'model-id', reasoningEffort: ReasoningEffort.High, autonomyLevel: AutonomyLevel.Low, disabledToolIds: ['Execute'], }); ``` Updatable settings include: - model and reasoning effort - interaction mode and autonomy level - spec-mode model settings - tool availability overrides Interaction mode controls whether Droid operates normally in Auto mode or produces a read-only plan in Spec mode. Autonomy level controls which actions require approval while Droid is in Auto mode. Tool availability overrides can enable, disable, or restrict the tools available to the session. Use `enterSpecMode()` to enter Spec mode. Return to Auto mode with `updateSettings({ interactionMode: DroidInteractionMode.Auto })`. The updated values are available through `session.settings`. ### Rename a session Use `rename()` after the first turn to replace the generated title. ```typescript await session.rename({ title: 'Authentication review' }); ``` ### List saved sessions `listSessions()` reads local Droid session storage. It does not start Droid or call the Factory API. List sessions for the current working directory: ```typescript import { listSessions } from '@factory/droid-sdk/node'; const sessions = await listSessions({ limit: 10 }); for (const session of sessions) { console.log(session.id, session.title, session.modifiedTime); } ``` List sessions across all working directories: ```typescript const sessions = await listSessions({ fetchOutsideCWD: true, limit: 10, }); ``` Results are sorted by `modifiedTime`, newest first. Each item includes its ID, title, owner, message count, creation time, modification time, and working directory when available. ### What persists Droid saves the session data needed to continue later. | Restored from the saved session | Attach again when resuming | | ------------------------------- | -------------------------- | | Conversation history | Permission handler | | Working directory | AskUser handler | | Title | SDK MCP servers | | Session settings | Observability sinks | | | Abort signal | Session settings include model, reasoning, Auto or Spec interaction mode, autonomy, spec-mode model settings, tool availability overrides, and the custom system prompt. Handlers, observability sinks, abort signals, and SDK MCP servers are runtime objects. Droid cannot serialize them with the session, so attach them again when calling `resumeSession()`. They do not need to be the same object instances or implementations used when the session was created. ### Close sessions safely Call `close()` in `finally` when your code creates or resumes a session. Closing releases the subprocess, subscriptions, and SDK-owned MCP servers. ```typescript const session = await createSession(); try { // Use the session. } finally { await session.close(); } ``` ## Streaming and results Each call to `session.stream()` starts one turn and returns an async iterable. By default, it yields complete messages and ends with a terminal `DroidResult`. ### Complete messages ```typescript for await (const message of session.stream('Find the failing test.')) { switch (message.type) { case DroidMessageType.Assistant: console.log(message.text); break; case DroidMessageType.ToolCall: console.log(`Tool: ${message.name}`); break; case DroidMessageType.Result: console.log(message.subtype); break; } } ``` Default message types: | Type | Purpose | | ------------- | ------------------------------------ | | `assistant` | Complete assistant message | | `user` | User message recorded by the session | | `tool_call` | Complete tool request | | `tool_result` | Complete tool result | | `hook` | Hook execution | | `error` | Runtime error event | | `result` | Terminal `DroidResult` | ### Partial events Enable partial events only when the application needs live text, thinking, tool progress, token usage, or state updates. ```typescript for await (const event of session.stream('Explain the test failure.', { includePartialMessages: true, })) { if (event.type === DroidMessageType.AssistantTextDelta) { process.stdout.write(event.text); } } ``` Partial streams can also include thinking deltas, tool-call deltas, tool progress, token updates, permission results, settings changes, working-state changes, and MCP status. ### Handle the result `DroidResult` is a discriminated union. Check `subtype` or `success` before using the response. ```typescript const result = await run('Run the test suite.'); switch (result.subtype) { case 'success': console.log(result.text); break; case 'interrupted': console.log('The run was interrupted.'); break; case 'error_during_execution': case 'error_structured_output': console.error( result.structuredOutputError?.message ?? result.error?.message ?? 'Unknown error' ); break; } ``` | Subtype | Meaning | | ------------------------- | ------------------------------------------------- | | `success` | The turn completed successfully | | `interrupted` | The turn was cancelled or permission declined | | `error_during_execution` | Droid could not complete the requested work | | `error_structured_output` | Structured-output generation or validation failed | Every result includes: ```typescript console.log(result.sessionId); console.log(result.durationMs); console.log(result.turnCount); console.log(result.messages); console.log(result.text); console.log(result.tokenUsage); ``` `run()` returns the terminal result directly. When consuming `session.stream()`, handle the message whose type is `DroidMessageType.Result`. ### Token and context usage `DroidResult.tokenUsage` reports raw token counts for the current SDK turn and, when available, the Factory Service Credits (FSC) charged for that turn. One turn is one `run()` or `session.stream()` call. This is not the session's lifetime total. ```typescript if (result.tokenUsage) { console.log(`input: ${result.tokenUsage.inputTokens}`); console.log(`output: ${result.tokenUsage.outputTokens}`); console.log(`cache read: ${result.tokenUsage.cacheReadTokens}`); console.log(`cache created: ${result.tokenUsage.cacheCreationTokens}`); if (result.tokenUsage.factoryCredits !== undefined) { console.log(`Factory Service Credits: ${result.tokenUsage.factoryCredits}`); } } ``` | Field | What it measures during the turn | | --------------------- | ------------------------------------------------------------------------------------------------------------------------------- | | `inputTokens` | Provider-reported input across model calls, which can include instructions, tools, conversation history, and the latest message | | `outputTokens` | Output generated across model calls, not only the final visible response | | `cacheReadTokens` | Input tokens reused from the prompt cache | | `cacheCreationTokens` | Input tokens written to the prompt cache | | `thinkingTokens` | Reasoning tokens reported separately when available | | `factoryCredits` | Factory Service Credits charged, when available | One turn can make several model calls while Droid uses tools. Their usage is combined in the result, along with usage from delegated work. Earlier conversation can be counted again when it is included as input to a later turn. Cached input is also reported through the separate cache fields. `tokenUsage` is `null` when usage is unavailable. Partial streams can emit `DroidMessageType.TokenUsageUpdate`. Unlike the per-turn result, these events contain cumulative committed usage for the session. `getContextStats()` measures current context occupancy, not cumulative token usage: ```typescript const stats = await session.getContextStats(); console.log(stats.used, stats.remaining, stats.limit, stats.accuracy); ``` | Field | Meaning | | ----------- | ----------------------------------------------------------------------------------------------------- | | `used` | Estimated tokens currently occupied by prompts, tools, instructions, and conversation history | | `limit` | Active model's maximum input window; unresolved model IDs fall back to the effective compaction limit | | `remaining` | `Math.max(0, limit - used)` | | `accuracy` | Currently `estimated`, because Droid approximates tokens from character counts | `remaining` estimates room in the model's input window. It does not necessarily mean tokens remaining before compaction. Droid can compact at a lower configured threshold, commonly 250,000 tokens, even when the active model's input window is larger. ### Errors An error event and a thrown exception mean different things: - An `error` message reports a runtime problem during the turn. A terminal `DroidResult` normally follows and describes the final outcome. - A failed `DroidResult` means the stream completed, but the requested work or structured output failed. - A thrown exception means the SDK operation itself could not finish, such as an abort, connection failure, protocol error, or concurrent stream attempt. Handle stream exceptions around the iteration: ```typescript try { for await (const message of session.stream('Run the tests.')) { if (message.type === DroidMessageType.Error) { console.error(message.message); } } } catch (error) { console.error('The stream could not finish:', error); } ``` `run()` follows the same distinction: agent failures are returned as failed results, while setup, transport, protocol, and abort failures are thrown. ### Stream concurrency One session handle can have one active stream. Starting another active stream throws `ConcurrentStreamError`. ### Cancellation and interruption Use an abort signal to cancel one turn: ```typescript const controller = new AbortController(); setTimeout( () => controller.abort(new Error('Timed out after 5 seconds')), 5_000 ); for await (const message of session.stream('Perform a long review.', { abortSignal: controller.signal, })) { // Handle messages. } ``` The stream interrupts the active turn and throws the abort reason. Use `session.interrupt()` when another part of the application needs to stop the active turn: ```typescript await session.interrupt(); ``` The stream normally continues to a terminal result with subtype `interrupted`. The session remains available for later prompts. Breaking out of the stream also interrupts the active turn: ```typescript for await (const message of session.stream('Investigate every failing test.')) { if (message.type === DroidMessageType.Assistant) { console.log(message.text); // Interrupt this turn instead of merely hiding its remaining output. break; } } ``` Breaking out of the loop before the turn finishes sends an interrupt request to Droid. This cancels the remaining model and tool work for that turn. The session remains open and can accept another prompt, but this loop does not receive the interrupted turn's terminal `DroidResult`. A session-level abort signal passed to `createSession()` closes the entire session when aborted. ## Models Model IDs are strings because availability depends on the account, organization policy, and daemon configuration. Omit `modelId` to use the Droid default. ### Discover available models Use `listModels()` to inspect the models available to the current account and project before creating a session. ```typescript import { listModels } from '@factory/droid-sdk/node'; const models = await listModels(); for (const model of models) { console.log(model.id, model.defaultReasoningEffort); } ``` `listModels()` starts a one-shot Droid process, requests the model catalog without creating a session, and closes the process before returning. Disabled models are hidden by default. Include them when building a model picker or explaining why a model cannot be selected: ```typescript const models = await listModels({ includeDisabled: true }); for (const model of models) { const status = model.disabled ? `disabled: ${model.disabledReason}` : 'available'; console.log(model.displayName, status); } ``` Applications connected to a Droid daemon can request the same catalog: ```typescript const models = await droid.models.list({ includeDisabled: true, }); ``` Both APIs return `ModelInfo[]`. The catalog includes built-in and configured custom models after applying feature flags, region availability, and organization policy. Custom endpoint and credential configuration is never returned. | Field | Meaning | | --------------------------- | ---------------------------------------------- | | `id` | Value accepted by `modelId` | | `displayName` | Full display name | | `shortDisplayName` | Compact display name | | `modelProvider` | Model provider | | `supportedReasoningEfforts` | Reasoning levels accepted by the model | | `defaultReasoningEffort` | Reasoning level used when none is supplied | | `isCustom` | Whether the model is user-configured | | `disabled` | Whether policy prevents selecting the model | | `disabledReason` | Explanation when `disabled` is `true` | | `noImageSupport` | Whether image input is unavailable | | `supportsImageGeneration` | Whether the model can generate images | Optional fields provide the model tier, token multiplier, classification, and display badges. ### Select a model Pass a configured model ID when creating a session or running one prompt. ```typescript import { createSession, DroidMessageType, ReasoningEffort, } from '@factory/droid-sdk/node'; const session = await createSession({ modelId: 'model-id', reasoningEffort: ReasoningEffort.High, }); try { for await (const message of session.stream('Review this repository.')) { if (message.type === DroidMessageType.Assistant) { console.log(message.text); } } } finally { await session.close(); } ``` For an existing session, call `updateSettings()`: ```typescript await session.updateSettings({ modelId: 'model-id', reasoningEffort: ReasoningEffort.High, }); ``` Set `reasoningEffort` only to a level supported by the selected model. Omit it to use the model's configured default. ### Use the Factory Router Set `modelId` to `"auto"` to let Factory choose a model for each turn. ```typescript import { run } from '@factory/droid-sdk/node'; const result = await run('Find the cause of the failing tests.', { modelId: 'auto', }); if (!result.success) { throw new Error(result.error?.message ?? `Run failed: ${result.subtype}`); } console.log(result.text); ``` The selected model can change between turns. Choose a fixed model ID when runs must use the same model. ### Configure mode-specific models The primary model handles normal turns. The spec-mode model handles turns in spec mode. ```typescript import { createSession, DroidInteractionMode, ReasoningEffort, } from '@factory/droid-sdk/node'; const session = await createSession({ modelId: 'auto', interactionMode: DroidInteractionMode.Spec, specModeModelId: 'model-id', specModeReasoningEffort: ReasoningEffort.High, }); try { for await (const _message of session.stream( 'Draft an implementation plan.' )) { // Consume messages. } } finally { await session.close(); } ``` You can also choose the model when entering spec mode: ```typescript await session.enterSpecMode({ specModeModelId: 'model-id', specModeReasoningEffort: ReasoningEffort.High, }); ``` ### Use a custom model Configure custom models in `~/.factory/settings.json`. See [Model Independence](/model-independence/byok) for provider, credential, and model settings. After configuration, pass the custom model ID anywhere the SDK accepts `modelId`. The ID uses `custom:` followed by the configured `model` value. ```typescript import { run } from '@factory/droid-sdk/node'; const result = await run('Summarize this repository.', { modelId: 'custom:gpt-4o-mini', }); if (!result.success) { throw new Error(result.error?.message ?? `Run failed: ${result.subtype}`); } console.log(result.text); ``` ## Inputs and outputs Pass images, documents, or an output schema in the options for the turn that should use them. Both `run()` and `session.stream()` accept these options. ### Images and documents ```typescript const result = await run('Describe this image.', { images: [ { type: 'base64', data: base64Png, mediaType: 'image/png', }, ], }); ``` `data` contains the base64-encoded image bytes. Supported media types are `image/jpeg`, `image/png`, `image/gif`, and `image/webp`. #### Documents ```typescript const result = await run('Summarize this document.', { files: [ { type: 'text', mediaType: 'text/plain', data: report, name: 'report.txt', }, ], }); ``` Document inputs support: | Source | `type` | `mediaType` | `data` | | ------------ | -------- | ----------------- | ----------------------------- | | Plain text | `text` | `text/plain` | The document text | | PDF document | `base64` | `application/pdf` | Base64-encoded PDF file bytes | `name` is optional for both forms. Attach content through `data`; a filesystem path by itself is not a document upload. ### Structured output Use `outputFormat` when another system needs a JSON object that follows a schema. ```typescript import { OutputFormatType, run } from '@factory/droid-sdk/node'; const result = await run('Return the repository name and test command.', { outputFormat: { type: OutputFormatType.JsonSchema, schema: { type: 'object', properties: { repository: { type: 'string' }, testCommand: { type: 'string' }, }, required: ['repository', 'testCommand'], additionalProperties: false, }, }, }); if (!result.success) { throw new Error( result.structuredOutputError?.message ?? result.error?.message ?? `Structured output failed (${result.subtype}).` ); } if (!result.structuredOutput) { throw new Error('No structured output was returned.'); } console.log(JSON.stringify(result.structuredOutput, null, 2)); ``` Invalid or missing schema output normally produces a failed result with subtype `error_structured_output`; `structuredOutputError` contains the available validation details. `structuredOutput` is typed as `unknown`, because a JSON Schema does not provide a TypeScript runtime validator. Check that it is present and validate or narrow its shape before reading fields or passing it to another system. A successful result can still omit structured output, so do not use `result.success` as the only check. ## Permissions and user input ### Autonomy Autonomy controls which actions require approval in Auto mode. ```typescript import { AutonomyLevel } from '@factory/droid-sdk/node'; const session = await createSession({ autonomyLevel: AutonomyLevel.Off, }); ``` | Level | Behavior | | -------- | --------------------------------------- | | `Off` | Ask before every action | | `Low` | Allow file edits and read-only commands | | `Medium` | Allow reversible commands | | `High` | Allow commands without approval prompts | When omitted, Droid uses the configured default. Autonomy does not select Auto or Spec mode; `interactionMode` does that. ### Permission handler Use `permissionHandler` to approve or reject requests that reach the client. The handler may be synchronous or asynchronous. ```typescript import { ToolConfirmationOutcome, ToolConfirmationType, } from '@factory/droid-sdk/node'; const session = await createSession({ autonomyLevel: AutonomyLevel.Off, permissionHandler({ toolUses, options }) { const createsOnly = toolUses.length > 0 && toolUses.every((use) => use.details.type === ToolConfirmationType.Create); const desired = createsOnly ? ToolConfirmationOutcome.ProceedOnce : ToolConfirmationOutcome.Cancel; const offered = options.find((option) => option.value === desired); if (!offered) throw new Error(`Outcome not offered: ${desired}`); return desired; }, }); ``` `toolUses` describes the pending actions. `options` contains the outcomes available for this request. Return one of those values. Common outcomes are: | Outcome | Effect | | --------------- | --------------------------------------- | | `ProceedOnce` | Approve this request | | `ProceedAlways` | Approve and persist the applicable rule | | `Cancel` | Reject the request | The available outcomes vary by request. Returning an unavailable value, throwing from the handler, or omitting the handler cancels the request. ### AskUser handler `askUserHandler` answers questions Droid asks during a turn. Connect it to a form, prompt, or other application UI. ```typescript const result = await run('Ask me which environment to deploy.', { askUserHandler({ questions }) { return { answers: questions.map((question) => ({ index: question.index, question: question.question, answer: question.options[0] ?? 'none', })), }; }, }); ``` Each answer must copy the question's `index` and `question`. To decline the questionnaire: ```typescript return { cancelled: true, answers: [] }; ``` Without a handler, AskUser requests are declined. ### Tool controls Disable tools when creating, resuming, or updating a session: ```typescript const session = await createSession({ disabledToolIds: ['Execute'], }); await session.updateSettings({ disabledToolIds: ['Execute', 'Edit'], }); ``` Tool IDs are strings. Use `listTools()` to inspect the tools available to the current session. ```typescript const tools = await session.listTools(); for (const tool of tools) { console.log(tool.id, tool.allowed); } ``` The public SDK provides subtractive `disabledToolIds`. It does not expose a restrictive allowlist. ## Extensions Use the extension that matches the job: | Extension | Use it for | | -------------------- | ------------------------------------------------ | | Skills | Reusable instructions and supporting files | | SDK MCP tools | In-process TypeScript functions exposed to Droid | | External MCP servers | Tools provided by another process or service | | Hooks | Commands that run at defined lifecycle events | ### Skills List the skills available to the session: ```typescript const { skills } = await session.listSkills(); for (const skill of skills) { console.log(skill.name, skill.enabled); } ``` Enable or disable a skill at the user or project level: ```typescript import { SettingsLevel } from '@factory/droid-sdk/node'; await session.setSkillDisabled({ skillName: 'skill-name', disabled: true, settingsLevel: SettingsLevel.Project, }); ``` Skills can come from project, personal, built-in, or automation settings. ### In-process MCP tools Use `tool()` to define an in-process TypeScript function with a Zod input schema. The SDK and Droid CLI must use compatible Factory protocol versions. ```typescript import { z } from 'zod'; import { createSdkMcpServer, createSession, tool, } from '@factory/droid-sdk/node'; const server = createSdkMcpServer({ name: 'review-tools', tools: [ tool( 'lookup_owner', 'Returns the owner of a file', { path: z.string() }, ({ path }) => `Owner for ${path}: platform-team` ), ], }); const session = await createSession({ mcpServers: [server], }); try { for await (const _message of session.stream( 'Find the owner of src/index.ts.' )) { // Consume messages. } } finally { await session.close(); } ``` The SDK starts an authenticated loopback MCP server and closes it with the session. Tool calls still follow the session's autonomy and permission rules. Attach SDK MCP servers again when resuming a saved session because they are runtime objects, not persisted session data. ### External MCP servers Pass external MCP configuration when creating or resuming a session: ```typescript const session = await createSession({ mcpServers: [ { name: 'docs', type: 'http', url: 'https://example.com/mcp', headers: [], }, ], }); ``` Supported session configuration: | Transport | `type` field | Required configuration | | ------------- | ------------ | ---------------------- | | Local process | Omit | `command` and `args` | | HTTP | `http` | `url` and `headers` | | SSE | `sse` | `url` and `headers` | Inspect server status and discovered tools: ```typescript try { const { servers, summary } = await session.listMcpServers(); const tools = await session.listMcpTools(); console.log(servers, summary, tools); } finally { await session.close(); } ``` `addMcpServer()`, `removeMcpServer()`, and `toggleMcpServer()` change the user's Droid configuration. Use `authenticateMcpServer({ serverName })` when a server requires authentication. ### Hooks Hooks run shell commands at defined lifecycle events. Configure them in `.factory/hooks.json`: ```json { "PreToolUse": [ { "matcher": "Execute", "hooks": [ { "type": "command", "command": "echo before Execute" } ] } ] } ``` Hooks can run before or after tools, when prompts are submitted, when notifications arrive, during compaction, and when sessions or subagents stop. Hook events are available in the stream: ```typescript for await (const message of session.stream('Run the tests.')) { if (message.type === DroidMessageType.Hook) { console.log(message.status, message.command); } } ``` The SDK exposes hook schemas and stream events, but it does not provide a programmatic hook-registration API. ## Session lifecycle Fork, compact, and rewind return ready successor sessions and retire the source wrapper. ### Fork `fork()` creates a new session from the current conversation. ```typescript const forked = await session.fork(); ``` ### Compact `compact()` summarizes older conversation history and returns the successor in `session`. ```typescript const { session: compacted, removedCount } = await forked.compact(); ``` ### Rewind `rewind()` returns the conversation to an earlier message and can restore or delete files. ```typescript const { session: rewound, restoredCount, deletedCount, } = await compacted.rewind({ messageId, filesToRestore, filesToDelete, forkTitle: 'Before the failed change', }); ``` ### Successor ownership After a successful replacement, active operations on the source wrapper throw `SessionReplacedError`. Its `id`, `settings`, and `cwd` remain readable, and `close()` is a no-op. The persisted source session can still be loaded later with `resumeSession(sourceId)`. Do not replace a session while it has an active stream. ## Spec mode Spec mode lets Droid inspect the codebase and propose a plan without changing files. Start a session in Spec mode or switch an existing session with `enterSpecMode()`: ```typescript const session = await createSession({ interactionMode: DroidInteractionMode.Spec, }); // Or switch an existing Auto session. await session.enterSpecMode(); ``` `enterSpecMode()` can also set the model used for planning: ```typescript await session.enterSpecMode({ specModeModelId: 'model-id', specModeReasoningEffort: ReasoningEffort.High, }); ``` ### Leave without approving a plan Call `updateSettings()` to return the session to Auto mode without approving a plan or starting implementation: ```typescript import { DroidInteractionMode } from '@factory/droid-sdk/node'; // Return to Auto mode without approving or implementing the plan. await session.updateSettings({ interactionMode: DroidInteractionMode.Auto, }); ``` Changing the interaction mode does not select a `ToolConfirmationOutcome` or approve an `ExitSpecMode` permission request. `updateSettings()` is available on Node `DroidSession` handles returned by `createSession()`, `resumeSession()`, and replacement operations such as `fork()`, `compact()`, and `rewind()`. Daemon clients use `droid.sessions.updateSettings()`. ### Approve a plan When the plan is ready, Droid sends an `ExitSpecMode` permission request. Its details contain the plan. Unlike changing the interaction mode with `updateSettings()`, approving it with `ProceedOnce` accepts the completed plan, returns to Auto mode, and continues implementation in the same session: ```typescript import { AutonomyLevel, createSession, DroidInteractionMode, DroidMessageType, ToolConfirmationOutcome, ToolConfirmationType, } from '@factory/droid-sdk/node'; const session = await createSession({ interactionMode: DroidInteractionMode.Auto, autonomyLevel: AutonomyLevel.Low, permissionHandler({ toolUses }) { const request = toolUses.find( ({ details }) => details.type === ToolConfirmationType.ExitSpecMode ); if (request?.details.type === ToolConfirmationType.ExitSpecMode) { console.log(request.details.plan); return ToolConfirmationOutcome.ProceedOnce; } return ToolConfirmationOutcome.Cancel; }, }); try { await session.enterSpecMode(); for await (const message of session.stream('Plan and update README.md.')) { if (message.type === DroidMessageType.Result) { console.log(message.subtype); } } } finally { await session.close(); } ``` The handler can be asynchronous. An interactive application should display the plan and wait for the user's decision before returning. Return `Cancel` to reject the plan. The turn ends with an `interrupted` result, and the session remains available for another prompt. Some `ExitSpecMode` requests offer `ProceedNewSession` variants. Return one of those offered values to hand implementation to a new session. Return only values present in the request's `options`. ## Observability and advanced APIs ### Observability ```typescript import { run } from '@factory/droid-sdk/node'; const result = await run('Check the repository status.', { observability: { logger: { log(event) { console.log(event.level, event.name, event.attributes); }, }, }, }); if (!result.success) { throw new Error(result.error?.message ?? `Run failed: ${result.subtype}`); } ``` Pass `logger`, `metrics`, or `tracing` sinks through `observability`. Sink methods must not throw. The SDK excludes prompts, messages, tool inputs, raw output, and stack traces from observability events. Process startup emits telemetry before the first turn, and `close()` emits a request log. To verify turn-specific logs, snapshot the count after creating the session and check it before closing. ### Raw notifications Use `onNotification()` only when `session.stream()` does not expose the event you need: ```typescript const unsubscribe = session.onNotification( (notification) => console.log(notification.type), { type: 'settings_updated' } ); unsubscribe(); ``` ## Examples | Example | Demonstrates | | --- | --- | | Multi-turn session | Keep context across prompts | | Initialization metadata | Read session metadata after initialization | | Session settings | Read and update active-session settings | | Rename a session | Change a saved session's title | | List sessions | Discover saved sessions | | Model discovery | Discover selectable and disabled models | | Session stream | Consume complete stream messages | | Result metadata | Inspect terminal result metadata | | Abort a stream | Cancel a turn with an abort signal | | Interrupt a session | Request interruption during a turn | | Image attachment | Send a local image with a prompt | | Structured output | Validate model output against a schema | | Permission handler | Respond to tool permission requests | | AskUser handler | Answer questions from Droid | | Tool controls | Restrict and configure native tools | | Enter Spec mode | Plan before implementation | | SDK MCP tool | Define an in-process MCP tool | | Hook execution | Observe configured hook events | | Observability | Capture SDK logs and telemetry | ## API reference ### Entrypoint Import functions, classes, enums, and types from `@factory/droid-sdk/node`. ### Functions | API | Returns | | ------------------------------------ | ---------------------------- | | `run(prompt, options?)` | `Promise` | | `createSession(options?)` | `Promise` | | `resumeSession(sessionId, options?)` | `Promise` | | `listSessions(options?)` | `Promise` | | `listModels(options?)` | `Promise` | | `createSdkMcpServer(options)` | `SdkMcpServer` | | `tool(...)` | `DroidTool` | ### `DroidSession` | Member | Purpose | | ------------------------- | -------------------------------------------- | | `id` | Session ID | | `cwd` | Live working directory | | `settings` | Live read-only settings | | `stream()` | Run one turn | | `interrupt()` | Stop the active turn | | `updateSettings()` | Update session settings | | `enterSpecMode()` | Change to spec mode | | `listTools()` | List normalized tools | | `listSkills()` | List skills | | `setSkillDisabled()` | Enable or disable a skill | | `addMcpServer()` | Add an MCP server | | `removeMcpServer()` | Remove an MCP server | | `toggleMcpServer()` | Enable or disable an MCP server | | `listMcpServers()` | List MCP servers | | `listMcpTools()` | List MCP tools | | `authenticateMcpServer()` | Start MCP authentication | | `getContextStats()` | Read context usage | | `getRewindInfo()` | Inspect rewindable file changes | | `rewind()` | Create a rewound successor | | `compact()` | Create a compacted successor | | `fork()` | Create a forked successor | | `rename()` | Change the session title | | `onNotification()` | Subscribe to raw notifications | | `close()` | Close the session and owned resources | ### Stream types | Type | Purpose | | -------------------- | ------------------------------------------------- | | `DroidStreamMessage` | Complete messages returned by the default stream | | `DroidStreamEvent` | Complete messages plus partial events | | `DroidResult` | Terminal result returned by `run()` | | `DroidMessageType` | Runtime constants for narrowing stream messages | | `TokenUsage` | Token counts and optional Factory Service Credits | | `ModelInfo` | Model catalog entry returned by discovery APIs | `session.stream()` yields `DroidStreamMessage` by default and `DroidStreamEvent` when `includePartialMessages` is `true`. ### Result states | Subtype | `success` | `interrupted` | Meaning | | ------------------------- | --------- | ------------- | ------------------------- | | `success` | `true` | `false` | Turn completed | | `interrupted` | `false` | `true` | Turn was stopped | | `error_during_execution` | `false` | `false` | Runtime failure | | `error_structured_output` | `false` | `false` | Structured-output failure | Every result includes `sessionId`, `durationMs`, `tokenUsage`, `messages`, `text`, and `turnCount`. Check `success` before using successful output. ### Main enums | Enum | Values | | ---------------------- | -------------------------------------------------------------------------------- | | `AutonomyLevel` | `Off`, `Low`, `Medium`, `High` | | `DroidInteractionMode` | `Auto`, `Spec` | | `ReasoningEffort` | `None`, `Dynamic`, `Off`, `Minimal`, `Low`, `Medium`, `High`, `ExtraHigh`, `Max` | | `OutputFormatType` | `JsonSchema` | Model IDs and tool IDs are strings rather than closed enums. ### Main errors | Error | Meaning | | ------------------------- | --------------------------------------------- | | `ConcurrentStreamError` | A session handle already has an active stream | | `SessionReplacedError` | A source wrapper was replaced | | `SessionReplacementError` | Successor loading or rollback failed | | `DroidClientError` | Base client error | | `ConnectionError` | Droid process connection failed | | `TimeoutError` | A request timed out | | `ProtocolError` | Protocol or API response failed | | `SessionError` | Base session error | | `SessionNotFoundError` | Saved session was not found | | `InvalidSessionCwdError` | Saved working directory is invalid | | `ProcessExitError` | Droid subprocess exited unexpectedly | ## Known limitations - New-session spec handoff IDs require raw notification inspection. - SDK MCP tools require compatible protocol versions in the npm package and Droid CLI. Build asyncio integrations with the Droid SDK. Run Droid non-interactively from the command line. # Droid Computers Persistent, long-lived compute environments for Droid. Droid Computers are persistent compute environments that Droid can connect to and work on across sessions. Unlike ephemeral cloud templates that are destroyed after each session, Droid Computers **retain state**: installed packages, files, running services, and configuration all persist between sessions. ## Droid Computer types Factory supports two ways to use Droid Computers: - **Bring Your Own Machine (BYOM)**: You register a machine you already manage, such as a VPS, workstation, or on-prem server - **Managed Droid Computers**: Factory provisions and manages the underlying cloud compute environment for you Both types appear in **Settings → Droid Computers** and can be used as long-lived targets for Droid sessions. ## Using Droid Computers in sessions ### Factory App In the new session modal, select the **Computer** tab, pick an active Droid Computer, set a working directory (for example, `/home/factory-user/projects/my-app`), and start your session. Connection status (connecting / connected / error) is shown in real time. ### Slack Droid Computers can also be selected when creating sessions via the Slack integration. ## Managing Droid Computers From **Settings → Droid Computers**, you can: - Configure your own local Droid Computer - Browse available Droid Computers - Rename a Droid Computer - Delete a Droid Computer - Open the detail page for status and connection information (Managed Droid Computers only) ## Bring Your Own Machine (BYOM) You can register any machine you own (VPS, cloud VM, on-prem server, etc.) as a Droid Computer. For the full BYOM setup flow, Remote Access guidance, and management notes, see [Bring Your Own Machine (BYOM)](/droid-computers/byom). ## Managed Droid Computers The fastest way to get started is to create a Droid Computer directly from the Factory App. 1. Navigate to **Settings → Droid Computers**. 2. Click **Create**. 3. Give your Droid Computer a name. Factory provisions a cloud Droid Computer (4 CPU, 8GB RAM, 6GB swap) automatically. Provisioning progress is tracked live in the UI through the following steps: 1. **Creating Droid Computer**: allocating the cloud Droid Computer 2. **Setting up user**: creating the `factory-user` account with sudo access 3. **Configuring environment**: writing environment config, SSH keys, and service files 4. **Installing Droid**: downloading and installing the Droid binary 5. **Starting services**: launching the SSH and Droid daemon services Once provisioning completes, the Droid Computer status changes to **Active** and it becomes available for use in sessions. Platform-managed Droid Computers auto-pause when idle and auto-resume when a new session targets them. They also expose provisioning details, resource metrics, and remote update flows. ## CLI commands ### General | Command | Description | |---------|-------------| | `droid computer list` | List all Droid Computers (shows name, ID, status; marks current machine) | | `droid computer ssh ` | Open an interactive SSH session to any Droid Computer | | `droid computer port-forward ` | Forward local ports to a Droid Computer over the relay | ### BYOM setup | Command | Description | |---------|-------------| | `droid computer register [name]` | Register the current machine as a BYOM Droid Computer | | `droid computer remove` | Unregister the current machine and clean up local config | ### SSH options The `ssh` subcommand supports the following flags: - `--proxy`: Run as a stdio proxy (for use as a `ProxyCommand`) - `--port `: Target port on the remote machine (default: 22) - `--debug`: Enable verbose connection logging ## Port forwarding `droid computer port-forward ` forwards local TCP ports to a Droid Computer over the same relay tunnel used by `droid computer ssh` -- no direct port exposure is required. Managed Droid Computers are woken automatically before the connection is established. Mappings can be passed multiple at a time in one command: - `LOCAL:REMOTE` -- bind `LOCAL` locally and forward to `REMOTE` (e.g. `8080:80`) - `REMOTE` -- use the same port on both sides (e.g. `5432`) - `:REMOTE` -- pick a random free local port (e.g. `:9000`) ```bash # Forward local 8080 to remote 80, and local 5432 to remote 5432 droid computer port-forward mycomputer 8080:80 5432 ``` The `port-forward` subcommand supports the following flags: - `--address ` -- Local address to bind (default: `127.0.0.1`) - `--debug` -- Enable debug output Only TCP is supported. ## Monitoring Click on any Droid Computer in **Settings → Droid Computers** to open its detail page. For platform-managed Droid Computers, the detail page shows live resource metrics: Percentage utilization over time. Used vs. total memory. Used vs. total disk space. The detail page also shows the current **daemon version** and connection status. Possible Droid Computer statuses are: Initial setup in progress. Ready for sessions. Provisioning or runtime failure. ## Updating the daemon The Droid Computer detail page displays the running daemon version alongside the latest available version. When an update is available, an **Update** button appears. Clicking it triggers a remote update: the daemon downloads the new binary, restarts, and reconnects automatically. ## Git credentials - **BYOM**: Droid uses whatever git credentials are already configured on the machine. - **Managed Droid Computers**: If you have a personal GitHub integration with Factory, credentials are added automatically for authenticated repository access. When you add new GitHub App installations or change repository access, credentials are refreshed automatically on the next session connection. ## SSH & IDE integration `droid computer ssh ` establishes a secure WebSocket tunnel through the daemon. No direct SSH port exposure is required. Factory generates and manages a dedicated Ed25519 SSH key pair (stored in `~/.factory/.ssh/`), separate from your personal SSH keys. The public key is injected on each connection for passwordless authentication. ### VS Code Remote-SSH You can use `droid computer ssh` as a `ProxyCommand` for VS Code Remote-SSH. Add the following to your `~/.ssh/config`: ```text Host factory-* ProxyCommand droid computer ssh %h --proxy User factory-user StrictHostKeyChecking accept-new UserKnownHostsFile ~/.factory/.ssh/known_hosts IdentityFile ~/.factory/.ssh/id_ed25519 IdentitiesOnly yes ``` The wildcard `factory-*` host applies to every Droid Computer, and `%h` passes the host name through to the CLI, so you connect to `factory-` (for example, `factory-mycomputer`) without adding a stanza per computer. Host keys are recorded on first connect (`StrictHostKeyChecking accept-new`) in a dedicated `~/.factory/.ssh/known_hosts` file, so a changed key surfaces an SSH warning instead of being accepted silently. This lets you open a full VS Code remote session on your Droid Computer. ## Security {/* sweep-allow: term-bullets */} - **Firewall**: Platform-managed Droid Computers use iptables rules to restrict inbound traffic to the daemon port only; all other ports are blocked by default. Note that `factory-user` has passwordless sudo access and could modify these rules. For a hard network-level boundary, use relay mode, which blocks all public traffic at the infrastructure level. - **SSH hardening**: Root login and password authentication are disabled. - **Relay mode**: BYOM Droid Computers route all traffic through Factory's relay service, so no public ports are exposed. Platform-managed Droid Computers can also be configured to use relay mode at the organization level for stricter network isolation. - **Git credentials**: Configured automatically per-user for authenticated repository access on platform-managed Droid Computers. Register your own machine as a Droid Computer. Use deprecated cloud-hosted templates as an alternative compute environment. # Bring Your Own Machine (BYOM) Enable and connect your own machine as a Droid Computer. Bring Your Own Machine (BYOM) lets you register your own Linux, macOS, or Windows machine as a Droid Computer. This is useful when you want Droid to work inside an environment you already manage, such as a VPS, cloud VM, workstation, or on-prem server. ## Before you start Make sure you have: - A machine running **Linux**, **macOS**, or **Windows** - The Droid CLI is installed on that machine - Authenticated with Factory (`/login` in an interactive session) - Network access to Factory APIs, specifically `relay.factory.ai` No inbound ports need to be opened. BYOM machines connect through Factory's relay service when you start the daemon with `--remote-access`. ## Register with the Factory App If you're using the Factory App, your machine will be registered automatically using your system's hostname. To enable it as a remote Droid Computer, open **Settings → Droid Computers** and switch on the **Remote Access** toggle. This will automatically connect your local machine to our relay service to be reachable from other devices, and the connection will remain active for as long as the Factory App is running. ## Register with the CLI ```bash droid daemon --remote-access ``` If not registered, this will prompt you for a Droid Computer name and register the machine automatically. If you want to register the machine manually first, you can still run `droid computer register [name]`, then start `droid daemon --remote-access` to connect through Factory's relay service. Only one Droid Computer registration is allowed per machine at a time. Run `droid computer remove` first, or delete the machine through the settings page, if you need to re-register. ## Troubleshooting availability Check **Settings → Droid Computers** to confirm the machine is listed. In **Settings → Droid Computers**, verify **Remote Access** is enabled for that machine. Make sure the daemon is running with `droid daemon --remote-access`. Contact an organization Manager, Owner, or Factory support. ## Managing and updating a BYOM machine - Use `droid computer list` to see registered Droid Computers - Use `droid computer ssh ` to open an SSH session to a Droid Computer - Use `droid computer remove` to unregister the current machine and clean up local config - Update the Droid CLI manually on the machine, then restart the daemon process BYOM Droid Computers use whatever git credentials are already configured on the machine. ## Security and networking notes {/* sweep-allow: term-bullets */} - **Relay mode**: BYOM machines route traffic through Factory's relay service, so no public ports are exposed - **Credentials**: Git credentials are not managed by Factory on BYOM machines; the machine's existing git configuration is used - **Ownership**: You are responsible for hardening and maintaining the underlying machine Return to the Droid Computers overview for CLI commands and management. Consider Factory-managed cloud templates as an alternative compute option. # Droid Computers API Provision and manage Factory computers: create, inspect, refresh, restart, and delete. ## List computers `GET /api/v0/computers` Returns the computers owned by the authenticated client, ordered newest first. Provisioning steps are included only when `includeProvisioningSteps=true`. ```bash curl 'https://api.factory.ai/api/v0/computers' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `hostId` (`string `) - Query parameter - `includeProvisioningSteps` (`string`) - Query parameter. Whether to include provisioning step details. Allowed values: true, false. - `includeUnlisted` (`string`) - Query parameter. Whether to include Computers hidden from normal listings. Allowed values: true, false. **Response:** `200` - Response for status 200 ## Create a computer `POST /api/v0/computers` Creates a computer. Managed computers are returned in a provisioning state and set up asynchronously; BYOM computers register an existing machine and are active immediately. ```bash curl -X POST 'https://api.factory.ai/api/v0/computers' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Response:** `201` - Response for status 201 ## Create computers in bulk `POST /api/v0/computers/bulk` Creates up to 20 identical managed computers with server-assigned names (`-`, lowest free suffixes). Creation is fail-fast without rollback: on a mid-batch failure the response is still 201 with the computers created so far plus an `error` describing the failure. ```bash curl -X POST 'https://api.factory.ai/api/v0/computers/bulk' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Response:** `201` - Response for status 201 ## Get a computer by name `GET /api/v0/computers/name/{name}` Returns the authenticated caller's computer with the given name. ```bash curl 'https://api.factory.ai/api/v0/computers/name/{name}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `name` (`string`, required) - Path parameter. Computer name **Response:** `200` - Response for status 200 ## List available computer providers `GET /api/v0/computers/providers` Returns the managed compute providers currently available for new computers in the authenticated organization. ```bash curl 'https://api.factory.ai/api/v0/computers/providers' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Response:** `200` - Response for status 200 ## Get a computer `GET /api/v0/computers/{computerId}` Returns a computer by ID. Callers can access their own computers; organization members with the manager role or higher can also access service-account-owned computers. ```bash curl 'https://api.factory.ai/api/v0/computers/{computerId}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `computerId` (`string`, required) - Path parameter. Computer ID **Response:** `200` - Response for status 200 ## Delete a computer `DELETE /api/v0/computers/{computerId}` Permanently deletes a computer. Deleting a managed computer also shuts down its hosted environment and archives its sessions; deleting a BYOM computer removes the registration but leaves the machine and its sessions intact. Automations running on the computer are paused. ```bash curl -X DELETE 'https://api.factory.ai/api/v0/computers/{computerId}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `computerId` (`string`, required) - Path parameter. Computer ID **Response:** `204` - Response for status 204 ## Update a computer `PATCH /api/v0/computers/{computerId}` Updates a computer's name or remote user. A host ID can be assigned to a computer that does not have one; host IDs are immutable once set and must be unique among the owner's computers. ```bash curl -X PATCH 'https://api.factory.ai/api/v0/computers/{computerId}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Parameters** - `computerId` (`string`, required) - Path parameter. Computer ID **Response:** `200` - Response for status 200 ## Complete a computer file upload `POST /api/v0/computers/{computerId}/files/{fileName}/complete-upload` Completes a file upload by writing the previously uploaded file into the computer's workspace. Returns `409 Conflict` if a file with that name already exists. ```bash curl -X POST 'https://api.factory.ai/api/v0/computers/{computerId}/files/{fileName}/complete-upload' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Parameters** - `computerId` (`string`, required) - Path parameter. Computer ID - `fileName` (`string`, required) - Path parameter. Workspace file name **Response:** `200` - Response for status 200 ## Create a computer file download `POST /api/v0/computers/{computerId}/files/{fileName}/create-download` Creates a download for a file in the computer's workspace and returns a time-limited download URL. ```bash curl -X POST 'https://api.factory.ai/api/v0/computers/{computerId}/files/{fileName}/create-download' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Parameters** - `computerId` (`string`, required) - Path parameter. Computer ID - `fileName` (`string`, required) - Path parameter. Workspace file name **Response:** `200` - Response for status 200 ## Create a computer file upload `POST /api/v0/computers/{computerId}/files/{fileName}/create-upload` Starts a file upload to a computer. Returns a time-limited upload URL and form fields; upload the file there, then call complete-upload to place it in the computer's workspace. ```bash curl -X POST 'https://api.factory.ai/api/v0/computers/{computerId}/files/{fileName}/create-upload' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Parameters** - `computerId` (`string`, required) - Path parameter. Computer ID - `fileName` (`string`, required) - Path parameter. Workspace file name **Response:** `200` - Response for status 200 ## Retry the install-dependencies step `POST /api/v0/computers/{computerId}/install-deps` Starts a new dependency-installation attempt on a managed computer that was configured to install dependencies. Installation runs in the background; returns `202 Accepted` with the updated computer, or `409 Conflict` when an install is already in progress. ```bash curl -X POST 'https://api.factory.ai/api/v0/computers/{computerId}/install-deps' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `computerId` (`string`, required) - Path parameter. Computer ID **Response:** `202` - Response for status 202 ## Get computer metrics `GET /api/v0/computers/{computerId}/metrics` Returns CPU, memory, and disk usage metrics for an active managed computer. Not supported for BYOM computers. ```bash curl 'https://api.factory.ai/api/v0/computers/{computerId}/metrics' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `computerId` (`string`, required) - Path parameter. Computer ID - `start` (`string `) - Query parameter. Start of time range (ISO 8601) **Response:** `200` - Response for status 200 ## Refresh a computer `POST /api/v0/computers/{computerId}/refresh` Refreshes the Git credentials and user-owned computer secrets on a managed computer. BYOM computers are not modified and return zero counts. ```bash curl -X POST 'https://api.factory.ai/api/v0/computers/{computerId}/refresh' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `computerId` (`string`, required) - Path parameter. Computer ID **Response:** `200` - Response for status 200 ## Restart a managed computer `POST /api/v0/computers/{computerId}/restart` Ensures an active managed computer is running, resuming it if necessary. Computers that are already running are left untouched and reported with `wasRestarted: false`. Not supported for BYOM computers. ```bash curl -X POST 'https://api.factory.ai/api/v0/computers/{computerId}/restart' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `computerId` (`string`, required) - Path parameter. Computer ID **Response:** `200` - Response for status 200 # Cloud Templates Deprecated cloud-hosted templates that mirror your local dev setup. **Cloud Templates have been superseded by [Droid Computers](/droid-computers/overview).** Cloud Templates are still supported, but we highly recommend switching for stability. Cloud templates let you code **anywhere** without the "works on my machine" dance. Each template is a pre-configured environment that lives in the cloud, boots in seconds and can be customized to run setup commands. ## Why use cloud templates? | Benefit | What it means for you | | ----------------- | ------------------------------------------------------------------------------------------------------- | | **Zero setup** | Open a session and start coding; no local installs or VM juggling. | | **Consistency** | Every teammate (and CI job) runs the _exact_ same environment. | | **Speed** | Heavy builds run on powerful cloud CPUs; your laptop fan stays silent. | | **Isolation** | Experiments live in disposable templates, keeping your local machine clean. | | **Collaboration** | Share a template link; reviewers jump into the _live_ environment with code and ports already running. | --- ## Installation & usage A **cloud template** is a fully-configured, on-demand development environment that lives in the cloud. Cloud templates give you the same tools and dependencies you'd expect locally, so you can build, test, and run code directly from Factory. To get the most out of cloud templates, configure environment variables and a setup script during template creation. The setup script installs dependencies and prepares your development environment automatically, so every team member gets an identical setup. ### System requirements - A repository enabled in Factory - **User** role or higher to create cloud templates 1. In Factory, click the **Settings** icon from the left sidebar. 2. Select **Cloud Templates**. 1. Click **Create Template**. 2. Enter the repository you want to use. 3. Give your template a friendly name (e.g., "frontend-template"). 4. (Optional) Configure a setup script to run during template initialization. 5. Click **Create**. Factory clones your repo and prepares the environment. This can take a minute for large projects. The new template appears in the list with a status indicator. Once it shows **Ready**, you can use it from any session. ### Launching a cloud template inside a session Join any Factory session as usual. 1. On the session start page, click the Machine Connection button. 2. Choose **Remote** tab. 3. Select the template you created earlier. 4. Factory attaches the cloud template to your session. ![Cloud template attachment flow in the Factory session setup UI](/docs-assets/images/web/machine-connection-start.gif) A green indicator and remote working directory appear on the top-right next to your profile dropdown menu. You're now interacting with the cloud template. ### Everyday usage Use the **Terminal** toolkit to execute commands like:
npm run dev
pytest
git status
Output streams live into chat and logs.
Open files from the repo, make changes, and save. Files persist in the cloud template and can be committed upstream when ready. Auto-save is disabled by default. Enable it from the **Session Settings** panel whenever you want live file syncing.
--- ## Setup script The setup script is a shell script that Factory runs during template creation, after your repository is cloned and before the template is activated. Use this feature to set up your template and give droid tools to work with your codebase. ### How to define a setup script 1. In the modal for template creation, in the "Setup Script (Optional)" section, add your initialization script. You can write a multi-line bash script with all the commands you need. 2. Submit. The script runs in the repo root exactly as provided. Add `set -euo pipefail` at the top of your script if you want strict error handling. Script failures will stop the build. 3. Keep your script non‑interactive and idempotent. Write commands that can be safely re-run. 4. Review build logs if anything fails to see detailed output from your script execution. Examples: **Node.js (Next.js):** ```bash #!/usr/bin/env bash set -euo pipefail npm ci npm run build ``` **PNPM monorepo:** ```bash #!/usr/bin/env bash set -euo pipefail pnpm -w i pnpm -w build ``` **Python:** ```bash #!/usr/bin/env bash set -euo pipefail pip install -r requirements.txt pytest -q ``` **Multi-language project:** ```bash #!/usr/bin/env bash set -euo pipefail # Install Node.js dependencies npm ci # Install Python dependencies pip install -r requirements.txt # Run setup script bash ./scripts/setup.sh ``` What happens under the hood: - The script executes after repository cloning, inside the build container at the repo root. - Environment variables specified in template settings are available during script execution. - Errors are surfaced clearly (e.g., `Setup script failed: ...`) for quick fixes. ### Setup script troubleshooting tips Check the build logs for specific error messages. Run the script locally to debug, add error handling, use non-interactive flags such as `-y`, then retry. Install required tools earlier in your script or ensure they're available in the base Ubuntu image. Make scripts executable (`chmod +x ./scripts/setup.sh`) or invoke via interpreter (`bash ./scripts/setup.sh`). Add it in Environment Variables section and reference it as `$VAR`. Avoid echoing secrets in your script. Keep your script minimal, prefer cached installs (`npm ci` over `npm install`), and avoid heavy, non-essential work. Scripts run at the repo root. Verify relative paths and that files exist after clone. --- ## Best practices Cloud templates let you spin up consistent, production-ready development environments in seconds. Below are field-tested practices that keep templates fast, predictable, and team-friendly. ### Smart setup script practices | Practice | Why it matters | How to do it | | ---------------------------------- | ----------------------------------------------------- | ----------------------------------------------------------------------------------------------------- | | **Order commands by dependency** | Later commands may depend on earlier installs. | Run package installation first: `npm ci && npm run build` or `pip install -r requirements.txt && pytest -q` | | **Use exact package managers** | Consistent lockfiles prevent version drift. | Use `npm ci` (not `npm install`), `pnpm -w i`, or `pip install -r requirements.txt` for reproducible builds | | **Add error handling** | Stops build on first failure, saves debugging time. | Start your script with `#!/usr/bin/env bash` and `set -euo pipefail` for proper error handling | | **Make scripts executable early** | Avoid permission errors mid-build. | Add `chmod +x ./scripts/setup.sh && bash ./scripts/setup.sh` or use `bash ./scripts/setup.sh` directly | | **Keep scripts idempotent** | Re-running setup shouldn't break things. | Use flags like `pip install --no-deps` or check for existing files before creating them | | **Minimize heavy operations** | Long builds slow down template creation. | Focus on essential setup; defer optional tools to manual installation later | > **Tip:** Test your setup script locally first. The script runs with `bash` at the repo root, and you can add `set -euo pipefail` for strict error handling. ### Workflow patterns that scale | Pattern | How to use it | Benefit | | --- | --- | --- | | **Spin-Up-Per-Task** | Treat remote sessions as disposable: create one per ticket or PR, then archive when merged. | Perfect isolation, zero "works on my machine" drift. | | **Parallel Environments** | Launch two separate sessions when you need to test multiple branches. | Switch context without killing processes. | ### Team collaboration tips | Tip | Details | | --------------------------- | ------------------------------------------------------------------------------------------------------------------- | | **Name templates clearly** | Name templates according to the tracked repository, e.g. `repo-name` to work on a repository named `repo-name`. | | **Document entry commands** | Add an `AGENTS.md` file with common tasks (`npm run dev`, `pytest`). Droid automatically reads this file. | --- ## Troubleshooting Even the smoothest cloud template can hit a snag. This section walks you through the quickest fixes for the most common cloud template issues. Diagnose clone failures from the error toast and Git exit code. For private repositories, verify the repo is enabled and displays **Connected** in Integrations, then refresh your OAuth token if prompted. In your session, click the **CPU** icon and ensure **Cloud Machine** is selected. If it shows **Local Machine**, switch to **Cloud Machine** and pick the template. Use Chrome or Edge, disable VPN or proxy tools that may block WebSocket upgrades, then reload the session tab (⌘R / Ctrl-R). Add `.dockerignore` entries for `node_modules`, `*.pyc`, and build outputs. Use a lighter base image such as Alpine when possible. Close unused browser tabs with heavy JavaScript, check local bandwidth (more than 5 Mbps recommended), and pause real-time spell-checker extensions. Disable file watchers in dev tools such as `nodemon` or `webpack --watch` unless needed, and use auto-save only when collaborating. Make sure the repo integration is connected; cloud templates inject HTTPS tokens automatically. Clear package caches (`npm cache clean --force`, `pip cache purge`), remove large build artifacts, or rebuild the template. The file is outside the repo template. Save inside `/workspaces//`. Use `npm install -g` or `pip install --user`. If still blocked, add commands to `postCreateCommand`. Anything outside `/workspaces` is cleared during rebuild. Commit important scripts or store them inside the repo. A quick rebuild fixes many cloud template issues. Click **Rebuild** when in doubt. Migrate to the current persistent compute environment. Register your own machine as an alternative to cloud templates. # IDE Integrations Run Droid from VS Code, Cursor, Windsurf, JetBrains IDEs, Zed, or any editor terminal. Droid works in any editor with an integrated terminal. For deeper editor context, use the dedicated integrations for VS Code-family editors, JetBrains IDEs, or Zed. ## Supported editors | Editor | Setup | Context and capabilities | | ------ | ----- | ------------------------ | | **VS Code, Cursor, Windsurf** | Run `droid` in the integrated terminal or install the [VS Code extension](https://marketplace.visualstudio.com/items?itemName=Factory.factory-vscode-extension). | Shares active file, selection, open files, diagnostics, and IDE-native diff/file actions. | | **JetBrains IDEs** | Install **Factory Droid** from JetBrains AI Agents or configure ACP manually. | Runs Droid inside JetBrains AI Chat with model and autonomy controls. | | **Zed** | Install the Factory Droid extension or configure a custom ACP agent. | Runs Droid in Zed Agent Panel, supports `@`-tagged file context, and can use Zed MCP context servers. | | **Other editors** | Run `droid` from the integrated terminal. | Uses normal CLI behavior. Reference files by path or paste context manually. | ## VS Code, Cursor, and Windsurf Open your editor's integrated terminal and run: ```bash droid ``` Droid detects VS Code-family terminals and can install or update the Factory extension automatically. You can also install it directly from the Marketplace. The extension gives Droid: - Active file path and selection. - Open file list. - Diagnostics from the editor. - IDE-native diff viewing. - File-opening actions from Droid responses. Use `/ide` inside Droid to inspect, install, update, or disconnect the integration: ```text /ide ``` ## JetBrains IDEs JetBrains IDEs use the Agent Client Protocol (ACP). Use this path for IntelliJ IDEA, PyCharm, WebStorm, Android Studio, GoLand, PhpStorm, RubyMine, DataGrip, Rider, and other JetBrains IDEs. ### Install from JetBrains AI Agents In any JetBrains IDE (2025.3+) with JetBrains AI: 1. Open **Settings → Tools → AI Assistant → Agents**, or choose **Install From ACP Registry...** from the agent picker. 2. Find **Factory Droid** and click **Install**. 3. Open the AI Chat panel. 4. Select **Factory Droid** from the agent dropdown. 5. If prompted, complete the browser login flow and confirm the device code. ### Configure manually Install the Droid CLI first with curl -fsSL https://app.factory.ai/cli | sh. Then edit `~/.jetbrains/acp.json`: ```json { "agent_servers": { "Factory Droid": { "command": "/path/to/droid", "args": ["exec", "--output-format", "acp"] } } } ``` To authenticate with an API key instead of browser login, add `FACTORY_API_KEY`: ```json { "agent_servers": { "Factory Droid": { "command": "/path/to/droid", "args": ["exec", "--output-format", "acp"], "env": { "FACTORY_API_KEY": "your_api_key" } } } } ``` Account creation, billing, and API key management happen in the Factory App, not inside JetBrains. After setup, open **View → Tool Windows → AI Assistant**, start a new chat, and choose **Factory Droid**. JetBrains handles session history from the AI Chat panel. The JetBrains integration is currently a rich ACP chat front-end; it does not yet share open files, selections, diagnostics, or IDE-native Droid diffs automatically. ## Zed Zed also uses ACP. The quickest setup is to install the Factory Droid extension, open the Agent Panel, click **+**, and choose **Factory Droid**. To configure Droid manually, install the CLI, then edit `~/.config/zed/settings.json`: ```json { "agent_servers": { "Factory Droid": { "type": "custom", "command": "/path/to/droid", "args": ["exec", "--output-format", "acp"] } } } ``` For API key authentication: ```json { "agent_servers": { "Factory Droid": { "type": "custom", "command": "/path/to/droid", "args": ["exec", "--output-format", "acp"], "env": { "FACTORY_API_KEY": "$FACTORY_API_KEY" } } } } ``` Open the Agent Panel with on macOS or on Linux and Windows, then start a new chat with Factory Droid. Zed does not restore past Factory Droid sessions from the Agent Panel. For longer work, keep the panel open or start a new chat with a short recap. Use Zed's `@`-tagging to add file context: ```text Refactor the state management in @src/components/TodoList.tsx to use a reducer instead of multiple useState hooks. ``` ### Add MCP servers in Zed Zed's `context_servers` can expose MCP tools to Droid. For example: ```json { "context_servers": { "chrome-devtools": { "command": "npx", "args": ["-y", "chrome-devtools@latest"] } } } ``` ## Settings Set `ideAutoConnect` when you want Droid to reconnect to the most recent compatible IDE even from an external terminal: ```json { "ideAutoConnect": true } ``` When `ideAutoConnect` is off, Droid only auto-connects from a detected IDE terminal. See [Settings](/droid-cli/settings) for the full configuration reference. ## Troubleshooting - Run `droid` from the editor's integrated terminal. - Make sure the matching shell command is available: `code` for VS Code, `cursor` for Cursor, or `surf` / `windsurf` for Windsurf. - If the command is missing, open the Command Palette and run the editor's **Shell Command: Install ... command in PATH** action. - Confirm the editor is allowed to install extensions. - Run `droid exec --output-format acp` in a regular terminal to confirm the CLI and authentication work. - Check that the `command` path points to the Droid binary. - Check that `args` includes `exec` and `--output-format acp`. - If using API key auth, confirm `FACTORY_API_KEY` is available to the IDE process. - Windows on ARM is not supported. In JetBrains settings, go to **Tools → Terminal**. Disable **Move focus to the editor with Escape**, or remove the **Switch focus to Editor** terminal keybinding. Verify the VS Code extension is installed, then run `/ide` to inspect connection state. Save files first. Unsaved buffers over 500 KB are skipped for performance. Refresh diagnostics from the VS Code Command Palette. Check shell startup scripts and make sure they do not auto-exit. Configure proxy settings or set `HTTP_PROXY` / `HTTPS_PROXY`. Run the same `command` and `args` outside Zed to debug environment or dependency issues. Contact [support@factory.ai](mailto:support@factory.ai) with logs from `~/.factory/logs/`. When editor context should feed a reviewed implementation plan. Expose editor-hosted context servers and other external systems. # Software Factory Use Software Factory to manage the automations that cover your software delivery lifecycle. Software Factory is a 24/7 autonomous system that connects AI agents across the software development lifecycle, from intake and triage through planning, execution, validation, shipping, monitoring, and recurring automation. Software Factory is in Private Preview. To request access, fill out the form at factory.ai/contact or contact your Factory account team. ![Software Factory overview dashboard](/docs-assets/images/software-factory-dashboard.webp) *Illustrative dashboard values; your numbers reflect your own connected automations.* ## What Software Factory is In the Factory App, Software Factory gives teams one place to track that system: live metrics, connected integrations, repository coverage, stage health, and persistent automations. | Stage | What it covers | | --- | --- | | **Triage** | Classify inbound work, reduce duplicates, and route items. | | **Code-gen** | Connect requests from Slack or issue trackers to implementation sessions. | | **Validate** | Run code review, security review, and QA before merge. | | **Release** | Track deployment gates, release checks, and ship workflows. | | **Document** | Keep repository knowledge current with AutoWiki. | | **Monitor** | Connect incidents, alerts, and agent effectiveness back to engineering work. | ### When to use Software Factory Use Software Factory as a centralized view into product health across your software delivery lifecycle: - Which SDLC stages already have automation coverage. - Which repositories, integrations, or stages need setup. - Whether review, QA, docs, triage, and incident workflows are healthy. - How much work Factory processed over time. ## Manage coverage across the SDLC Open Software Factory to see your delivery lifecycle as an automation coverage map. Each stage shows what is already connected, what is healthy, and where to add coverage next. Use the dashboard to: - Switch between different time ranges. - Scan impact metrics and stage health. - Filter by integration, including GitHub, Slack, Linear, and GitLab. - Open a stage to review templates, automations, and covered repositories. - Create missing automations from the stage detail page or the [Automations](/software-factory/automations) page. Repo-backed stages show repository coverage, so you can see which repos are covered by workflows like **Code Review**, **QA**, **Security Audit**, and **AutoWiki**. | Stage | Templates and setup | Coverage shown | | --- | --- | --- | | **Triage** | Triage | Intake automation coverage | | **Code-gen** | Slack and Linear delegation | Integration coverage | | **Validate** | Code Review, QA, Security Audit | Repository coverage | | **Release** | Release checks and deployment gates | Release status and gate coverage | | **Document** | AutoWiki | Repository coverage | | **Monitor** | Incident Response, Agent Effectiveness | Incident and efficiency coverage | ## Metrics and coverage ### Automation metrics Software Factory gives you a live view of automation impact across the SDLC with metrics such as: ![Software Factory automation metrics](/docs-assets/images/software-factory-automation-metrics.png) | Metric | Dashboard tag | What it indicates | | --- | --- | --- | | **Tickets Triaged** | Queue burn | Tickets classified, deduplicated, or routed by triage workflows. | | **PR Validations** | Merge gate | QA, security, and code-review checks run against pull requests. | | **PRs Merged** | Ship rate | Pull requests merged in the selected time range. | | **Incidents Processed** | Reliability | Incidents, errors, or alerts handled by incident-response workflows. | Use these metrics to see where automation is moving work forward and where your team may need more coverage. ### Stage metrics Use stage metrics to spot which parts of the Software Factory are moving work, waiting for setup, or blocked by handoffs. | Stage | Example metric | What it answers | | --- | --- | --- | | Triage | Throughput and backlog | Are incoming requests being classified and routed quickly? | | Code-gen | Lines changed and session output | How much implementation work is flowing through Droid? | | Validate | Pass rate and review volume | Are QA, security, and code-review gates healthy? | | Release | Deployments and cycle time | Is validated work reaching users on schedule? | | Document | Docs updated and repository coverage | Is AutoWiki keeping knowledge current? | | Monitor | Incident response and token efficiency | Are incidents and agent effectiveness improving? | ### Health and coverage Health and coverage show where Software Factory is ready to run and where setup is missing. | Signal | Meaning | | --- | --- | | Stage health | Whether the stage has the integrations and automations it needs. | | Repository coverage | Which repositories are covered for Validate, Release, and Document workflows. | | Integration coverage | Which intake systems, channels, and deployment tools are connected. | | Setup gaps | The repos, channels, or gates that still need an automation or integration. | {/* sweep-allow: term-bullets */} - **Stage metrics** show throughput, pass rate, queue depth, and cycle time for each SDLC stage. - **Health** shows whether a stage has the integrations and automations it needs. - **Coverage** shows where those automations apply, such as which repositories are covered for **Validate**, **Release**, or **Document**. If something is missing, open the stage to connect the integration, add the repo or channel, or create the expected automation. Create and manage automations that run across your SDLC stages. Classify and route inbound work into the right workflow. # Triage Classify inbound work, reduce duplicates, and route items to the right owner or workflow. Triage is the intake stage of the [Software Factory](/software-factory/overview). It turns incoming issues, threads, and alerts into routed engineering work with enough context for a Droid or teammate to act. Use Triage when the team needs a repeatable path from "someone reported something" to "the right workflow owns it." Triage is in Private Preview. To request access, fill out the form at factory.ai/contact or contact your Factory account team. ## What Triage does Triage connects intake signals to Software Factory stages: 1. **Ingest** work from connected systems such as [Linear](/software-factory/linear), [Slack](/software-factory/slack), GitHub, and support channels. 2. **Classify** the request by type, priority, repository, owning team, and likely workflow. 3. **Reduce duplicates** by comparing new work against recent issues, related threads, and open pull requests. 4. **Route** the item to a teammate, queue, Droid session, or [Automation](/software-factory/automations). 5. **Measure** throughput, backlog, and cycle time in the [Software Factory dashboard](/software-factory/overview). ## When to use it Classify and route Linear issues before they become stale backlog. Turn team threads into tracked engineering work with the right context attached. Group duplicate reports, identify the likely owning surface, and start remediation. Send repeatable triage decisions into scheduled or trigger-based workflows. ## Triage metrics The Software Factory dashboard tracks both intake volume and routing health. | Metric | What it indicates | | --- | --- | | Throughput | How many requests Triage classified in the selected time range. | | Backlog | Items still waiting for an owner, duplicate decision, or next workflow. | | Cycle time | Time from intake signal to routing decision. | | Health | Whether the connected intake sources and automations are ready to run. | ## What to connect first Start with the systems where work already arrives: | Source | What Triage reads | Common route | | --- | --- | --- | | Linear | Issue title, description, comments, labels, project, and status | Owning team, implementation session, or duplicate closure | | Slack | Channel, thread, links, and requester context | Incident response, bug investigation, or follow-up issue | | GitHub | Repository, issue, pull request, and CI context | Code review, QA, security review, or AutoWiki | After the first source is connected, add a narrow automation before broadening coverage. A good first workflow is "classify new bugs in one team queue and route only high-confidence matches." Turn Slack threads into tracked engineering work. Route triage decisions into automations. # Linear Guide to connect Factory with your Linear workspace Connect Factory to Linear so Droid can reference Linear project context during development workflows. ## Prerequisites - Admin access to your Linear workspace - A Factory account with the **Manager** or **Owner** role ## Integration steps Log in to your Factory account and navigate to the Integrations section in your Settings. Locate and select the Linear integration option. You'll be redirected to Linear. Review the requested permissions and click "Authorize" to allow Factory access to your Linear workspace. Select the teams and projects you want Factory to interact with. You can adjust these settings later if needed. Ensure that Factory has the necessary permissions to view and create issues, manage workflows, and access relevant project data. After granting permissions, you'll be redirected back to Factory. Verify that the integration status shows as "Connected". Create a new session directly from Linear with this ticket in context by clicking the **Open in Factory** link on the **Factory** attachment that appears on the issue. - Note that clicking this link will always create a **new session**. ## Verification To ensure the integration is working correctly: 1. Create a test issue in one of the authorized Linear projects. 2. Use Factory to interact with this issue (e.g., comment on it or change its status). 3. Verify that the changes are reflected in Linear. ## Best practices - Regularly review and audit the permissions granted to Factory. - Use Linear's team-level settings to manage access efficiently. - Keep your Linear workspace's security settings up-to-date. ## Troubleshooting Ensure you have the **Manager** or **Owner** role in Factory and admin rights in Linear. Check Linear's audit logs for any permission-related issues. Verify that the integration has access to the intended teams and projects. Contact Factory support with specific error messages or screenshots. Visit Factory's Trust Center for compliance documents, certifications, and security resources Classify and route Linear issues before they go stale. Automate Linear-driven workflows across your team. # Slack Step-by-step guide to connect Factory with your Slack workspace Connect Factory to Slack so Droid can work with Slack conversations and team context. If the Factory Slack app is already installed in your workspace, rerun the Slack install flow below to enable Droid to ingest images from Slack threads and post videos, generated artifacts, and other files back to Slack. An admin may have to approve this depending on your Slack workspace settings. ## Prerequisites - A Factory account with the **Manager** or **Owner** role - Admin access to your Slack workspace - Ability to install apps in your Slack workspace ## Integration steps Log in to your Factory account, open Settings → Organization, then find **Slack** under **Org Integrations**. ![Factory Organization settings page showing Slack under Org Integrations](/docs-assets/images/slack_org_settings.png) Click **Connect**, or **Manage** if Slack is already installed, to start the integration process. If you are updating an existing Slack install for the latest permissions, click **Reconnect** from the integration details to rerun the install flow. ![Factory Slack integration details showing the Reconnect button](/docs-assets/images/slack_manage_settings.png) You'll be redirected to Slack. Review the requested permissions and click "Allow" to authorize Factory access to your Slack workspace. If you belong to multiple workspaces, select the workspace you want to connect to Factory. After authorization, you'll be redirected back to Factory. Verify that the integration status shows as "Connected". In your Slack workspace, add the Factory app to relevant channels by typing `/invite @Factory` in each channel. ## Verification To ensure the integration is working correctly: 1. Mention `@Factory` in a thread within a channel where the app has been added. 2. Verify that Factory responds with a link to open the conversation in Factory. 3. Click the link and confirm that the Slack thread content appears in your new Factory session. 4. If you reran the install flow for the latest permissions, test with a thread that includes an image attachment and confirm Droid can use it as context. 5. If your workflow posts results back to Slack, confirm Droid can upload a generated file or short result video to the thread. ## Capabilities With the Slack integration, you can: - Mention `@Factory` in any Slack thread to start a Droid session from that thread - Continue the session on Web or Desktop - Import full Slack thread context into Factory by clicking the provided link - Include supported image attachments from Slack as session context - Have Droid post messages, generated files, artifacts, and result videos back to Slack - Send follow-up messages and attachments from Slack into an active Droid session - **Use any model you want, no lock-in.** Factory is model-agnostic: choose Claude, GPT, Gemini, open-source/Droid Core, or any of 30+ supported models - Pick the model when you start a session, set a default **model** per channel for Auto-Run and service-account workflows, or let Factory Router pick the best model automatically - Reference Slack threads in Factory by pasting a thread URL into a Factory chat ![Factory Session Settings dialog showing the model picker with Factory models](/docs-assets/images/slack-model-selection.png) When creating a PR, the Slack integration will also use a subagent to handle CI so the agent can respond faster. When a Slack thread is imported, Factory has access to the entire conversation history and uses it as context. ## Slack settings After Slack is connected, click **Manage** from Settings → Organization to configure Slack channel settings. These settings are only needed for Incident Response and service-account workflows; standard `@Factory` mentions work without configuring channels here. In channel settings, you can enable a channel, expand it, and configure **Run as**, **Incident Response**, **Machine Type**, **Computer** or **Workspace**, **Session Visibility**, **Model**, and **Custom Prompt**. Only channels where the Factory app has been invited appear in this list. ## Service accounts Use [service accounts](/enterprise/identity-and-access#service-accounts) when Slack sessions should run from a shared identity and a preconfigured Droid Computer instead of the Slack user who mentions Factory. This is useful for incident, operations, or release channels that need consistent credentials and tool access. Before using a service account with Slack, make sure: - Service accounts are enabled for your organization. - The service account is active. - The service account owns at least one Droid Computer. - The Factory app has been invited to the Slack channel. In Factory, open **Settings → Organization**, find **Slack** under **Org Integrations**, and click **Manage** to open channel settings. Enable the channel and expand its settings. If the **Run as** row appears, select the service account. Selecting a service account makes the channel run with that service account's identity and computers. Service account Slack sessions run on Droid Computers. When a service account is selected, Factory locks the target to **Computer** and only shows computers owned by that service account. For automated channel workflows, turn on **Auto-Run** and configure the prompt, model, and visibility settings for that channel. Mention `@Factory` in the channel, or send a test alert if Auto-Run is enabled. Factory starts the session as the selected service account. For ad hoc Slack sessions, the computer picker can also show a **Run as** selector. Choose a service account, pick one of its computers, and optionally save it as your default for future Slack mentions. ## Best practices - Add the Factory app only to channels where development discussions occur. - Use threads rather than channel messages when mentioning Factory. - Provide sufficient context in the Slack thread before mentioning Factory. - Regularly review the permissions granted to the Factory app in your Slack settings. ## Troubleshooting Ensure you have the **Manager** or **Owner** role in Factory and admin rights in your Slack workspace. Verify that the Factory app has been added to the channel where you're mentioning it. Rerun the Slack install flow above to refresh the Factory Slack app's permissions. The Factory bot needs to be invited to that channel. Use `/invite @Factory` in the channel to resolve the issue. Check that your organization's firewall isn't blocking webhook communications. Contact Factory support with specific error messages. Visit Factory's Trust Center for compliance documents, certifications, and security resources Review Factory's Privacy Policy and Terms of Service. Route Slack requests into tracked engineering work. Set up service accounts for shared Slack channel runs. # Local Code Review Use the /review command to analyze local code changes with AI-powered review workflows ## Overview The `/review` command provides a local workflow for analyzing code changes with AI-powered insights. It offers multiple review modes to fit different development scenarios, from reviewing uncommitted changes to analyzing full branches or specific commits. When you run `/review`, droid guides you through selecting a review type, configuring parameters, and then performs a comprehensive analysis of your code changes based on industry-standard review guidelines. ## Quick start In a local droid session, type: ```bash /review ``` Choose from four review presets: - **Review against a base branch** - PR-style review comparing your branch to a base branch - **Review uncommitted changes** - Analyze working directory changes (staged, unstaged, and untracked) - **Review a commit** - Examine a specific commit from history - **Custom review instructions** - Define your own review criteria Depending on your selection: - For **base branch reviews**: Select the target branch (e.g., `main`, `develop`) - For **commit reviews**: Choose a commit from the interactive list - For **custom reviews**: Enter your review instructions Droid analyzes the code changes and provides: - Prioritized findings with severity levels [P0-P3] - Specific file locations and line numbers - Suggested fixes with code suggestions - Overall assessment of the changes ## Review types ### Review against a base branch Compare your current branch against a base branch (like a pull request review). This is ideal for pre-PR reviews or checking what changes would be merged. **How it works:** 1. Select "Review against a base branch" 2. Choose your target base branch from the list (local and remote branches shown) 3. Droid finds the merge base and reviews the diff **Use cases:** - Pre-commit PR reviews - Checking branch changes before creating a pull request - Validating feature branch against main/develop **Example workflow:** ```text > /review # Select: Review against a base branch # Choose: origin/main # Droid reviews all changes that would be merged ``` ### Review uncommitted changes Analyze all current working directory changes: staged files, unstaged modifications, and untracked files. **Use cases:** - Quick sanity check before committing - Reviewing work in progress - Finding issues early in development **Example workflow:** ```text > /review # Select: Review uncommitted changes # Droid immediately reviews all working directory changes ``` ### Review a commit Examine the changes introduced by a specific commit in your repository history. **How it works:** 1. Select "Review a commit" 2. Browse commits with hash, message, author, and date 3. Select a commit to review **Use cases:** - Reviewing recent commits for issues - Understanding changes in a specific commit - Post-merge review of teammate's work **Example workflow:** ```text > /review # Select: Review a commit # Browse and select: abc1234 - "Add user authentication" # Droid reviews that specific commit's changes ``` ### Custom review instructions Define your own review criteria for specialized analysis. **Use cases:** - Performance analysis - Checking specific coding standards - Domain-specific validations **Example workflow:** ```text > /review # Select: Custom review instructions # Enter: "Focus on performance regressions and unnecessary re-renders" # Droid performs a targeted review ``` ## Review guidelines All code reviews follow a structured rubric designed to produce actionable, high-quality feedback. ### Severity levels | Priority | Meaning | Typical response | | :------- | :------ | :--------------- | | `[P0]` | Critical issue blocking release or operations. | Fix immediately before merge or deploy. | | `[P1]` | Urgent issue that should be addressed in the next cycle. | Fix before the work is considered complete. | | `[P2]` | Normal priority issue to fix eventually. | Track or fix when practical. | | `[P3]` | Low-priority nice-to-have improvement. | Consider if it aligns with nearby work. | ### Bug detection criteria The AI flags an issue as a bug only when all of these are true: | Criterion | Meaning | | :-------- | :------ | | Meaningful impact | Affects accuracy, performance, security, or maintainability. | | Discrete and actionable | Clear, specific issue with a clear fix. | | Appropriate rigor | Does not demand more rigor than the rest of the codebase. | | Introduced in changes | Added by the reviewed changes, not pre-existing. | | Worth fixing | The author would likely fix it if made aware. | | No unstated assumptions | Based on verifiable facts, not speculation. | | Provably affected | Identifies specific affected code, not a theoretical risk. | | Not intentional | Clearly not a deliberate design choice. | ### Comment and output standards | Area | Standard | | :--- | :------- | | Reasoning | Explain why the issue matters and which conditions trigger it. | | Severity | Use `[P0]` through `[P3]` to communicate impact accurately. | | Brevity | Keep each finding to one paragraph. | | Code snippets | Limit Markdown code chunks to three lines. | | Tone | Stay matter-of-fact and avoid accusatory language or excessive flattery. | | Finding title | Use a clear title, 80 characters or fewer, in imperative mood. | | Location | Include file paths and line numbers when applicable. | | Suggested fix | Include a concrete replacement only when it is safe and specific. | | Overall assessment | State whether the changes are correct or incorrect, plus a one to three sentence summary. | ## Tips and best practices ### Making the most of reviews - Run reviews frequently during development, not just before commits - Use custom instructions for domain-specific concerns your team cares about - Treat P0/P1 findings seriously, they represent real issues worth addressing - Review the overall assessment for context on whether changes are sound ### Navigating the review UI - Press Esc to go back at any step in the review flow - From the preset selection, Esc closes the review overlay - Use arrow keys to navigate branch/commit lists - Type to filter branches by name in real-time ## Automated reviews in CI The `/review` command is meant for local CLI sessions. For automated PR reviews in CI/CD, use the [Automated Code Review](/software-factory/code-review-ci) workflow, which supports: {/* sweep-allow: term-bullets */} - **Automatic PR reviews** triggered by pull request events or comments - **Review depth** (`deep` or `shallow`) to control thoroughness and cost - **Custom review guidelines** for repository-specific checks Set it up by running `droid` and entering `/install-code-review`, or run a one-off review with: ```bash droid exec "Review the changes in this PR for correctness and performance regressions" ``` See the [Automated Code Review guide](/software-factory/code-review-ci) for full configuration options. Complete command reference including `/review`. Wire Droid review into GitHub Actions for every pull request. # Automated Code Review Set up automated pull request and merge request reviews with Droid Set up automated code review for GitHub or GitLab repositories. Droid analyzes pull requests and merge requests, identifies issues, and posts feedback as inline comments. For GitHub repositories, the setup flow checks the Factory Droid GitHub App installation as part of configuring review automation. Both platforms run on Droid Action, maintained at [`Factory-AI/droid-action`](https://github.com/Factory-AI/droid-action). The GitHub Action declares its inputs in `action.yml`. The GitLab CI/CD component declares its own in `templates/droid-review.yml`, and `docs/gitlab-setup.md` covers the include line GitLab projects need. This is the CI/CD counterpart to [Local Code Review](/software-factory/code-review). Use the `/review` command for on-demand reviews of local changes before you push; use the Droid Review workflow here to review pull requests automatically in GitHub Actions or GitLab CI. Factory Droid bot posting a code review summary with issues found Factory Droid bot posting inline code review comment on specific lines ## Setup Use the `/install-code-review` command to set up automated code review for GitHub or GitLab: Start Droid: ```bash droid ``` Then enter the setup command: ```text > /install-code-review ``` The guided flow will: 1. Detect your SCM platform (GitHub or GitLab) 2. Verify prerequisites (CLI tools, permissions) 3. Walk you through review configuration (depth, security, triggers) 4. Create a PR/MR with the workflow files To wire this up by hand on GitHub, copy `droid.yml` (on-demand `@droid` commands) and `droid-review.yml` (automatic reviews) from `.github/workflows/` in the [action repository](https://github.com/Factory-AI/droid-action). On GitLab, add `factory/droid-review.yml` plus one `include` line in `.gitlab-ci.yml`, both shown in the repository's `gitlab/examples/`. ## How it works Once enabled, the Droid Review workflow: 1. Triggers on pull request events (opened, synchronize, reopened, ready for review) 2. Skips draft PRs to avoid noise during development 3. Fetches the PR diff and existing comments 4. Analyzes code changes for bugs, security issues, and correctness problems 5. Posts inline comments on problematic lines 6. Submits an approval when no issues are found ## Authentication Automated review needs two separate kinds of access: permission to run Droid, and permission to post on your pull requests. You set them up independently. ### Factory API key (run Droid) Droid runs using your Factory API key. Create one in the Factory API keys settings, then add it to your repository or organization as a secret named `FACTORY_API_KEY`. The workflow passes it in like this: ```yaml - uses: Factory-AI/droid-action@main with: factory_api_key: ${{ secrets.FACTORY_API_KEY }} ``` This is required for every run. ### GitHub access (post reviews) To leave comments and approvals on your PRs, Droid needs a GitHub token. There are two ways to provide one: - **Factory Droid GitHub App (default, recommended).** If you don't supply a token, the action securely requests one for the installed Factory Droid GitHub App. For most teams this is all you need: install the app on your repositories from your organization settings and you're done. It requires the `id-token: write` permission so the action can request the token: ```yaml permissions: contents: write pull-requests: write issues: write id-token: write # required for GitHub App auth ``` - **Your own token (override).** If you'd rather use a personal access token or your own GitHub App, for example on GitHub Enterprise or to control which account posts comments, pass it as `github_token`. When set, Droid uses it directly and skips the app. The token needs write access to pull requests and repository contents. ```yaml - uses: Factory-AI/droid-action@main with: factory_api_key: ${{ secrets.FACTORY_API_KEY }} github_token: ${{ secrets.MY_GITHUB_TOKEN }} ``` On GitLab, the same two pieces apply: set `FACTORY_API_KEY` and `GITLAB_TOKEN` as CI/CD variables. The `/install-code-review` flow configures both for you. For the security architecture behind the GitHub App, see [Securing the GitHub Action and CI](/enterprise/compliance-audit-and-monitoring#securing-the-github-action-and-ci). ## Review depth The `review_depth` input controls the thoroughness and cost of each review. You choose the depth during `/install-code-review` setup, or set it directly in your workflow. - **`deep`** (default): Thorough analysis with higher reasoning effort. Catches more subtle bugs but costs more per review. Best for production code and security-sensitive repos. - **`shallow`**: Faster, more cost-effective reviews that cover surface-level issues. Good for high-volume repos, draft PRs, or teams watching spend. ```yaml with: automatic_review: true review_depth: deep # or shallow ``` You can also override the model or reasoning effort directly with `review_model` and `reasoning_effort`, which take precedence over the depth preset. ## Security review Security review is a dedicated workflow for STRIDE, OWASP, OWASP LLM Top 10, and supply-chain analysis. See [Security Review](/software-factory/security-review) for automatic PR security reviews, scheduled scans, and local full-codebase audits with the built-in `security-review` skill. ## What Droid reviews The automated reviewer focuses on clear bugs and issues: - Dead/unreachable code - Broken control flow (missing break, fallthrough bugs) - Async/await mistakes - Null/undefined dereferences - Resource leaks - SQL/XSS injection vulnerabilities - Missing error handling - Off-by-one errors - Race conditions It skips stylistic concerns, minor optimizations, and architectural opinions. ## Customizing the workflow After the workflow is created, you can customize it by editing `.github/workflows/droid-review.yml` in your repository. ### Change the trigger conditions Modify when reviews run: ```yaml on: pull_request: types: [opened, synchronize, reopened, ready_for_review] paths: - 'src/**' # Only review changes in src/ - '!**/*.test.ts' # Skip test files ``` ### Custom review guidelines Add repository-specific review guidelines by creating a `.factory/skills/review-guidelines/SKILL.md` file in your repo: ```markdown Additional checks for this codebase: - React hooks rules violations - Missing TypeScript types on public APIs - Prisma query performance issues ``` These guidelines are automatically picked up and injected into every review run. No workflow changes needed. ### Change the model Set `review_model` to any public model ID from [Models](/models). It overrides whichever model the `review_depth` preset would have chosen: ```yaml with: automatic_review: true review_model: claude-sonnet-4-6 reasoning_effort: high ``` ### Skip certain PRs Add conditions to skip reviews for specific cases: ```yaml jobs: code-review: # Skip bot PRs and PRs with [skip-review] in title if: | github.event.pull_request.draft == false && !contains(github.event.pull_request.user.login, '[bot]') && !contains(github.event.pull_request.title, '[skip-review]') ``` ### Limit comment count Adjust the maximum number of comments in the prompt: ```text Guidelines: - Submit at most 5 comments total, prioritizing the most critical issues ``` ## All workflow inputs | Input | Default | Description | |-------|---------|-------------| | `automatic_review` | `false` | Automatically review PRs without `@droid review` | | `review_depth` | `deep` | Review preset: `deep` (thorough) or `shallow` (fast) | | `review_model` | (from depth) | Override model for code review | | `reasoning_effort` | (from depth) | Override reasoning effort | | `include_suggestions` | `true` | Include code suggestion blocks in comments | These defaults are the GitHub Action's. Security review inputs are documented in [Security Review](/software-factory/security-review#configuration). The remaining inputs, including trigger phrases, bot allowlists, and debugging overrides, are declared in `action.yml`. The GitLab component declares its own set, with different defaults and extra inputs for organization-wide guidelines, in `templates/droid-review.yml`. Add a STRIDE and OWASP security pass to the same pull requests. Security architecture for the GitHub App integration. Running Droid in CI/CD environments. Source for the GitHub Action and the GitLab CI/CD component. # Security Review Run security-focused PR reviews and full-codebase audits with Droid using STRIDE, OWASP, and supply-chain methodology. Droid security review is a dedicated security workflow for finding high-confidence vulnerabilities in pull requests or across an entire repository. It can run locally from the CLI or automatically in GitHub Actions. Review only the pull request diff, trace changed data flows, and post inline security findings with severity and suggested fixes. Audit every source file in the repository, group files for parallel review, and produce a structured report of validated findings. ## Run a full-codebase audit For the most thorough security results, run the audit inside a [Mission](/missions/overview). Missions plan the audit upfront, fan out work across orchestrated agents, and validate findings at each milestone, which produces dramatically deeper coverage than a single-session run. From any Droid session, enter a mission and kick off the security review: ```text /missions /security-review deep audit ``` ### Periodic scan in CI Run the same mission-based audit on a schedule by invoking `droid exec --mission` from a workflow. The audit writes its full output under `~/security-audits/-/` on the runner, so add an `actions/upload-artifact` step to preserve findings after the runner exits: ```yaml on: schedule: - cron: '0 6 * * 1' jobs: audit: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - name: Install Droid CLI run: | curl -fsSL https://app.factory.ai/cli | sh echo "$HOME/.local/bin" >> "$GITHUB_PATH" - name: Run deep security review env: FACTORY_API_KEY: ${{ secrets.FACTORY_API_KEY }} run: | droid exec --mission --auto high -m claude-opus-4-7 \ "/security-review across the entire repository" - name: Upload security review output if: always() uses: actions/upload-artifact@v4 with: name: deep-security-review-${{ github.run_id }} path: ~/security-audits/ if-no-files-found: warn retention-days: 90 ``` ## Run locally on a diff To review the current diff in your working tree or branch from the CLI, run the built-in skill in any Droid session: ```text /security-review local diff ``` When invoked on a diff, Droid traces changed data flows across authentication, authorization, validation, database, network, filesystem, and LLM boundaries, and reports validated findings inline with severity and suggested fixes. ## Run in GitHub CI on pull requests With [Droid Action](https://github.com/Factory-AI/droid-action), comment on a pull request to trigger an on-demand security review: ```text @droid security ``` To run security review automatically on every non-draft PR, add `automatic_security_review: true` to your review workflow: ```yaml - name: Run Droid Auto Review uses: Factory-AI/droid-action@main with: factory_api_key: ${{ secrets.FACTORY_API_KEY }} automatic_review: true automatic_security_review: true ``` When `automatic_review` and `automatic_security_review` are both enabled, Droid runs the security pass alongside the standard code review and includes the security summary in the PR feedback. ## Configuration These are the Droid Action security inputs currently wired for the workflows documented on this page: | Input | Default | Description | | --- | --- | --- | | `automatic_security_review` | `false` | Run security review automatically on PRs without requiring `@droid security`. | | `security_model` | `""` | Override the model used for security review candidate generation and full-repository scans. Falls back to `review_model` if unset. | Severity threshold, blocking, team notification, and scheduled-scan inputs are declared in `action.yml`. ## Methodology Security review uses the built-in `security-review` skill. In PR automation, Droid Action runs a dedicated `security-reviewer` subagent that loads this methodology before reading files, then traces changed data flows across authentication, authorization, validation, database, network, filesystem, and LLM boundaries. The methodology applies multiple security frameworks together: {/* sweep-allow: term-bullets */} - **STRIDE threat modeling**: Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service, and Elevation of Privilege. - **OWASP Top 10:2021**: Broken access control, cryptographic failures, injection, insecure design, misconfiguration, vulnerable components, authentication failures, integrity failures, logging failures, and SSRF. - **OWASP Top 10 for LLM Applications:2025**: prompt injection, sensitive information disclosure, insecure LLM output handling, excessive agency, vector/embedding weaknesses, and other AI-specific risks when the codebase uses LLMs. - **Supply-chain analysis**: dependency manifest and lockfile review, including typosquatting signals, install scripts, overly broad version ranges, and newly published packages. - **Repository threat-model context**: if `.factory/threat-model.md` exists, Droid uses it as the attack-surface map. ### Review pipeline Security review uses a two-pass workflow: 1. **Candidate generation**: Droid reads the diff or codebase, identifies security-relevant areas, traces untrusted input across trust boundaries, and produces candidate vulnerabilities. 2. **Validation**: Droid re-checks each candidate for reachability, exploitability, existing controls, and false positives before reporting it. Findings are reported only when there is a realistic exploit path, such as an injection vulnerability, missing authentication or authorization on a sensitive operation, hardcoded secret, data exposure, unsafe LLM output handling, or risky supply-chain change. ### Severity levels | Severity | Priority | Examples | | --- | --- | --- | | Critical | `P0` | RCE, hardcoded production secret, auth bypass, unauthenticated admin endpoint | | High | `P1` | SQL injection behind auth, stored XSS, sensitive-data IDOR, very new dependency | | Medium | `P2` | CSRF on state-changing operations, information disclosure, prompt injection behind auth | | Low | `P3` | Minor security hardening with a concrete but low-impact exploit path | Plan and orchestrate large multi-step work, including thorough audits. How to invoke and customize skills. Security architecture and controls for the GitHub Action integration. Pair security review with day-to-day correctness and quality checks. # Automated QA Set up end-to-end automated quality assurance that tests your application as a real user would, across web, CLI, and API surfaces, with visual evidence, CI integration, and failure learning built in. Automated testing of your changes is one of the most important things you can do before shipping code. Unit tests and linters catch syntax-level issues, but they can't tell you whether your login flow actually works, whether your CLI renders correctly after a refactor, or whether a new API endpoint returns the right data. The QA skill fills that gap: it tests your application the way a real user would and produces a structured report with visual evidence. Factory ships two built-in skills that work together: - **`/install-qa`**: A one-time setup skill that analyzes your codebase, asks targeted questions, and generates a complete QA skill tailored to your project. - **`/qa`**: The generated skill that runs on every PR. It reads the git diff, identifies affected apps, and executes only the relevant test flows. The install-qa process is thorough and interactive. It performs deep codebase analysis, runs a multi-phase questionnaire, and generates multiple files. Expect it to take some time and to prompt you with questions -- quality assurance is foundational, and we take the time to get it right. ## Quick start Start Droid: ```bash droid ``` Then enter the setup command: ```text > /install-qa ``` The skill walks through three phases: 1. **Deep codebase analysis**: Detects apps, tech stack, auth, environments, feature flags, integrations, CI/CD, and existing tests. 2. **Interactive questionnaire**: Asks about what it couldn't auto-detect (QA target, user personas, critical flows, cleanup strategy). 3. **Skill generation**: Produces an orchestrator, per-app sub-skills, config, report template, and optionally a GitHub Actions workflow. After setup, invoke QA at any time with `/qa` or let it run automatically on PRs via the generated workflow. ## What gets generated ```text .factory/skills/qa/ SKILL.md # Orchestrator: reads diff, routes to relevant sub-skills config.yaml # All env/auth/integration config (single source of truth) REPORT-TEMPLATE.md # Standardized report template .factory/skills/qa-/ SKILL.md # One sub-skill per testable app (e.g., qa-web, qa-cli) ``` The `config.yaml` is **auto-generated** by `/install-qa` based on codebase analysis and your questionnaire answers. Once generated, you can edit it like any other checked-in file. Example: ```yaml project: MyProject environments: development: url: https://dev.example.com production: url: https://example.com restrictions: [read-only only, never create data] default_target: development auth: method: email-password provider: WorkOS personas: - name: admin test_focus: [settings, user-management, billing] - name: viewer test_focus: [dashboards, reports] cannot_do: [edit-settings, manage-users] apps: web: path_patterns: ["apps/web/**"] skill: qa-web test_tool: agent-browser cli: path_patterns: ["apps/cli/**"] skill: qa-cli test_tool: tuistory failure_learning: suggest_in_report ``` ## How QA runs 1. **Load config**: Reads `config.yaml` for environments, personas, and app definitions. 2. **Analyze the diff**: Maps changed files to apps using `path_patterns`. 3. **Scope the test run**: Only runs sub-skills for affected apps. CLI-only changes skip web tests entirely. 4. **Execute test flows**: Runs relevant flows plus generates targeted tests based on the specific diff. 5. **Capture evidence**: Screenshots (web), terminal snapshots (CLI), or API response data. 6. **Generate report**: Structured report with pass/fail/blocked results, posted as a PR comment. ## Testing tools ### Web apps: agent-browser Drives a real browser: navigates pages, fills forms, clicks buttons, captures accessibility tree snapshots and screenshots. ```bash agent-browser open https://dev.example.com/login agent-browser snapshot -i # Discover interactive elements agent-browser fill @e1 "user@example.com" agent-browser fill @e2 "password123" agent-browser click @e3 agent-browser screenshot result.png ``` If ImageMagick is installed, QA can also generate animated GIF diffs showing before/after UI states for visual regression testing. ### CLI/TUI apps: tuistory Launches the app in a virtual terminal, sends keystrokes, and captures the terminal state as text snapshots and PNG screenshots. ```bash tuistory launch "./my-cli" -s qa-test --cols 110 --rows 36 tuistory -s qa-test wait-idle --timeout 8000 tuistory -s qa-test snapshot --trim # Text snapshot (inline in PR comment) tuistory -s qa-test type "/help" tuistory -s qa-test press enter tuistory -s qa-test screenshot --format png -o /tmp/help.png # Visual evidence ``` ### API testing For backend services without a UI, QA uses standard `curl` commands to test endpoints and validate responses. ## CI integration If your project has a `.github/` directory, install-qa will offer to generate a GitHub Actions workflow that: - Triggers on pull requests (and after preview deployments if using Vercel/Netlify) - Installs tools (tuistory, ImageMagick) and runs `droid exec` with the QA skill - Uploads evidence as build artifacts and posts a QA report as a PR comment - Can be configured as a **required** or **optional** check The generated workflow needs GitHub secrets for credentials referenced in `config.yaml`. The install-qa skill will list exactly which secrets to add. ## Failure learning When QA encounters new failure patterns, it can feed that knowledge back: | Strategy | Behavior | | ------------------------------- | ---------------------------------------------------------------- | | **Suggest in report** (default) | Includes copy-paste snippets in the report for manual review. | | **Auto-commit** | Automatically commits updates to sub-skill files after each run. | | **Open a PR** | Opens a draft PR with failure catalog updates. | ## Real-world examples ### CLI app (Go TUI): [glow](https://github.com/nizar-test/glow/pull/1) A PR updating help text and flag descriptions in a terminal markdown renderer. QA built the Go binary and tested it with tuistory: | # | Test Case | Result | Notes | | --- | ------------------------------------- | ----------------------- | -------------------------------------------------- | | 1 | Help text shows updated description | :white_check_mark: PASS | `--help` includes new text | | 2 | Line-numbers flag description updated | :white_check_mark: PASS | Shows "rendered output" instead of "TUI-mode only" | | 3 | CLI renders markdown correctly | :white_check_mark: PASS | Headers, lists, code blocks render | | 4 | Width flag wraps at specified column | :white_check_mark: PASS | `-w 40` wraps correctly | | 5 | Stdin pipe rendering | :white_check_mark: PASS | Piped markdown renders | | 6 | Error on nonexistent file | :white_check_mark: PASS | Exits code 1 with clear message | | 7 | TUI browser launch | :no_entry: BLOCKED | CI PTY environment inconsistent | Evidence included inline terminal snapshots: ```text $ ./glow --help Render markdown on the CLI, with pizzazz! Now with improved word wrapping and line number support. ``` ### Full-stack web app (FastAPI + React): [full-stack-fastapi-template](https://github.com/nizar-test/full-stack-fastapi-template/pull/1) A PR adding a "Remember me" checkbox and footer version badge. QA spun up PostgreSQL, Mailcatcher, FastAPI backend, and React frontend in CI, then drove the UI with agent-browser: | # | Test Case | Result | Notes | | --- | ------------------------------------- | ----------------------- | ------------------------------------------------- | | 1 | Login page shows Remember Me checkbox | :white_check_mark: PASS | Form has Email, Password, checkbox, Log In button | | 2 | Login with Remember Me checked | :white_check_mark: PASS | Redirected to dashboard | | 3 | Login without Remember Me | :white_check_mark: PASS | Also works unchecked | | 4 | Invalid credentials (negative test) | :white_check_mark: PASS | Toast error, stays on login | | 5 | Footer shows v2.0 | :white_check_mark: PASS | Version badge visible on all pages | Here are actual screenshots captured by agent-browser during the QA run: ![Login page with the new 'Remember me for 30 days' checkbox, captured by agent-browser](/docs-assets/images/qa-examples/qa-login-page.png) ![Dashboard after successful login, showing the v2.0 footer badge](/docs-assets/images/qa-examples/qa-dashboard.png) ## Tips {/* sweep-allow: term-bullets */} - **Be detailed during the questionnaire.** The quality of the generated QA skill is directly proportional to the detail you provide. Describe user roles, critical flows, auth mechanics, and edge cases thoroughly. The more context install-qa has, the more targeted the generated test flows will be. - **Describe success criteria clearly.** Don't just say "login works." Say "user enters email and password, clicks Sign In, gets redirected to /dashboard, and sees a welcome message." Specificity produces test flows that verify the right thing. - **Mention known quirks.** If your login form renders differently in certain locales, if a checkout takes 15 seconds, or if your dev server needs a specific start command, say so. These become Known Failure Modes that prevent false failures. - **Iterate after the first run.** Review the report, refine sub-skill flows based on what passed, what was blocked, and what was missed. Schedule QA runs and trigger them from GitHub events. View QA coverage and validation health across repos. # Droid Control Terminal, browser, and desktop automation. Record demos, verify behavior claims, and run QA flows. Droid Control lets Droids *operate* software: launch apps, type commands, click buttons, record what happens, and produce polished video evidence of it. Built by Droids, for Droids. ## What you get Test whether a behavior claim is true and produce evidence either way. No staging, no advocacy, just investigation. Drive terminal CLIs, web apps, or Electron apps through end-to-end flows. Report pass/fail with annotated screenshots. Generate polished before/after comparison videos of PRs, complete with title cards, keystroke overlays, and window chrome. ## Get started Run `/plugins` in a Droid session. If `factory-plugins` is not registered, open **Marketplaces** and add `Factory-AI/factory-plugins`. Then install **droid-control** from **Available**. ```bash droid plugin marketplace add https://github.com/Factory-AI/factory-plugins droid plugin install droid-control@factory-plugins --scope user ``` For video rendering, Remotion dependencies need a one-time install after adding the plugin: ```bash droid plugin list --scope user cd /remotion && npm install ``` You also need the runtime tools for your use case (tuistory, agent-browser, ffmpeg, etc.). See [Prerequisites](#prerequisites) for per-use-case install commands. ## Commands Droid Control adds three slash commands. Each handles the full workflow end-to-end: planning, execution, recording, and reporting. Record a demo video of a feature or PR. ``` /demo pr-1847 ``` Accepts a PR number, GitHub URL, or free-text description. Comparison PRs get side-by-side layout by default; new features get single-branch. Add flags for extra polish: ``` /demo pr-1847 -- showcase, keys ``` | Flag | Effect | |------|--------| | `showcase` | Cinematic preset with warm backgrounds and film grain | | `keys` | Keystroke overlay pills showing user actions | ### How it works Fetches the PR description, diff, and linked ticket. For each change, identifies what needs to be proven and what a viewer could confuse it with. Scripts a sequence of actions that produces visible evidence the feature works. Both branches run identical interactions so only the behavior differs. Presents the plan and waits for your approval before recording. Launches recorded sessions on the baseline and candidate branches in parallel using worker subagents. Renders a polished video via Remotion with title cards, window chrome, and effects. Six visual presets range from cinematic (`factory`) to utilitarian (`minimal`). Checks the final video against the original commitments before delivering. Test a specific behavior claim and report findings with evidence. ``` /verify "ESC cancels streaming in bash mode" ``` Also accepts a PR reference with an optional claim: ``` /verify 11386 -- the fork flag creates a new session ``` If given a PR number alone, Droid fetches the PR and identifies the most important testable claim. The droid is framed as an **investigator**, not an advocate. If the claim is false, that's a valid finding. Anti-fabrication rules prevent staging evidence to match expected outcomes. ### How it works Identifies the specific behavior to observe and what evidence type is needed: text snapshots for functional claims, screenshots for visual claims, or raw byte captures for encoding claims. Launches the app, runs the minimal interaction sequence that demonstrates the behavior, and captures the result. If the behavior contradicts the claim, that is evidence, not an error. Delivers a structured report with a **CONFIRMED**, **REFUTED**, or **INCONCLUSIVE** conclusion, along with all captured evidence inline. Run automated QA against terminal CLIs, web apps, or Electron apps. ``` /qa-test https://app.example.com -- login, create a project, invite a member ``` Also accepts a CLI command, Electron app name, PR reference, or free-text description. Test steps after `--` are optional. Droid designs a reasonable flow if none are provided. ### How it works Determines the target (web, terminal, or Electron), designs test steps from your instructions or the app's UI, and identifies what evidence to capture at each step. Launches the app and executes each step, capturing screenshots (browser) or text snapshots (terminal) along the way. If a step fails, it records the failure and continues for maximum coverage. Delivers a step-level pass/fail table with inline evidence and a summary of any issues found. ### Example output Every video below was planned, recorded, and rendered entirely by a Droid. ## Automation drivers Droid Control supports four automation backends. The right one is selected automatically based on what you're targeting. **Virtual PTY automation.** Default for terminal work. Playwright-style CLI with asciinema recording and forced truecolor output. **Real terminal emulator.** Headless Wayland compositor (Linux), KVM/QEMU (Windows), or QEMU monitor (macOS). For when you need real rendering evidence. **Web and Electron apps.** Playwright-backed CLI with Chrome DevTools Protocol support. Navigates pages, fills forms, clicks buttons, captures screenshots. **Native desktop apps.** Accessibility-tree snapshots, element or pixel actions, and per-window screenshots via [trycua/cua](https://github.com/trycua/cua) `cua-driver`, all without stealing focus. macOS and Windows are production tier; Linux is pre-release. ## Video rendering Demo and showcase videos are rendered with [Remotion](https://www.remotion.dev/), a React-based video engine. The plugin includes 23 visual components and 6 presets. ### Visual presets | Preset | Look | Best for | |--------|------|----------| | `factory` | Warm black, traffic lights, amber glow | Official Factory content | | `factory-hero` | Same + gradient background | Landing pages, social | | `hero` | Cool gradient, generous margins | Non-Factory marketing | | `macos` | Dark, clean frame | General-purpose demos | | `presentation` | Black, generous margins | Slide decks, talks | | `minimal` | No window bar, tight margins | Docs embeds, inline clips | ### Automatic layers - Warm radial backgrounds, floating particles, film grain overlay, color grading - Configurable title-to-content transition (`motion-blur`, `flash`, `whip-pan`, `light-leak`, `glitch-lite`) - Animated window chrome with traffic lights and glassmorphic borders - Auto-scaled title/subtitle text ### Effect layers - Spotlight overlays to highlight specific regions - Directed zoom for small text or details - Keystroke pills showing user actions - Section headers and transition sweeps - Syntax-highlighted code annotations for source-change overlays ## Architecture The plugin uses a composition architecture with three layers: {/* sweep-allow: term-bullets */} - **Orchestrator**: Routes each request through three independent lookups (target, stage, artifact) to determine which skills to load. - **11 atom skills**: Self-contained background knowledge loaded on demand, split into drivers, targets, stages, and polish. - **3 commands**: Parse arguments into commitments, then delegate to atoms via hybrid handoffs. Every workflow flows through **capture → compose → verify**. Commands declare *what* to produce; atoms own *how*. Skills chain through explicit handoffs rather than hardcoded pipelines, so the droid follows the flow naturally. Design rationale: UX for droids, waterfall routing, task delegation, and hybrid handoffs. ## Prerequisites Only install what you need for your use case. ### Terminal demos (tuistory) ```bash npm install -g tuistory # virtual PTY driver pip install asciinema # terminal recording cargo install --git https://github.com/asciinema/agg # .cast → .gif converter sudo apt-get install -y ffmpeg # video processing ``` ### Web/Electron automation (agent-browser) ```bash agent-browser install # downloads Chromium ``` ### Native desktop apps (desktop-control) ```bash curl -fsSL https://raw.githubusercontent.com/trycua/cua/main/libs/cua-driver/scripts/install.sh | bash cua-driver skills install # upstream skill pack (deep tool reference) ``` Windows installs via PowerShell: `irm https://raw.githubusercontent.com/trycua/cua/main/libs/cua-driver/scripts/install.ps1 | iex`. macOS additionally needs `cua-driver permissions grant` (Accessibility + Screen Recording). ### Real terminal emulator (true-input) | Platform | Required tools | |----------|----------------| | Linux/Wayland | `cage`, `wtype`, a Wayland-compatible terminal | | Windows (KVM) | `libvirt`, `qemu`, KVM VM with SSH | | macOS (QEMU) | `qemu`, `socat`, macOS VM with SSH | ### Video composition (showcase) Requires Node.js >= 18, Chrome/Chromium, `ffmpeg`, `ffprobe`, and `agg`. Full plugin source: skills, commands, scripts, and Remotion components. Learn how plugins work, how to install them, and how to build your own. Quick start, command reference, and prerequisites. Persistent environments for running terminal and browser automation. # Release Track deployment gates and ship workflows as they become available. Release is the shipping stage of the [Software Factory](/software-factory/overview). It tracks deployment gates, release workflows, and ship status so teams can see whether validated work is ready to reach users. Use Release when shipping depends on more than a merge button: versioning, rollout checks, approval gates, deployment health, and follow-up documentation. Release is in Private Preview. To request access, fill out the form at factory.ai/contact or contact your Factory account team. ## What Release tracks Release connects the final Software Factory stages: 1. **Readiness** from code review, QA, security review, and CI. 2. **Release preparation** such as changelog updates, version notes, migration checks, and ownership handoffs. 3. **Deployment gates** including rollout state, failed checks, and manual approval points. 4. **Post-ship feedback** from incidents, alerts, customer reports, and follow-up tasks. ## When to use it Keep recurring release steps visible before a ship window starts. Track deployment constraints, environments, and rollout boundaries. Connect review and QA outcomes to release readiness. Move repeatable release checks into persistent automations. ## Release metrics The Software Factory dashboard shows release throughput and gate health alongside the rest of the SDLC. | Metric | What it indicates | | --- | --- | | Deployments | How many release workflows shipped in the selected time range. | | Pass rate | The share of release gates that completed without intervention. | | Cycle time | Time from validated work to shipped release. | | Health | Whether release checks, deployment gates, and handoffs are configured. | *Illustrative dashboard values; your numbers reflect your own connected automations.* ## Release workflow | Step | What Droid can help with | What stays visible | | --- | --- | --- | | Prepare | Draft release notes, summarize merged work, check migration instructions | Owners, linked pull requests, and release scope | | Verify | Run validation, review failures, compare rollout blockers | Gate status, pass rate, and unresolved blockers | | Ship | Execute approved release steps or hand off to the deployment system | Deployment count, rollout state, and cycle time | | Learn | Capture incidents, regressions, and follow-up work | Post-ship signals routed back into Triage or Validate | For controlled headless release workflows. Route post-ship signals back into automated investigation. # AutoWiki AutoWiki generates comprehensive, structured Wikis for your repositories. AutoWiki generates comprehensive, structured Wikis for your repositories. It analyzes each codebase and produces browsable documentation with architecture overviews, module breakdowns, and cross-linked pages, then keeps each Wiki current as your code changes. ![AutoWiki repository overview in the Factory App](/docs-assets/images/wiki/autowiki-list.png) Run `/wiki` in the Droid CLI to produce a Wiki for any repo Run `/install-wiki` to refresh Wikis on every push Read, search, and export Wikis in the Factory App ## How it works Run `/wiki` inside a Droid session. Droid analyzes the codebase and produces a set of structured markdown pages covering architecture, modules, APIs, and conventions. Cloud Sync stores the generated Wiki in the Factory App, where it becomes browsable at app.factory.ai/wiki. Each sync is versioned so you can compare changes over time. For GitHub-hosted repositories, AutoWiki can push generated Wikis to the repo's built-in GitHub Wiki tab. Internal links are rewritten, a sidebar is generated, and the content is pushed to `{repo}.wiki.git`. Run `/install-wiki` to create a CI workflow (GitHub Actions or GitLab CI) that regenerates the Wiki on every push to the default branch, or trigger a refresh from the Factory App at any time. ## GitHub Wiki sync When AutoWiki generates a Wiki for a GitHub-hosted repository, Factory can optionally sync the content to GitHub Wiki. Your team can browse the same documentation directly from the repository page without visiting the Factory App. For details on how the sync works, see [Initialize a Wiki](/software-factory/wiki/generate#github-wiki-sync). ![Generated AutoWiki content synced to GitHub Wiki](/docs-assets/images/wiki/github-wiki-sync.png) GitHub requires GitHub Wiki to be initialized before Factory can sync to it. If you haven't used GitHub Wiki before, create the first page manually at `https://github.com/{owner}/{repo}/wiki`, then re-run `/wiki`. ## Enterprise Controls ### AutoWiki Cloud Sync Organizations that need to prevent AutoWiki content from being stored in the Factory App can disable **AutoWiki Cloud Sync**. This is an org-level enterprise control available in the Factory App under **Settings > Enterprise Controls**. ![AutoWiki Cloud Sync enterprise control in Factory settings](/docs-assets/images/wiki/autowiki-enterprise-controls.webp) When AutoWiki Cloud Sync is **disabled**: - All AutoWiki API endpoints return 403 for your organization - The `/wiki` command will not store content in the Factory App - GitHub Wiki sync is also skipped from the Droid CLI - The AutoWiki page in the Factory App shows no content AutoWiki Cloud Sync is **enabled by default**. Only org managers can toggle this setting. For more on how Enterprise Controls work, see [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control). Generate a Wiki for your repository from the Droid CLI. Browse, search, and export generated Wikis in the app. # Initialize a Wiki Use the /wiki command to generate a comprehensive Wiki with AutoWiki from the Droid CLI. The `/wiki` command analyzes your repository and generates a structured Wiki covering architecture, modules, APIs, and conventions. Cloud Sync stores the Wiki in the Factory App, and AutoWiki can optionally sync it to GitHub Wiki. ## Quick start Navigate to your project directory and start Droid: ```bash cd /path/to/your/project droid ``` ``` > /wiki ``` Droid analyzes your codebase and generates markdown files into a local `droid-wiki/` folder in your project directory. Once generation is complete, Cloud Sync stores the contents in the Factory App. After Cloud Sync completes, Droid prints a link to your Wiki: ``` Wiki uploaded successfully: https://app.factory.ai/wiki/{wikiRunId} ``` Open the link to browse your documentation, or visit app.factory.ai/wiki to see all your Wikis. ## What happens during generation When you run `/wiki`, Droid performs several steps: 1. **Codebase analysis**: Reads source files, configuration, tests, and documentation to understand the project structure 2. **Page generation**: Produces a set of structured markdown files organized by topic (architecture, modules, APIs, setup, etc.) 3. **Visual screenshots**: If the repository has a [QA skill](/software-factory/automated-qa) installed, Droid launches the application and captures screenshots of web UIs and TUIs to include in the Wiki. Run `/install-qa` to set one up if your repo doesn't have one yet. 4. **Cloud Sync**: Stores the pages and images in the Factory App 5. **GitHub Wiki sync**: For GitHub repos, flattens pages and pushes to GitHub Wiki Wiki generation uses your current branch and working directory state. For the most accurate documentation, run `/wiki` from a clean checkout of your default branch. The first Wiki generation for a repository can take a while because Droid needs to analyze the entire codebase from scratch. Subsequent runs are significantly faster because they use incremental mode, only updating pages affected by code changes since the last generation. ## GitHub Wiki sync After Cloud Sync stores the Wiki in the Factory App, the Droid CLI automatically syncs it to GitHub Wiki for GitHub-hosted repositories. The sync: ![Generated AutoWiki content synced to GitHub Wiki](/docs-assets/images/wiki/github-wiki-sync.png) - Flattens the hierarchical page tree into GitHub Wiki's flat format (e.g., `overview/architecture.md` becomes `overview--architecture.md`) - Rewrites internal links to use the flat filenames - Generates `_Sidebar.md` with a navigable table of contents - Generates `Home.md` from your Wiki root page - Clones `{repo}.wiki.git`, replaces all content, and pushes ### Prerequisites for GitHub sync - Your repository must be hosted on GitHub - GitHub Wiki must be initialized. If you've never used it, create the first page at `https://github.com/{owner}/{repo}/wiki` - Your Git credentials must have push access to the GitHub Wiki repository - [AutoWiki Cloud Sync](/software-factory/wiki/overview#autowiki-cloud-sync) must not be disabled for your organization If GitHub Wiki isn't initialized, the Droid CLI prints a message with a link to create the first page and skips the sync without failing. ## Versioning Each Wiki generation creates a new **Wiki Run** with metadata including: - Commit hash and branch name - Whether there were local uncommitted changes - Droid CLI version used - Timestamp You can browse previous versions from the [AutoWiki page in the Factory App](/software-factory/wiki/web-viewer) using the version dropdown. Browse the generated Wiki and compare versions in the app. Enrich Wikis with screenshots and tests of your application. # Wiki Browser Browse, search, and manage Wikis from the Factory App. The AutoWiki view in the Factory App at app.factory.ai/wiki lets you browse, search, and manage generated Wikis for all your repositories in one place. ## What you'll see The top-level Wiki view lists every generated repository Wiki, with freshness indicators and controls for searching, filtering, and refreshing documentation. ![AutoWiki repository list in the Factory App](/docs-assets/images/wiki/autowiki-list.png) Open a repository to read generated pages with navigation, page outlines, search, export, and refresh controls. ![AutoWiki reader showing generated architecture documentation](/docs-assets/images/wiki/autowiki-reader.png) When GitHub Wiki sync is enabled, the same generated pages are also available from the repository's GitHub Wiki tab. ![Generated AutoWiki content synced to GitHub Wiki](/docs-assets/images/wiki/github-wiki-sync.png) ## Browsing your Wikis The main AutoWiki page shows a table of repositories with generated Wikis. Each row displays the repository name, last refresh date, and branch. Click a repository to open its Wiki. ### Wiki reader The Wiki reader provides: {/* sweep-allow: term-bullets */} - **Markdown rendering**: Full-fidelity rendering of the generated documentation with syntax-highlighted code blocks - **Table of contents**: A sidebar outline of the current page for quick navigation - **Breadcrumbs**: Hierarchical navigation showing your position in the page tree - **Search**: Full-text search across all generated pages - **Version history**: A dropdown to browse previous Wiki generations and compare how documentation has evolved - **Export**: Download the generated documentation as markdown files ## Refreshing a Wiki Click the **Refresh** button on a Wiki page to regenerate documentation. A modal presents two options: Regenerate the Wiki once using the latest code on the default branch. Install a CI action that refreshes the Wiki on every push to the default branch. After choosing a mode, select how to run the generation: | Method | Description | |--------|-------------| | **Local** | Copies the Droid CLI command (`/wiki` or `/install-wiki`) for you to paste into a local Droid session | | **Cloud** | Runs the generation on a cloud template linked to the repository | ### Batch refresh From the top-level AutoWiki page at app.factory.ai/wiki, you can refresh multiple repositories at once: 1. Click **Refresh** to enter selection mode 2. Select the repositories you want to refresh 3. Choose **One-time** or **Recurring** 4. Select a **Droid Computer** to run the work ![Batch refresh selection mode in AutoWiki](/docs-assets/images/wiki/autowiki-batch-refresh.png) Droid clones each repository on the selected Droid Computer and generates Wikis, or sets up recurring generation, running the jobs in parallel where possible. Generate or refresh a Wiki for any repository. Learn how AutoWiki keeps documentation current. # Automatic Refresh Use /install-wiki to set up a CI action that regenerates a Wiki on every push to the default branch. The `/install-wiki` command creates a CI workflow that automatically regenerates a Wiki every time code is pushed to the default branch. It detects your CI framework (GitHub Actions or GitLab CI) and generates the appropriate configuration. This keeps your documentation in sync with your codebase without manual intervention. ## Quick start ```bash cd /path/to/your/project droid ``` ``` > /install-wiki ``` Droid creates a workflow file and opens a pull request for you to review. Add `FACTORY_API_KEY` as a secret in your CI settings: - **GitHub**: Repository **Settings > Secrets and variables > Actions** (or at the organization level under **Organization Settings > Secrets and variables > Actions**) - **GitLab**: Repository **Settings > CI/CD > Variables** (or at the group level under **Group Settings > CI/CD > Variables**) Generate a key in the Factory API keys settings. Once the secret is configured, merge the workflow PR. The Wiki will now refresh automatically on every push to the default branch. ## Generated workflows Droid detects your CI framework and creates the appropriate configuration. ### GitHub Actions Creates `.github/workflows/droid-wiki-refresh.yml` (replace `main` with your default branch if different): ```yaml name: Droid AutoWiki Refresh on: push: branches: [main] # change to your default branch if not main jobs: wiki-refresh: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - name: Install Factory Droid run: curl -fsSL https://app.factory.ai/cli | sh - name: Generate Wiki run: droid exec --auto high "/wiki" env: FACTORY_API_KEY: ${{ secrets.FACTORY_API_KEY }} ``` ### GitLab CI Appends a job to your existing `.gitlab-ci.yml`: ```yaml droid-wiki-refresh: stage: deploy before_script: - curl -fsSL https://app.factory.ai/cli | sh script: - droid exec --auto high "/wiki" rules: - if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH variables: FACTORY_API_KEY: $FACTORY_API_KEY ``` Both workflows: {/* sweep-allow: term-bullets */} - **Trigger on push** to the default branch - **Install the Droid CLI** using the official installer - **Run `/wiki`** in headless mode with high autonomy, generating the Wiki and storing it in the Factory App with Cloud Sync You can customize the workflow after installation. For example, change the trigger branch, add path filters to only regenerate when source files change, or adjust the timeout. ## Prerequisites - A GitHub or GitLab hosted repository - A Factory API key (generate one in the Factory API keys settings) - Permission to add secrets and merge PRs in the target repository ## Refreshing from the Factory App You can also set up recurring refreshes from the Factory App: 1. Go to app.factory.ai/wiki 2. Click **Refresh** on the Wiki page 3. Select **Recurring (on push)** 4. Choose a **Local** or **Cloud** method: - **Local**: Copies the `/install-wiki` command for you to run in a local clone - **Cloud**: Selects a cloud template or Droid Computer to run the setup remotely For batch operations across multiple repositories, the Factory App lets you select several repos and run the setup on a Droid Computer in a single operation. See [Wiki Browser](/software-factory/wiki/web-viewer) for more on the Factory App interface. ## Troubleshooting Verify the `FACTORY_API_KEY` secret is set correctly in your repository's **Settings > Secrets and variables > Actions**. Generate a new key at Factory API keys settings if needed. Check that the workflow file exists at `.github/workflows/droid-wiki-refresh.yml` and that the branch trigger matches your default branch. View the Actions tab in your repository to see workflow run logs. The GitHub Wiki tab must be initialized before Factory can push to it. Create the first page manually at `https://github.com/{owner}/{repo}/wiki`. Also verify that your organization hasn't disabled [AutoWiki Cloud Sync](/software-factory/wiki/overview#autowiki-cloud-sync). Headless mode used by the generated workflow. How Factory generates and publishes repository documentation. # AutoWiki API Generate and manage AutoWiki runs: create runs, browse pages, search, and export. ## List the latest wiki runs `GET /api/v0/wiki` ```bash curl 'https://api.factory.ai/api/v0/wiki' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Response:** `200` - Response for status 200 ## Create a wiki run `POST /api/v0/wiki` ```bash curl -X POST 'https://api.factory.ai/api/v0/wiki' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Response:** `200` - Response for status 200 ## List wiki run history for a repository `GET /api/v0/wiki/history/{repoUrl}` ```bash curl 'https://api.factory.ai/api/v0/wiki/history/{repoUrl}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `repoUrl` (`string`, required) - Path parameter **Response:** `200` - Response for status 200 ## Create a wiki video upload URL `POST /api/v0/wiki/video-upload-url` ```bash curl -X POST 'https://api.factory.ai/api/v0/wiki/video-upload-url' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Response:** `200` - Response for status 200 ## Get a wiki run `GET /api/v0/wiki/{wikiRunId}` ```bash curl 'https://api.factory.ai/api/v0/wiki/{wikiRunId}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `wikiRunId` (`string`, required) - Path parameter **Response:** `200` - Response for status 200 ## Delete a wiki run `DELETE /api/v0/wiki/{wikiRunId}` ```bash curl -X DELETE 'https://api.factory.ai/api/v0/wiki/{wikiRunId}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `wikiRunId` (`string`, required) - Path parameter **Response:** `200` - Response for status 200 ## Export a wiki run as a zip `GET /api/v0/wiki/{wikiRunId}/export` ```bash curl 'https://api.factory.ai/api/v0/wiki/{wikiRunId}/export' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `wikiRunId` (`string`, required) - Path parameter **Response:** `200` - Response for status 200 ## Get a wiki page `GET /api/v0/wiki/{wikiRunId}/pages/{pageId}` ```bash curl 'https://api.factory.ai/api/v0/wiki/{wikiRunId}/pages/{pageId}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `wikiRunId` (`string`, required) - Path parameter - `pageId` (`string`, required) - Path parameter **Response:** `200` - Response for status 200 ## Update wiki run privacy level `POST /api/v0/wiki/{wikiRunId}/privacy` ```bash curl -X POST 'https://api.factory.ai/api/v0/wiki/{wikiRunId}/privacy' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `wikiRunId` (`string`, required) - Path parameter **Response:** `200` - Response for status 200 ## Search a wiki run `GET /api/v0/wiki/{wikiRunId}/search` ```bash curl 'https://api.factory.ai/api/v0/wiki/{wikiRunId}/search' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `wikiRunId` (`string`, required) - Path parameter **Response:** `200` - Response for status 200 # Incident Response Set up Droid to respond to production alerts in Slack and resolve or investigate incidents autonomously. Incident Response is in Private Preview. To request access, fill out the form at factory.ai/contact or contact your Factory account team. ## Overview When an alert lands in a configured Slack channel, an Incident Response automation automatically starts a linked Droid session so Droid can investigate the incident with the right context, machine, and instructions. Droid can work toward an RCA, open a PR for a fix when appropriate, and update an incident runbook so future investigations of the same alert type start faster. ## Quickstart: set up from the Factory App Incident Response is set up as an [automation](/software-factory/automations) in the Factory App. In Slack, open the channel where alert bots post incidents and run: ```text /invite @Factory ``` If you have not connected Slack yet, set up the [Slack integration](/software-factory/slack) first. Only channels where Factory has been invited appear in the channel picker. In the Factory App, open All Automations under **Software Factory** in the sidebar and click **New automation**. From the template gallery, select the **Incident Response** template under Software Factory templates. ![Automation template gallery with the Incident Response template](/docs-assets/images/incident-response-templates.png) Give the automation a name, then under **Triggers**, search for and select the incident channel. **Trigger on messages from** defaults to **Bots only**, which matches alert bots posting incidents; change it only if humans should also trigger investigations. ![New Incident Response automation form](/docs-assets/images/incident-response-create.png) - **Run as** sets the identity and billing the automation runs under. A [service account](/enterprise/identity-and-access#service-accounts) is recommended for incident response so the computer and session interactivity are shared across the team. - **Run on** sets the computer where Droid investigates. Pick the computer with the right repositories, observability tools, and anything else you want to add that can help the agent. Learn more about [Droid Computers](/droid-computers/overview). Use **Configure MCP Servers** to add and authenticate the observability tools and context sources Droid needs to respond to incidents. MCP servers are configured per computer and **require connection to the computer to set up**; if you have issues, retry a few times or change the computer. The template pre-fills a prompt that is injected into every session triggered from the channel: ```text Use the `/incident` skill for Slack incidents and run an RCA if applicable. If the thread message is an incident or asking for RCA/root-cause analysis, invoke the `incident` skill before investigating. ``` Keep the default or add channel-specific instructions, and optionally select a model. Choose whether triggered sessions are visible to your **Team** or **Private**, review any **Additional Settings**, then click **Create**. Post a test top-level alert from the same kind of bot that will send real incidents. Confirm the automation starts a Droid session and links it from the Slack thread. You can review runs from the automation's detail view. ## Configuration details ### Customizing the prompt The automation prompt is injected into every session triggered from the channel. Keep it specific to incident work: - Tell Droid to run RCA and use the `/incident` workflow when appropriate. - Mention the primary services or repositories for that channel. - Point Droid at runbooks, dashboards, or common alert sources. - State escalation expectations, such as when to summarize uncertainty instead of taking action. Avoid prompts that make the automation a general-purpose Slack surface. Incident Response works best when the prompt is narrow and operational. For broader Slack workflows, create a separate [Slack automation](/software-factory/automations) instead. ### Channel selection and trigger source Incident Response is designed for channels where incident alerts arrive as top-level messages. Thread replies are not used to start new sessions, which prevents ordinary discussion from repeatedly launching Droids. For best results, use a dedicated channel such as `#incidents`, `#alerts-production`, or a service-specific incident channel. The channel picker only shows channels where Factory has been invited. **Trigger on messages from** defaults to **Bots only** so ordinary messages in the channel do not trigger investigations; widen it only if humans should also be able to start incident sessions. ### Session privacy and model **Session privacy** controls who can see triggered sessions. **Private** sessions are visible to the run identity. **Team** sessions can be reviewed by other members, which is useful for incident handoff and postmortem review. Use the default model unless your team has a known preference for incident analysis. For complex production incidents, choose a stronger reasoning model if available. ### Managing the automation Open the automation from All Automations to see recent runs and statuses. Depending on your permissions, you can run now, pause, resume, share, fork, rename, edit, or delete the automation. Pause the automation instead of deleting it when you want to temporarily stop incident sessions. ## How the incident flow works Once configured, Incident Response follows this flow: 1. An alert bot posts a top-level message in the configured Slack channel. 2. The automation starts a Droid session using its run identity, computer, session privacy, model, and prompt. 3. Droid receives the Slack alert context and checks for any matching `incident-guidelines` runbook guidance. 4. Droid investigates with the configured tools and should invoke the `/incident` workflow when RCA is appropriate. 5. Factory posts status and session links back into Slack, and Droid can save reusable learnings for similar incidents. The `incident-guidelines` runbook is stored locally at `.factory/skills/incident-guidelines/SKILL.md` for team reuse or `~/.factory/skills/incident-guidelines/SKILL.md` for personal reuse. It should store reusable investigation guidance, not secrets or one-off RCA details. ## Troubleshooting Invite Factory to the channel with `/invite @Factory`, then search for the channel again in the automation's trigger settings. Private channels must invite Factory explicitly. Make sure the automation has a selected **Run on** computer. If **Run as** is set to a service account, that service account must be active and own at least one computer. Confirm the message is a top-level alert in the configured channel, not a thread reply, and that the sender matches the **Trigger on messages from** setting. Also check that the automation is not paused. Check that the selected computer has the right repository, tooling, and credentials. Update the automation prompt with links to runbooks or dashboards Droid should use first. Add more channel-specific context to the prompt: service names, alert sources, dashboard links, log query examples, repository paths, and expected RCA format. Understand how Droid runs work without repeated approvals. Configure persistent environments for investigations. Learn how reusable workflows like incident investigation guide Droid. Schedule and trigger Droid work, including incident-response runbooks. # Custom Automations Create and manage Factory automations that run workflows across your software delivery lifecycle. Automations turn Droid into a teammate for repeatable engineering work. They run from schedules, Slack messages, or GitHub events, covering code review, QA, AutoWiki, security audits, triage, incident response, Slack intake, and scheduled checks. ![Automations list in the Factory App](/docs-assets/images/automations-list.webp) ## What Automations are An automation has: {/* sweep-allow: term-bullets */} - **A trigger**: when it starts. - **Instructions**: the prompt that's delegated to Droid. - **A run target**: where Droid runs, such as a computer, Slack channel, or GitHub workflow. - **Visibility**: who can view or manage it. ### When to use Automations Use Automations when a workflow should run repeatedly, start from an event, or be shared with a team instead of handled manually each time. ## How to use Automations ### View your automations Open All Automations under **Software Factory** in the Factory App sidebar. From the list, you can: - Search by title, description, or SDLC stage. - Switch between **Private** and **Shared**. - Filter by trigger type: **Scheduled**, **Slack**, or **GitHub**. - See owner, status, run target, and SDLC stage. ### Create a new automation Click **New automation**, then choose a template. Software Factory templates map to SDLC stages, while custom templates give you flexible starting points for your own triggers and workflows. ![Automation template gallery](/docs-assets/images/automations-templates.png) #### Software Factory templates Use these templates to add automation to specific SDLC stages in Software Factory. | Template | Use it for | | --- | --- | | **Code Review** | Review pull requests automatically. | | **QA** | Generate and run QA checks. | | **AutoWiki** | Refresh repository documentation. | | **Security Audit** | Run scheduled deep security reviews. | | **Triage** | Scan intake sources and route work. | | **Incident Response** | Investigate alerts from Slack channels. | #### Custom templates Use these templates to create automations around your own channels, schedules, and GitHub events. | Template | Use it for | | --- | --- | | **Slack** | Trigger Droid sessions from Slack messages. | | **Scheduled** | Run a prompt on a recurring schedule. | | **GitHub** | Run Droid from custom GitHub workflow events. | ### Configure triggers After you select a template, Factory tailors setup to the trigger and destination. ![Slack automation create form](/docs-assets/images/automations-slack-create.png) | Trigger type | What you configure | | --- | --- | | **Scheduled** | Schedule, prompt, run target, model, and optional working directory. | | **Slack** | Channels, message source, prompt, run target, privacy, keywords, and model. | | **GitHub** | Repositories, events, workflow settings, and template-specific options. | ### Manage an automation Select an automation row to open its detail view. Depending on the trigger type and your permissions, you can run now, pause, resume, share, fork, rename, edit, or delete the automation. ## Automation settings Use automation settings to control what starts the run, where it executes, who owns it, and how results are tracked. | Setting area | Options | What it controls | | :----------- | :------ | :--------------- | | Trigger type | Scheduled, Slack, GitHub | What starts the automation and what event context Droid receives. | | Scheduled trigger | Natural-language cadence or cron schedule | When a recurring prompt runs. | | Slack trigger | Channels, message source, keywords, privacy, and prompt | Which Slack messages start sessions and how they are scoped. | | GitHub trigger | Repositories, events, workflow settings, and template options | Which repository events, such as pull requests, pushes, comments, labels, or checks, start the automation. | | Run target | Run as, run on, working directory, and model | The user or service account, Droid Computer, repo path, and model used for execution. | | Visibility | Private or shared | Whether only the owner or the organization can see the automation. | | Ownership actions | Run now, pause, resume, share, fork, rename, edit, or delete | What owners can do from the automation detail view. Non-owners can fork shared automations when they need a copy. | | Run history | Recent runs, statuses, and generated dashboard output | How teams inspect automation results over time. | For shared team workflows, use [service accounts](/enterprise/identity-and-access#service-accounts) and [Droid Computers](/droid-computers/overview) to keep runs on a stable identity and environment. Add end-to-end QA as a Software Factory automation. See automation coverage and stage health in one dashboard. # CI Automations API Manage CI automation workflows: scan repositories, list jobs, and open workflow PRs. ## Open a CI workflow PR (create / edit / delete) `POST /api/v0/automations/ci/edit` ```bash curl -X POST 'https://api.factory.ai/api/v0/automations/ci/edit' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Response:** `200` - Response for status 200 ## List CI automation jobs `GET /api/v0/automations/ci/jobs` ```bash curl 'https://api.factory.ai/api/v0/automations/ci/jobs' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Response:** `200` - Response for status 200 ## List CI-eligible repositories `GET /api/v0/automations/ci/repositories` ```bash curl 'https://api.factory.ai/api/v0/automations/ci/repositories' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Response:** `200` - Response for status 200 ## List CI repository owners `GET /api/v0/automations/ci/repository-owners` ```bash curl 'https://api.factory.ai/api/v0/automations/ci/repository-owners' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Response:** `200` - Response for status 200 ## List GitHub Actions runs for CI automation workflows `GET /api/v0/automations/ci/runs` ```bash curl 'https://api.factory.ai/api/v0/automations/ci/runs' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Response:** `200` - Response for status 200 ## Scan repositories for CI workflows `GET /api/v0/automations/ci/scan` ```bash curl 'https://api.factory.ai/api/v0/automations/ci/scan' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Response:** `200` - Response for status 200 # Agent Readiness Understand the Agent Readiness Model across Factory App, Droid CLI, dashboard, and API surfaces. Agent Readiness measures whether repositories have the structure, validation, and operational signals that let autonomous agents work reliably. The Agent Readiness Model scores that foundation so teams can find and remediate gaps before they delegate larger work to Droid. --- ## Getting started There are four ways to interact with Agent Readiness: 1. **Factory App or Droid CLI:** Run the `/readiness-report` [slash command](/agent-readiness/readiness-report) in a Factory App web or desktop session, or in the Droid CLI, to evaluate a repository's readiness level 2. **Factory App dashboard:** View your organization's readiness scores in the [Agent Readiness dashboard](/agent-readiness/dashboard) 3. **API:** Programmatically access readiness reports via the [Readiness Reports API](/api-reference/readiness-reports) 4. **Remediation:** Automatically fix failing criteria with the `/readiness-fix` [slash command](/agent-readiness/readiness-report#remediation-with-readiness-fix) in a Factory App web or desktop session, or in the Droid CLI --- ## What autonomous development looks like What does it look like when an organization reaches high agent readiness? Here are concrete examples of workflows that become possible when the technical foundation is in place. ### Code from conversation A developer describes what they need built, and the system executes through deployment. **Input:** "Refactor the authentication module to support OAuth2 with PKCE, maintaining backward compatibility." **The system:** - Generates idiomatic code following established patterns - Validates against linters, type checkers, and test suites - Handles the pull request and code review process - Updates documentation and notifies stakeholders - Deploys and monitors for issues ### Design to implementation A designer shares a mockup, and the system implements it without handoffs. **Input:** Figma mockup of a new dashboard. **The system:** - Interprets the visual specification and design system - Implements UI components with proper styling and responsiveness - Connects to existing APIs or creates new endpoints - Adds loading states, error handling, and accessibility - Generates tests for user interactions - Deploys to staging for design review ### Bug to deployed fix A customer reports an issue, and the system diagnoses, fixes, and deploys autonomously. **Input:** Bug report through support system. **The system:** - Triages based on error logs and impact - Creates a ticket with reproduction steps and diagnostic context - Identifies the root cause from code analysis - Generates a fix and comprehensive tests - Opens a PR and assigns a developer for review - Notifies support when the fix is deployed --- ## The 5 readiness levels Repositories progress through five distinct levels, each representing a qualitative shift in how autonomous agents can operate within your codebase. | Level | Name | Description | Example Criteria | | :---- | :----------- | :------------------------------------------------------------------------------------------------------------------------------------- | :------------------------------------------------------------------- | | **1** | Functional | Code runs, but requires manual setup and lacks automated validation. Basic tooling that every repository should have. | README, linter, type checker, unit tests | | **2** | Documented | Basic documentation and process exist. Workflows are written down and some automation is in place. | AGENTS.md, devcontainer, pre-commit hooks, branch protection | | **3** | Standardized | Clear processes are defined, documented, and enforced through automation. Development is standardized across the organization. | Integration tests, secret scanning, distributed tracing, metrics | | **4** | Optimized | Fast feedback loops and data-driven improvement. Systems are designed for productivity and measured continuously. | Fast CI feedback, regular deployment frequency, flaky test detection | | **5** | Autonomous | Systems are self-improving with sophisticated orchestration. Complex requirements decompose automatically into parallelized execution. | Self-improving systems | --- ## How scoring works ### Level progression To unlock a level, you must pass **80% of the criteria** from the previous level. This creates a gated progression system: 1. All repositories start at Level 1 2. Pass 80% of Level 1 criteria → Unlock Level 2 3. Pass 80% of Level 2 criteria → Unlock Level 3 4. And so on... ### Evaluation scopes Criteria are evaluated at two different scopes: - **Repository Scope:** Evaluated once for the entire repository (e.g., CODEOWNERS file exists, branch protection enabled) - **Application Scope:** Evaluated per application in monorepos (e.g., linter configured, unit tests exist for each app) For monorepos with multiple applications, application-scoped criteria show scores like `3 / 4` (3 out of 4 apps pass). --- ## The technical pillars The Agent Readiness Model organizes criteria into nine technical pillars that form the foundation for autonomous operation. | Pillar | Why it matters | Example criteria | | :----- | :------------- | :--------------- | | **Style & Validation** | Linters, type checkers, and formatters catch obvious errors instantly. Agents avoid wasting cycles on syntax errors, style inconsistencies, and type mismatches. | Linter configuration, type checker, code formatter, pre-commit hooks | | **Build System** | Clear, deterministic build commands let agents verify their changes compile and run before committing. No guessing about which commands to run or which flags to pass. | Build command documented, dependencies pinned, VCS CLI tools | | **Testing** | Fast unit and integration tests create tight feedback loops. Agents learn whether their changes work correctly by running tests, seeing failures, and iterating. | Unit tests exist, integration tests exist, tests runnable locally | | **Documentation** | Explicit instructions house the tribal knowledge that "everyone just knows." How to set up the environment, run tests, deploy changes, debug issues. Agents need written instructions that are discoverable, accurate, and maintained. | AGENTS.md, README, documentation freshness | | **Development Environment** | Reproducible environments ensure consistency. When developers and agents work in identical environments, entire classes of problems disappear. No more "works on my machine." | Devcontainer, environment template, local services setup | | **Debugging & Observability** | Structured logging, tracing, and metrics give agents runtime visibility into what code actually does. Good observability turns "it failed" into "it failed because X was null when calling Y after receiving Z." | Structured logging, distributed tracing, metrics collection | | **Security** | Branch protection, secret scanning, and code owners prevent agents from introducing security issues or bypassing required reviews. Agents move fast. Automated guardrails ensure they move fast safely. | Branch protection, secret scanning, CODEOWNERS | | **Task Discovery** | Infrastructure for agents to find and scope work autonomously. Well-structured issues and templates help agents understand what needs to be done. | Issue templates, issue labeling system, PR templates | | **Product & Experimentation** | Tools for measuring impact, running experiments, and understanding user behavior. Agents can see whether features are actually used and measure the impact of their changes. | Product analytics instrumentation, experiment infrastructure | Evaluate your repository with the `/readiness-report` slash command. Track readiness scores across your organization in the Factory App. # Readiness Report Command Evaluate your repository's agent readiness with the /readiness-report slash command. The `/readiness-report` slash command evaluates your current repository against the Agent Readiness Model, scoring it across five readiness levels and providing actionable recommendations to improve. --- ## Prerequisites - Run the command from inside a git repository with an `origin` remote configured (the repository URL is used to associate the report with your project in the Factory App). --- ## Usage To use the command in the Factory App (web or desktop), open a session for the repository. In the Droid CLI, navigate to the repo you'd like to evaluate and start Droid: ```bash droid ``` Then enter the slash command: ```text > /readiness-report ``` The evaluation runs against your current repo directory. --- ## What happens When you run `/readiness-report`, the droid performs a comprehensive evaluation of the repository: 1. **Language Detection**: Identifies repository languages (JavaScript/TypeScript, Python, Rust, Go, Java, Ruby) based on configuration files and source code 2. **Sub-application Discovery**: Determines whether the repository is a mono-repo, or a single service/package/library. For mono-repos, this step identifies all independently deployable applications within the repo 3. **Criteria Evaluation**: Checks the criteria across all five readiness levels 4. **Report Storage**: Persists the evaluation results for visualization in the Factory App 5. **Summary Output**: Prints a human-readable report with results from the evaluation --- ## Understanding the output After evaluation, `/readiness-report` prints a structured report with the repository's readiness level, discovered applications, criterion scores, and recommended next actions. ### Report sections | Section | What it shows | Example | | :------ | :------------ | :------ | | Level achieved | Current repository readiness level from 1 to 5. | `Level 3: Standardized` | | Applications discovered | Each independently deployable app or package found in a monorepo. | `apps/web - Main Next.js application` | | Criteria results | Score and rationale for each evaluated criterion. | `Linter Configuration: 2/2 - ESLint configured in both applications` | | Action items | Two to three highest-impact recommendations for reaching the next level. | `Add pre-commit hooks with husky to enforce linting and formatting` | ### Readiness levels | Level | Name | Meaning | | :---- | :--- | :------ | | 1 | Functional | Basic tooling is in place. | | 2 | Documented | Process and documentation are established. | | 3 | Standardized | Security and observability are configured. | | 4 | Optimized | Fast feedback and continuous measurement are in place. | | 5 | Autonomous | Systems can improve themselves with agent support. | Scores use `numerator/denominator`: the numerator is the number of sub-applications that pass the criterion, and the denominator is the number evaluated. ### Criterion difficulty Each criterion also carries a **difficulty**: `Basic`, `Intermediate`, or `Advanced`. Difficulty is a separate axis from the readiness level: it reflects how much effort a criterion typically takes to satisfy, not which level it belongs to. When closing gaps, start with `Basic` criteria (the highest-leverage foundations every repository should have), then work through `Intermediate` and `Advanced`. --- ## Viewing historical reports All readiness reports are automatically saved and can be viewed in the [web dashboard](/agent-readiness/dashboard). This allows you to: - Track readiness progression over time - Compare scores across repositories - Share results with your team Run `/readiness-report` periodically (e.g., after major infrastructure changes) to track your organization's progress toward higher readiness levels. --- ## Remediation with `/readiness-fix` Once you've generated a readiness report, the `/readiness-fix` slash command lets the droid automatically remediate the failing signals from the latest report. That turns the report from a diagnostic tool into an automated improvement workflow. ### What it does `/readiness-fix` fetches the most recent stored readiness report for your repository and starts an agent session that works through the failing criteria, implementing fixes such as adding pre-commit hooks, creating an `AGENTS.md`, configuring CI checks, or improving documentation. ### Usage In the Factory App (web or desktop), open a session for the same repository you ran `/readiness-report` against. In the Droid CLI, navigate to that repository and start Droid: ```bash droid ``` Then enter the slash command: ```text > /readiness-fix ``` You can also pass additional natural-language instructions to focus the run: ```text > /readiness-fix prioritize style & validation; do not touch CI configuration ``` ### What to expect 1. **Report lookup**: The command resolves your repository from `git remote get-url origin` and fetches the latest readiness report stored for that repo. 2. **Agent session**: The droid plans and applies changes to address failing criteria, one at a time, just like a normal coding session. You can review and approve changes as they happen. 3. **Verification**: After fixes are applied, run `/readiness-report` again to re-score the repository and confirm signals have moved. `/readiness-fix` requires that a readiness report has already been generated for the repo. If no prior report is found, run `/readiness-report` first. Treat `/readiness-fix` like any other agent run: review the proposed changes, run your tests, and commit only what looks good for your codebase. View and track readiness scores across your repositories. Understand the five readiness levels and how scoring works. # Readiness Dashboard Track and analyze your organization's Agent Readiness scores in the Factory App. The Agent Readiness dashboard provides a centralized view of your organization's readiness scores across all repositories, with historical trends and detailed breakdowns. ![Agent Readiness dashboard](/docs-assets/images/web/readiness-dashboard.png) --- ## Accessing the dashboard Navigate to **Settings → Analytics → Agent Readiness** in the Factory App. --- ## Main dashboard view The dashboard shows an organization-level overview with three main sections: ### Summary cards At the top, you'll see key metrics: | Metric | Description | | :----------------------- | :-------------------------------------------------------------------------- | | **Organization Score** | Average readiness level across all measured repositories (rounded down) | | **Repositories Tracked** | Number of repositories with readiness reports vs. total enabled repositories | | **Last Updated** | Time since the most recent readiness evaluation | ### Progress graph A time-series chart showing your organization's readiness level over time. Use the period filters to view different time ranges: Last 7 days. Last month. Last 6 months. Last year. All time. ### Repositories table A searchable, paginated table of all repositories showing: Click to view details. Current readiness level (1-5). Percentage complete toward next level. When the repository was last evaluated. Use the search bar to filter repositories by name or URL. --- ## Repository detail page Click any repository row to view its detailed readiness breakdown. ![Repository Detail Page](/docs-assets/images/web/readiness-criteria.png) ### Header Shows the repository name, current level achieved, last evaluation time, and a **Refresh** button to trigger a new evaluation. ### Level accordions Each readiness level (Functional through Autonomous) has an expandable accordion section showing: - **Percentage complete**: How much of that level's criteria are passing - **Lock status**: Levels are locked until the previous level reaches 80% ### Criterion rows Within each level, individual criteria display: | Element | Description | | :------------ | :------------------------------------------------- | | **Name** | The criterion being evaluated | | **Score** | Format `[X/Y]`: numerator/denominator | | **Status** | Pass (green) or fail (red) indicator | | **Rationale** | Click to expand and see the evaluation explanation | ### Remediation Planned Failing criteria will display a **Fix** button that triggers automated remediation. Select the criteria you want to fix, and the system will implement the necessary changes to your repository. --- ## Triggering a refresh There are three ways to refresh a readiness evaluation: ### From the Factory App dashboard 1. Navigate to the repository detail page 2. Click the **Refresh** button in the header 3. A new session starts to evaluate the repository 4. Once done, the dashboard will update with results from the latest report ### From a Factory App session Open a web or desktop session for the repository, then enter the `/readiness-report` [slash command](/agent-readiness/readiness-report): ```text > /readiness-report ``` ### From the Droid CLI Start Droid while in the repository directory: ```bash droid ``` Then enter the `/readiness-report` [slash command](/agent-readiness/readiness-report): ```text > /readiness-report ``` Re-evaluations run the full readiness assessment against the current state of your repository. This is useful after making infrastructure improvements or merging changes that can affect a readiness criterion. --- ## Understanding the metrics ### Organization level calculation The organization-level score is calculated as the **average of all repository levels, rounded down**. For example: - Repo A: Level 3 - Repo B: Level 2 - Repo C: Level 3 Organization Level = floor((3 + 2 + 3) / 3) = floor(2.67) = **Level 2** ### Repository level calculation A repository's level is determined by the **80% threshold system**: 1. Start at Level 1 2. If 80% of Level 1 criteria pass → achieve Level 2 3. If 80% of Level 2 criteria pass → achieve Level 3 4. Continue through Level 5 The percentage shown for each level indicates progress toward completing that level's criteria. --- ## Best practices {/* sweep-allow: term-bullets */} - **Regular evaluations:** Run readiness reports after significant infrastructure changes - **Focus on current level:** Address failing criteria in your current level before jumping ahead to focus on foundational improvements first - **Track trends:** Use the progress graph to monitor improvement over time - **Prioritize high-impact fixes:** The action items in CLI reports highlight the most impactful improvements Run `/readiness-report` to evaluate a repository from a Factory App or CLI session. Learn the five readiness levels and how scoring works. # Agent Readiness Reports API REST API for programmatic access to Factory agent readiness reports. The Agent Readiness Reports API returns Factory readiness evaluations for repositories in your organization. --- ## Authentication All requests require a Factory API key in the `Authorization` header. ```bash Authorization: Bearer fk-your-api-key ``` [Factory API keys settings](https://app.factory.ai/settings/api-keys) --- ## Base URL `https://app.factory.ai` --- ## List readiness reports Retrieves readiness reports for your organization. Filter reports by repository ID Maximum number of reports to return (must be positive) Report ID for pagination cursor ```json { "reports": [ { "reportId": "550e8400-e29b-41d4-a716-446655440000", "createdAt": 1701792000000, "repoUrl": "https://github.com/org/repo", "apps": { "apps/web": { "description": "Main Next.js application" }, "apps/api": { "description": "Backend API service" } }, "report": { "lint_config": { "numerator": 2, "denominator": 2, "rationale": "ESLint configured in both applications" }, "type_check": { "numerator": 2, "denominator": 2, "rationale": "TypeScript strict mode enabled" } }, "commitHash": "abc123def456", "branch": "main", "hasLocalChanges": false, "hasNonRemoteCommits": false, "modelUsed": { "id": "claude-sonnet-4-5-20250929", "reasoningEffort": "high" }, "droidVersion": "0.30.0" } ] } ``` ```bash curl -X GET "https://app.factory.ai/api/organization/agent-readiness-reports?limit=10" \ -H "Authorization: Bearer fk-your-api-key" ``` --- ## Readiness report schema Unique identifier for the report (UUID) Unix timestamp in milliseconds when the report was created Repository URL that was evaluated Map of application paths to description objects Map of criterion IDs to evaluation results Git commit hash at time of evaluation Git branch name at time of evaluation Whether uncommitted changes existed Whether unpushed commits existed Model configuration used for evaluation CLI version that generated the report Brief description of what the application does Number of sub-applications passing the criterion (0 to denominator) Number of sub-applications on which the criterion was evaluated (minimum 1) Explanation of the evaluation result Model identifier Reasoning effort level (`low`, `medium`, `high`, `off`) --- ## Pagination For large result sets, use cursor-based pagination: 1. Make initial request with desired `limit` 2. Get the `reportId` of the last item in the response 3. Pass that ID as `startAfter` in the next request ```bash # First page curl "https://app.factory.ai/api/organization/agent-readiness-reports?limit=10" # Next page (using last reportId from previous response) curl "https://app.factory.ai/api/organization/agent-readiness-reports?limit=10&startAfter=550e8400-e29b-41d4-a716-446655440000" ``` --- ## Use cases ### CI/CD integration Track readiness scores over time by fetching reports after each evaluation: ```bash # Get latest report for a specific repository curl "https://app.factory.ai/api/organization/agent-readiness-reports?repoId=123&limit=1" \ -H "Authorization: Bearer $FACTORY_API_KEY" ``` ### Custom dashboards Build internal dashboards by fetching all reports and calculating aggregate metrics: ```javascript const response = await fetch( "https://app.factory.ai/api/organization/agent-readiness-reports", { headers: { Authorization: `Bearer ${apiKey}` } } ); const { reports } = await response.json(); // Calculate average level const avgLevel = reports.reduce((sum, r) => sum + calculateLevel(r), 0) / reports.length; ``` ### Automated alerting Set up alerts when readiness scores drop below thresholds: ```bash # Fetch recent reports and check for regressions reports=$(curl -s "https://app.factory.ai/api/organization/agent-readiness-reports?limit=50" \ -H "Authorization: Bearer $FACTORY_API_KEY") # Process and alert on regressions... ``` --- ## Errors | Status | Description | | :----- | :------------------------- | | `400` | Invalid request parameters | | `401` | Missing or invalid API key | | `500` | Internal server error | # Agent Effectiveness Measure how much faster your engineering organization ships with Factory, and connect agent spend to the work it produces. Agent Effectiveness measures how much faster your engineering organization ships with Factory, and connects agent spend to the work it produces. Most teams that adopt AI coding agents can tell that work is moving faster, but they have a hard time proving it. Survey-based estimates of "how long would this have taken before AI" are unreliable, and existing delivery metrics show how fast work moves without showing how much of that speed came from agents. Agent Effectiveness answers that question from data your organization already produces: sessions in Factory, and issues, projects, and pull requests in the tools your teams already use. --- ## How it works Agent Effectiveness reads from your connected project management, issue-tracking, and source control integrations, then links agent activity to the work those tools track. As agent spend is applied, cycle times on projects, issues, and pull requests change measurably, and teams use the recovered time either to take on more work or to raise the bar on the work they already have. Analysis happens at the organization level and requires no per-repository configuration. --- ## What it measures How cycle times on projects, issues, and pull requests change as agent spend is applied, and which projects, pods, and users the change is concentrated in. Where effort is going, using session intents to classify work and compare the split of engineer time and Factory Standard Credits against your targets. The link between an individual session, the work stream it belongs to, and the issues, projects, and artifacts it produced. See the [Effectiveness dashboard](/agent-effectiveness/dashboard) for how each view is presented in the Factory App. --- ## Session intents A session intent is a classification of what a session was for, such as feature development, maintenance, bug fixing, or exploration. Intents let you compare planned allocation against actual allocation during the quarter rather than after it. When the mix drifts from your targets, the Output view surfaces the drift while there is still time to act on it. --- ## Attribution and local signals Attribution connects a session to the concrete work it produced. It relies on local signals, which record the association between a session and the repository work it touched, so an organization-level number can be traced back to specific issues, projects, and artifacts. Attribution and the Output view only cover the systems you connect. Connect every issue tracker and source control provider your teams use so coverage is complete. --- ## Other usage and adoption metrics Agent Effectiveness focuses on delivery speed and the link between spend and shipped work. Adjacent measures are documented separately: | Measure | Where it lives | | :------ | :------------- | | Factory Standard Credits consumption, tool usage, user activity, productivity, and per-user metrics | [Analytics API](/api-reference/analytics) | | Self-hosted OpenTelemetry metrics export | [Telemetry & Analytics](/enterprise/telemetry) | | Cost management and productivity measurement guidance | [Cost & Productivity](/agent-effectiveness/cost-and-productivity) | --- ## Requirements Agent Effectiveness requires organization-level integrations plus the Advanced Analytics enterprise control. See [Enable Agent Effectiveness](/agent-effectiveness/setup) for the full setup. Agent Effectiveness is in Private Preview. To request access, fill out the form at factory.ai/contact or contact your Factory account team. Connect integrations and turn on Advanced Analytics. Explore the Throughput, Output, and Attribution views in the Factory App. Query credits consumption, tool usage, and per-user productivity metrics. Export metrics to your own observability stack over OpenTelemetry. # Enable Agent Effectiveness Connect your integrations and turn on Advanced Analytics to enable Agent Effectiveness. Agent Effectiveness depends on two things: integrations that describe where your work lives, and the Advanced Analytics enterprise control that enables collection and analysis. Setup requires the Owner or Manager role. --- ## Prerequisites - The Owner or Manager role in the Factory App. - At least one issue tracking or source control integration configured at the organization level. --- ## Setup Configure organization-level integrations for the tools your teams use. | Category | Supported integrations | | :------- | :--------------------- | | Issue tracking and project management | Jira, Linear | | Source control | GitHub, GitLab | These integrations let Factory map sessions to issues, projects, and pull requests. Coverage is limited to the systems you connect, so connect every tool your teams track work in. Turn on the Advanced Analytics enterprise control to make the Throughput, Output, and Attribution views available to your organization. Advanced Analytics includes the local signals that associate sessions with the repository work they touched and automatically backfills effectiveness data to the start of the account. --- ## Verifying setup After Advanced Analytics is enabled, open the [Effectiveness dashboard](/agent-effectiveness/dashboard). Populated Throughput and Attribution views confirm that sessions are being linked to your connected systems. If the views are empty, confirm that the relevant integration is connected at the organization level rather than only for an individual user, and that sessions have run since Advanced Analytics was enabled. Agent Effectiveness is in Private Preview. To request access, fill out the form at factory.ai/contact or contact your Factory account team. Understand what Throughput, Output, and Attribution measure. Explore the dashboard views once setup is complete. # Effectiveness Dashboard Explore the Throughput, Output, and Attribution views for Agent Effectiveness in the Factory App. The Agent Effectiveness dashboard presents organization-level analysis across three views: Throughput, Output, and Attribution. --- ## Accessing the dashboard Navigate to **Settings → Analytics → Agent Effectiveness** in the Factory App. The dashboard requires the Advanced Analytics enterprise control. If you do not see it, see [Enable Agent Effectiveness](/agent-effectiveness/setup). --- ## Throughput The Throughput view shows how delivery speed changes as agent spend is applied, and where the change is concentrated. Project, issue, and pull request cycle times over the selected period. Which projects, pods, and users consume the most resources. Consumption measured against your organization's targets for the period. Use the Throughput view to check whether resource consumption matches organizational intent. Concentration in a small number of projects is not inherently a problem, but it should be a deliberate choice rather than a surprise. --- ## Output The Output view maps spend to intent, so you can see what the work actually was rather than only how much of it there was. Sessions are classified by intent, such as feature development, maintenance, bug fixing, or exploration. The view compares the resulting split of engineer time and Factory Standard Credits against your planned allocation for the period. Reviewing this view during the quarter, rather than at the end, is what makes it actionable: a drift between planned and actual allocation can still be corrected while the quarter is in progress. --- ## Attribution The Attribution view traces an individual session to the work stream it belongs to and the output it produced. The originating Factory session. The issue or project the session maps to in your connected tracker. The pull requests and repository changes the session produced. Attribution depends on local signals, which are included with Advanced Analytics. When Advanced Analytics is enabled, Factory automatically backfills effectiveness data to the start of the account. --- ## Interpreting the data Cycle time compression is a trend, not a single number. Compare periods of comparable length and account for release cadence, holidays, and headcount changes before attributing a shift to agent adoption. Attribution coverage is bounded by your connected integrations. A work stream tracked in a system you have not connected will not appear, which can make coverage look lower than actual usage. Agent Effectiveness is in Private Preview. To request access, fill out the form at factory.ai/contact or contact your Factory account team. Understand what Throughput, Output, and Attribution measure. Connect integrations and turn on Advanced Analytics. # Cost & Productivity Control LLM spend with model policy and autonomy defaults, and measure Droid's productivity impact with the measurement surfaces you already have. Droid gives you two measurement surfaces: [OTEL metrics exported](/enterprise/telemetry) to your own observability stack, and the hosted [Analytics API](/api-reference/analytics). This page is guidance for using them to control spend and to measure what the spend produces. ## Cost management strategies LLM cost control is a combination of **model policy**, **usage patterns**, and **observability**. ### Constrain the model catalog Use org-level policies to limit which models are available. - Prefer smaller models for everyday tasks; reserve large models for complicated refactors or design work. - Disable experimental or high-cost models by default. - Enforce model choices per environment, such as cheaper models in CI. See [Models](/models) for the current model catalog. ### Tune autonomy and context usage Higher autonomy and larger context windows consume more tokens. - Set reasonable defaults for autonomy level and reasoning effort. - Use hooks to cap context size or block unnecessary large prompts. - Encourage teams to iterate with tighter scopes, such as specific directories instead of entire monorepos. ### Monitor activity and cost Combine both measurement surfaces: - Feed [exported activity metrics](/enterprise/telemetry/data-reference), including tool invocations, code activity, and git activity, into your observability stack to build per-team and per-tool dashboards. - Use the [Analytics API](/api-reference/analytics) for token consumption and cost estimates, which are not exported as customer OTEL metrics. - Alert on unusual spikes and compare trends before and after policy changes. ## Measuring productivity impact Cost only matters in the context of outcomes. You can correlate Droid usage with **software delivery and quality metrics** you already track. Common approaches: - Build dashboards from exported activity metrics (files and lines modified, commits, pull requests, tool invocations) per team and repository. - Pull aggregated adoption and productivity signals from the [Analytics API](/api-reference/analytics) for leadership reporting. - Measure how often Droid is involved in changes that reduce incidents, resolve alerts, or improve test coverage. - Use [Agent Effectiveness](/agent-effectiveness/overview) to connect agent spend to cycle-time changes on the issues, projects, and pull requests it touched. These analyses run entirely in your existing observability and analytics stack; Factory's role is to provide clean, structured signals from Droid. Measure how much faster your organization ships with Factory. Query credits consumption, tool usage, and per-user productivity metrics. Export OTEL metrics to your own observability stack. Every metric and attribute available for your dashboards. # Factory Missions Use Factory Missions to plan and execute large, multi-feature projects with structured orchestration. Describe your goal, collaborate on the plan, and let Droid manage the work. ![Mission Control orchestration view](/docs-assets/images/mission-control.webp) ## What Missions do Factory Missions are structured workflows for taking on large, multi-feature work with Droid. Instead of tackling everything in a single session, you collaborate with Droid upfront to build a plan (features, milestones, and the skills needed to accomplish each part), then hand off execution to an orchestration layer that manages the work. Access Missions with the `/missions` command (also available via `/mission`). Factory Missions are the orchestration feature. Mission Mode is the session state Droid enters after you approve a mission plan; see [Interaction Modes](/autonomy-and-safety/specification-mode) for how Mission Mode relates to Normal and Spec modes. ## When to use Missions Missions are best for bounded, multi-feature efforts that benefit from upfront planning, milestones, and ongoing validation. As a rough guide, the sweet spot is **about 1–500 features**: {/* sweep-allow: term-bullets */} - **A small number of features:** Use a normal session when the work is straightforward and can be completed without a separate planning and orchestration phase. - **About 1–500 features:** Start a new Mission so Droid can decompose the work into milestones, coordinate workers, and validate progress as it goes. - **More than 500 features:** Split the work into multiple Missions, or use a longer-term operating workflow instead of treating it as one Mission. These numbers are a planning heuristic, not a hard limit. The more important question is whether the work has a clear outcome and can be divided into meaningful milestones. Start a Mission when you recognize that the work needs this structure. Do not run a normal session for a long time and then convert it into a Mission. Starting a new Mission preserves the upfront planning, feature breakdown, validation strategy, and orchestration context that Missions are designed to provide. Work with Droid to define features, milestones, and success criteria before any code is written. Droid reuses your existing skills and develops new specialized skills for each part of the work. Mission Control manages execution across agents, tracking progress through your plan. MCP integrations, skills, hooks, and custom droids all work inside Missions. ## For optimal outcomes For the best results, your repository should be at [Agent Readiness](/agent-readiness/overview) **Level 4 (Optimized) or above**. As a mission works, it runs user-facing QA testing against your application to validate each feature and self-correct as it goes. For this to work in an existing project, your codebase needs an automated, scriptable way to exercise the app the way a user would (for example, a script to stand up or mock all dependencies of the app to simulate all potential user flows). Without it, the mission cannot reliably verify its own work. Not there yet? Re-run your readiness evaluation with [`/readiness-report`](/agent-readiness/readiness-report), then close the gaps with [`/readiness-fix`](/agent-readiness/readiness-report#remediation-with-readiness-fix). ## How it works Start by running `/missions` in any Droid session. Droid interacts with you back and forth to understand your goal. It asks clarifying questions, probes for constraints, and works with you to define what you actually want built. This is a conversation, not a one-shot prompt. Based on the conversation, Droid constructs a structured plan: a set of features organized into milestones. Each milestone represents a meaningful checkpoint in the work. Droid pulls in your existing skills where they apply, and develops specialized skills for parts of the work that need them. This means the execution is tailored to your project and workflow, not generic. Once the plan is approved, Droid enters Mission Control, the Missions orchestration view that manages execution of the plan. You can monitor progress, see which features are being worked on, and intervene when needed. For details on getting the plan right, see [Planning & Validation](/missions/planning). To run and steer an approved mission, see [Running in the CLI](/missions/running-cli) or [Running in the Factory App](/missions/running-app). ## What Missions are good for We have built and tested Missions across a range of work: {/* sweep-allow: term-bullets */} - **Full-stack development**: Building complete applications with frontend, backend, database, and deployment. - **Research**: Deep investigation tasks that require exploring multiple approaches, synthesizing findings, and producing structured output. - **Brownfield migrations**: Modernizing existing codebases, swapping frameworks, or restructuring large projects while preserving existing behavior. - **Ambitious prototypes**: Product experiments that need to be functional, not just sketched out. The common thread: work that benefits from upfront planning and structured decomposition rather than ad-hoc prompting. ## Open questions Missions are still evolving. There are fundamental questions we are working through: {/* sweep-allow: term-bullets */} - **Is parallelization necessary?** Running multiple agents in parallel sounds good in theory, but does it actually produce better results than sequential execution? We are testing this. - **How do you maximize correctness?** Long-running plans accumulate errors. What validation and correction strategies work best at each stage? - **Cost vs. quality tradeoffs**: How aggressive should the orchestrator be? More planning and validation means higher cost but potentially better output. Where is the right balance? We want your feedback on these. Use Missions, push the workflow hard, and tell us what works and what does not. ## Troubleshooting Missions are not fire-and-forget. The orchestrator is an agent, and you can talk to it. When something goes wrong, pause the orchestrator, describe what you are seeing in plain language, and ask it to recover. Tell the orchestrator the mission appears frozen, describe the last visible activity, and ask it to re-assess and continue. For example: *"The mission seems frozen, the last worker finished 10 minutes ago and nothing new has started. Re-assess and continue."* Pause the orchestrator and tell it to mark the current item as complete and move on. You can revisit that item later or handle it manually. For example: *"The worker on the auth integration has been stuck for 20 minutes. Mark it as complete and move to the next feature."* Ask the orchestrator to re-assess the remaining work and identify what is blocking progress. It can re-plan around the obstacle, reorder features, or adjust milestone scope. Pause and explain the new requirement or scope change. The orchestrator can update the plan, re-scope milestones, and continue from the new direction. Inspect the latest worker for that feature to understand why it keeps failing. If the failures come from user-testing validation, make sure your project has a script that stands up or mocks dependencies so QA can simulate real user flows. See [Planning & Validation](/missions/planning#development-scripting-for-qa). If the project does not need QA-style validation, disable it in Mission Control settings. Make sure the project can be started or mocked from a script, and that logs are written to the filesystem where Droid can inspect them. See [Planning & Validation](/missions/planning#development-scripting-for-qa). If QA-style validation is not useful for the project, disable it in Mission Control settings. Scope features, milestones, and validation frequency before execution begins. Monitor and steer an approved mission from the terminal. # Planning & Validation The planning phase is where Missions deliver the most value. Learn how to scope features and milestones, set validation frequency, and estimate cost and duration. ## The planning phase matters most The biggest value we have found in Missions is in the planning phase. Getting the upfront plan right (the features, the ordering, the milestones, the skills involved, and how the work gets validated) is what determines whether the execution succeeds. Droid will push back, ask questions, and iterate with you until the plan is solid. This is intentional. A well-scoped plan with clear milestones produces dramatically better results than jumping straight into execution on a vague goal. ## Validation Milestones define validation frequency. Validation workers run at the end of each milestone, verifying its work. For simple projects, one milestone is often enough; for longer or complex projects, more frequent milestone validation helps keep the foundation stable as work scales. For smaller, straightforward projects, a single milestone is often enough. For larger or longer-running projects, more granular milestones can prevent drift and reduce expensive rework later. If your project is not one that requires QA-style validation, you can disable it in the mission settings inside Mission Control. ## Development scripting for QA Missions validate their own work by exercising your running application, so one of the most valuable things you can prepare is reliable scripting that lets Droid stand up, drive, and observe the app. Good patterns we have seen: {/* sweep-allow: term-bullets */} - **One command to start the app.** For a web app, provide a single script that starts both the backend and frontend (plus any required services) so a worker can bring the whole stack up reproducibly. - **Route logs to the filesystem.** Send application logs to files on disk so Droid can read and inspect them after each action. Logs that only stream to a terminal are much harder for a worker to use. - **Keep resource usage modest.** Make sure running the app does not consume too many resources (RAM, CPU, disk). Workers run alongside the app, and a heavy local stack slows down or destabilizes the mission. - **Provide a way to send input.** Give Droid a programmatic way to drive the app the way a user would. We ship some tooling by default (`tuistory` and `agent-browser`) to enable QA testing of web apps, Electron apps, and terminal UI applications. If your app does not fall into one of these categories, we strongly encourage finding a way to give Droid a custom toolchain for driving it. Logs written to disk can capture secrets (credentials, tokens, session IDs, or PII), especially on shared machines or when logs are uploaded as CI artifacts. Redact sensitive fields, restrict file permissions on the log files, and make sure they are excluded from version control (for example via `.gitignore`) and from CI artifact collection. ## Estimating cost and duration As a rough planning heuristic, mission duration and cost scale with the number of worker runs: - **Feature workers:** roughly one run per feature - **Validator workers:** 2 runs per milestone, assuming that validation passes on the first go. So an initial estimate is approximately: `total runs ≈ #features + 2 * #milestones` In practice, this is a floor rather than a ceiling. Validation may surface issues that require follow-up work, and the orchestrator can create additional fix features during execution. Monitor the approved plan visually in Mission Control. Steer and intervene during mission execution from the terminal. # Running in the CLI Once a plan is approved, Droid enters Mission Control. Monitor progress, unblock workers, and redirect the orchestrator from the terminal. ## Working with Mission Control Once the plan is approved, Droid enters Mission Control, the orchestration view that manages execution. From here you can track progress across features and milestones, see which agents are working on what, and intervene when things need adjustment. Prefer a visual dashboard? The Factory App provides a richer Mission Control experience. See [Running in the Factory App](/missions/running-app). ## Intervening and redirecting Missions are not fire-and-forget. The orchestrator is an agent, and you can talk to it. The most effective way to use Missions is to treat yourself as the project manager: monitor progress, unblock workers, and redirect when the plan needs to change. When something goes wrong (the mission freezes, a worker or milestone gets stuck, or you need to change direction), pause the orchestrator and tell it what you are seeing. See [Troubleshooting](/missions/overview#troubleshooting) for common scenarios and example prompts. ## A new kind of debugging The skillset for working with Missions looks less like traditional debugging and more like **project management of agents**. You are not stepping through code line by line. You are monitoring a team of workers, unblocking them when they get stuck, redirecting them when priorities change, and making judgment calls about when to push through versus when to re-plan. This is a meaningfully different way of working with AI. The core skill is knowing when and how to intervene, not writing the code yourself. Run missions non-interactively in CI or scheduled environments. Use the visual Mission Control dashboard for richer monitoring. # Running in the Factory App Use Factory Missions in the Factory App to plan and execute large, multi-feature projects with structured orchestration. ![Mission Control dashboard in the Factory App](/docs-assets/images/mission-web.webp) ## The Mission Control dashboard While you can run Missions from the CLI, the Factory App provides a rich, visual Mission Control dashboard. This interface helps you act as a project manager by giving you greater visibility into what your orchestrator, workers, and validators are doing. ## Starting and resuming Missions {/* sweep-allow: term-bullets */} - **Start a new mission:** Start a new session in **Mission Mode**, or start a mission directly from the Mission Control page. Start the Mission before doing the work you want it to plan and orchestrate; do not use a long-running normal session as a substitute and convert it afterward. - **Continue a past mission:** Navigate to the Mission Control tab to find past missions, review their current state, and resume execution exactly where you left off. - **Run on remote machines:** When starting a mission, you can select a **Droid Computer** as your target environment. This allows long-running missions to execute safely in the background on persistent remote machines without tying up your local workspace. ## Navigating the UI The Mission Control dashboard gives you real-time visibility into your multi-agent workflows: {/* sweep-allow: term-bullets */} - **Track Progress:** See exactly what has been done so far. The top bar tracks overall time and credits, while the right sidebar maintains a live progress log of every worker and milestone. - **Select Models:** Use the model dropdowns in the right sidebar to configure different models for the Orchestrator, Worker, and Validator agents on the fly. - **Inspect Features:** The main view provides a high-level summary of the active milestone. You can drill down into any feature to review its specific validation criteria and commits. - **Examine Workers:** The left sidebar lists all active and completed workers. Select a specific worker to view its detailed terminal output, thought process, and any sub-tasks it completed. - **Intervene & Redirect:** If a Mission appears stuck or priorities change, pause the orchestrator directly from the UI, instruct it to re-assess or re-plan, and then resume execution. Tune mission models, reasoning effort, and enterprise access policy. Monitor and steer missions from the terminal instead of the app. # Missions Configuration & Reference Configuration inheritance, headless execution with droid exec --mission, mission settings, and enterprise access policy. ## Configuration inheritance Missions inherit your existing Droid configuration: {/* sweep-allow: term-bullets */} - **MCP integrations**: Workers can use your connected tools (Linear, Sentry, Notion, etc.) - **Custom skills**: Your existing skills are available and new ones can be developed during planning. - **Hooks**: Lifecycle hooks fire during mission execution. - **Custom droids**: Subagents configured in your project are available to workers. - **AGENTS.md**: Workers follow your project conventions and coding standards. ## Headless mission execution Missions can also run non-interactively via `droid exec --mission`. This is useful for CI, scheduled jobs, and any environment where you want the orchestrator to plan and execute without a live TUI. ```bash droid exec --mission -f mission.md ``` You can override the models and reasoning effort used by the orchestrator's worker and validator agents: ```bash droid exec --mission \ --worker-model claude-sonnet-4-6 \ --worker-reasoning-effort medium \ --validator-model claude-opus-4-7 \ --validator-reasoning-effort high \ -f mission.md ``` | Flag | Description | | ----------------------------------- | ------------------------------------------------------------ | | `--mission` | Run `droid exec` in Mission Mode (multi-agent orchestration). | | `--worker-model ` | Model used for mission worker agents. | | `--worker-reasoning-effort ` | Reasoning effort for mission worker agents. | | `--validator-model ` | Model used for mission validator agents. | | `--validator-reasoning-effort ` | Reasoning effort for mission validator agents. | See [`droid exec`](/droid-exec/overview) for the full headless reference. ## Configuration Missions are tuned through the `missionModelSettings` object and a few top-level keys. Set these in your global or project [settings](/droid-cli/settings) file. | Setting | Description | | ------------------------------------------------------ | ---------------------------------------------------------------------------- | | `missionModelSettings.workerModel` | Default model used by mission worker subagents. | | `missionModelSettings.workerReasoningEffort` | Reasoning effort for mission workers (`off`, `none`, `low`, `medium`, `high`). | | `missionModelSettings.validationWorkerModel` | Model used by validation workers (scrutiny and user-testing). | | `missionModelSettings.validationWorkerReasoningEffort` | Reasoning effort for validation workers. | | `missionModelSettings.skipScrutiny` | Skip scrutiny validation milestones during missions. | | `missionModelSettings.skipUserTesting` | Skip user-testing validation milestones during missions. | | `missionOrchestratorModel` | Model used by the mission orchestrator. | | `missionOrchestratorReasoningEffort` | Reasoning effort for the mission orchestrator. | | `keepSystemAwakeDuringMissions` | Prevent the OS from sleeping while a mission is running. Defaults to `true`. | Pairing a strong orchestrator model with a faster worker model is a common cost-quality tradeoff: planning and validation benefit most from extra reasoning, while routine worker tasks can use a lighter model. ## Enterprise: restricting mission access Organizations can restrict who is allowed to launch missions through the `missionPolicy` org-level setting: ```json { "missionPolicy": { "restrictedAccess": true, "allowedUserIds": ["user_123", "user_456"] } } ``` When `restrictedAccess` is `true`, only members listed in `allowedUserIds` can start new missions. See the [settings reference](/droid-cli/settings) for the full enterprise policy surface. Run missions non-interactively with `droid exec --mission`. Govern mission access and model policy at the organization level. # Factory Router Automatically select the best model for each task. Factory Router automatically chooses the best model for each Droid task. Instead of locking a session to one model, it evaluates the work in front of it and routes each task to the model with the best balance of quality, latency, and cost. It runs in the Droid CLI and Factory App, under the same Enterprise Controls as every other model. ## Why use Factory Router? Use it when you want strong results without hand-picking a model for every workflow. | Benefit | What Factory Router does | | :------ | :----------------------------------------------- | | **Best performance at lower cost** | In production, it delivers **43%** aggregate cost savings versus pricing the same workload at top-tier model rates, while holding frontier-level performance in Factory evaluations. | | **Improved model selection** | Droid sends routine steps to faster, lower-cost models and reserves stronger models for work that needs deeper reasoning. | | **Enterprise-ready controls** | The same org, project, and user-level model controls used across the Factory platform apply without change. | ## Availability Factory Router is generally available in the Droid CLI and Factory App model selector, with no setup required. In some product surfaces, the same router appears as **Auto Model**; the canonical product and docs name is **Factory Router**. ## Reliability Factory Router routes across multiple models, providers, and capacity sources, so sessions keep running when a provider degrades, hits rate limits, or constrains capacity. Factory evaluations put request reliability at 99.9%+, more than any single-provider path. {/* sweep-allow: term-bullets */} - **Provider failover.** If one provider path is unavailable, sessions continue through another healthy path. - **Dedicated throughput.** Enterprise customers get reserved TPM for critical work instead of relying only on shared public capacity. - **US-hosted open-source models.** Eligible work can route to US-hosted open-source models for cost-efficient or controlled options. ## Enterprise Controls Admins manage Factory Router through the same governance system used for every other Factory-supported model, allowing or restricting it at the organization level like any other option. For more on enterprise model controls, see [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control). ## FAQ It routes per task, combining session-level and per-request decisions, and weighs prompt cache maintenance and the resulting savings heavily in its optimization. If the selected model struggles with a task or a provider degrades, it moves the session to a more capable model or another healthy provider path. You can enable or disable it like any other model today, and Enterprise Controls add governance on top. Yes. On Factory's enterprise engineering benchmarks it reaches 99% of Claude Opus 4.7's pass rate at 20% lower cost per session on Terminal-Bench 2, and 96% of its pass rate at 25% lower cost on [Legacy-Bench](/benchmarks/legacy-bench). In production, measured as billed usage versus the same workload priced at top-tier model rates, aggregate cost is **43%** lower. The median session costs 81% less, and 61% of sessions are at least 80% cheaper, with frontier-level performance throughout. On benchmarks, cost per successful run is about 80% of a Claude Opus 4.7 baseline (80.5% on Terminal-Bench 2, 78.0% on Legacy-Bench). Connect your own API keys, open source providers, or local models. Govern which models are available to your organization. # Custom Models (BYOK) Connect your own API keys, use open source models, or run local models The Droid CLI supports custom model configurations through BYOK (Bring Your Own Key). Use your own OpenAI or Anthropic keys, connect to any open source model providers, or run models locally on your hardware. For Factory-managed models and multipliers, see [Available Models](/models). For org-level restrictions on custom models, base URLs, and provider access, see [Enterprise Controls](/enterprise/hierarchical-settings-and-org-control). Your API keys remain local and are not uploaded to Factory servers. Custom models are available in the Droid CLI and the desktop app, which read your local `settings.json`; they don't appear in Factory's hosted web or mobile platforms. [Install the CLI with the 5-minute quickstart →](/droid-cli/quickstart) --- ## Configuration reference Add custom models to `~/.factory/settings.json` under the `customModels` array: ```json { "customModels": [ { "model": "your-model-id", "displayName": "My Custom Model", "baseUrl": "https://api.provider.com/v1", "apiKey": "${PROVIDER_API_KEY}", "provider": "generic-chat-completion-api", "maxOutputTokens": 16384 } ] } ``` In `settings.json` (and `settings.local.json`), `apiKey` supports environment variable references using `${VAR_NAME}` syntax. For example, `"apiKey": "${PROVIDER_API_KEY}"` reads from the environment variable named `PROVIDER_API_KEY` (for example: `export PROVIDER_API_KEY=your_key_here`). `apiKey` is optional for HTTP custom models. Omit it entirely for a **keyless** endpoint (for example, a gateway that authenticates via `extraHeaders` or network-level trust), or use [`apiKeyHelper`](#dynamic-credentials-with-apikeyhelper) to mint a short-lived token at request time. Some OpenAI-compatible servers require a non-empty `apiKey` even if they ignore its value; in that case, set any placeholder. `baseUrl` is still required for non-Bedrock models. **Legacy support**: Custom models in `~/.factory/config.json` using snake_case field names (`custom_models`, `base_url`, etc.) are still supported for backwards compatibility. Both files are loaded and merged, with `settings.json` taking priority. Environment variable expansion for `apiKey` does not apply to legacy `config.json`. ### Supported fields Model identifier sent via API (e.g., `claude-sonnet-4-5-20250929`, `gpt-5.3-codex`, `qwen3:4b`). Human-friendly name shown in model selector. API endpoint base URL. Your API key for the provider. Optional: omit it for a keyless endpoint, or use `apiKeyHelper` instead. If set, it can't be empty. Supports environment variable references; see the Tip above. Shell command whose stdout is used as the API key, resolved at request time and refreshed on a TTL and after a `401`. Takes precedence over `apiKey`. Only honored from org-managed settings; see [Dynamic credentials](#dynamic-credentials-with-apikeyhelper). Refresh interval for `apiKeyHelper` output, in milliseconds. Defaults to 5 minutes; the `FACTORY_API_KEY_HELPER_TTL_MS` environment variable overrides it. Credential transport for HTTP Anthropic models. Omit it or use `provider-default` to send the credential as `x-api-key`. Use `bearer` to send it as `Authorization: Bearer`. Bearer mode requires `provider: "anthropic"` and isn't supported for Bedrock models. Provider type that determines API compatibility. See the Understanding Providers section below. Maximum output tokens for model responses. Set to `true` to disable image inputs for this model. Additional provider-specific arguments to include in API requests. Additional HTTP headers to send with requests. AWS Bedrock routing options (`awsRegion`, `awsProfile`, `bedrockBaseUrl`, `awsAuthRefresh`, `awsCredentialExport`, `requestMetadata`). See [AWS Bedrock](#aws-bedrock). ### Using extraArgs Pass provider-specific parameters like temperature or top_p by adding an `extraArgs` object: ```json { "extraArgs": { "temperature": 0.7, "top_p": 0.9 } } ``` ### Using extraHeaders Add custom HTTP headers to API requests by adding an `extraHeaders` object: ```json { "extraHeaders": { "X-Custom-Header": "value", "Authorization": "Bearer YOUR_TOKEN" } } ``` ### Dynamic credentials with apiKeyHelper Some gateways issue short-lived tokens (for example, an internal LLM gateway fronted by an identity provider) rather than a long-lived static key. Instead of a static `apiKey`, point the model at an `apiKeyHelper` command whose standard output is the token. Droid runs it at request time, caches the result, and refreshes it automatically. ```json { "customModels": [ { "model": "your-model", "displayName": "Gateway Model", "baseUrl": "https://llm-gateway.internal.example.com/v1", "provider": "generic-chat-completion-api", "apiKeyHelper": "/usr/local/bin/mint-gateway-token.sh", "apiKeyHelperTtlMs": 300000 } ] } ``` Behavior: - The command's trimmed stdout is used as the credential (sent as the provider-appropriate auth header). `apiKeyHelper` takes precedence over any static `apiKey`. - The token is cached and re-minted when it expires. TTL resolution order: the `FACTORY_API_KEY_HELPER_TTL_MS` environment variable, then the per-model `apiKeyHelperTtlMs`, then a 5-minute default. - The token is also refreshed once after a `401` response. If the command fails, it enters a short cooldown and fails fast rather than retrying on every request. - The token, command output, and command text are never logged. `apiKeyHelper` runs a shell command, so it is only honored from **org-managed (trusted) settings**. It is stripped from user, project, and folder settings so an untrusted repository cannot execute commands on your machine. When adding custom models through **org-managed settings**, net-new models cannot use a static `apiKey` — use `apiKeyHelper`, a keyless endpoint, or a Bedrock model instead. Existing static-key models continue to work and can still rotate their key. This mirrors Claude Code's `apiKeyHelper` (the `FACTORY_API_KEY_HELPER_TTL_MS` variable is the analog of `CLAUDE_CODE_API_KEY_HELPER_TTL_MS`). ### Anthropic bearer authentication By default, `provider: "anthropic"` sends the credential from `apiKey` or `apiKeyHelper` in the `x-api-key` header. Some Anthropic Messages-compatible gateways require `Authorization: Bearer` instead. Set `authMode` to `bearer` for those gateways: ```json { "customModels": [ { "model": "claude-compatible-model", "displayName": "Claude-compatible gateway", "baseUrl": "https://llm-gateway.internal.example.com/anthropic", "provider": "anthropic", "apiKeyHelper": "/usr/local/bin/mint-gateway-token.sh", "authMode": "bearer" } ] } ``` `authMode: "bearer"` is only valid for non-Bedrock models using `provider: "anthropic"`. --- ## AWS Bedrock To route an Anthropic or OpenAI custom model through AWS Bedrock, add a `bedrock` object to the model config. Credentials are resolved through the standard AWS SDK provider chain (environment variables, shared config/credentials files, or an SSO/IAM profile). ```json { "customModels": [ { "model": "anthropic.claude-sonnet-4-5-20250929-v1:0", "displayName": "Sonnet 4.5 [Bedrock]", "provider": "anthropic", "apiKey": "not-used-for-bedrock", "bedrock": { "awsRegion": "${AWS_REGION}", "awsProfile": "${AWS_PROFILE}" } } ] } ``` ### Bedrock fields | Field | Type | Description | |-------|------|-------------| | `awsRegion` | `string` | AWS region for Bedrock requests. Optional; when omitted, Droid falls back to the AWS default region chain. Supports `${VAR_NAME}` interpolation. | | `awsProfile` | `string` | AWS profile name to resolve credentials and (as a last resort) a region. Supports `${VAR_NAME}` interpolation. | | `bedrockBaseUrl` | `string` | Explicit Bedrock endpoint override. Accepts a full URL or a `${VAR_NAME}` reference that expands to one. | | `awsAuthRefresh` | `string` | Shell command run to refresh AWS credentials before requests. | | `awsCredentialExport` | `string` | Shell command whose output exports AWS credentials into the request environment. | | `requestMetadata` | `object` | Key-value tags attached to every Bedrock inference call for this model, so Droid usage is attributable in Bedrock model-invocation logs and cost reports. Keys and values support `${VAR_NAME}` interpolation. See [Request metadata](#request-metadata). | ### Request metadata `requestMetadata` is a map of string keys to string values that Droid sends on every Bedrock inference call for that model. AWS surfaces the tags in model-invocation logs and cost allocation, which is how you attribute Bedrock spend in your own AWS account back to Droid, a team, or a cost center: ```json { "customModels": [ { "model": "anthropic.claude-sonnet-4-5-20250929-v1:0", "displayName": "Sonnet 4.5 [Bedrock]", "provider": "anthropic", "apiKey": "not-used-for-bedrock", "bedrock": { "awsRegion": "us-west-2", "requestMetadata": { "tool": "droid", "team": "platform-eng", "cost_center": "${COST_CENTER}" } } } ] } ``` Bedrock carries this parameter differently on each of its wire formats, so what Droid sends depends on the model's `provider`: | `provider` | Bedrock API | How the tags are sent | | --- | --- | --- | | `anthropic` | InvokeModel / InvokeModelWithResponseStream | As the `X-Amzn-Bedrock-Request-Metadata` header, included in the SigV4 signature. | | `bedrock-converse` | Converse / ConverseStream | As the `requestMetadata` field in the request body. | | `openai` | Bedrock's OpenAI-compatible endpoint | Not sent. AWS does not support request metadata there, so the setting has no effect. | AWS constrains the tags, and exceeding a limit fails every call on that model: - At most **16 entries**. - Keys are 1-256 characters; values are at most 256 characters. - Keys and values may contain only alphanumerics, whitespace, and `: _ @ $ # = / + , . -`. Droid checks these limits per model, at request time, on the two APIs that actually send the tags. A violation fails only that model with a message naming the offending key, rather than invalidating your whole settings file. If the block itself cannot be represented at all (not an object of string values, or two keys that collapse onto the same name after `${VAR_NAME}` expansion), Droid logs a warning naming the key, drops `requestMetadata` for that model, and keeps the rest of your settings. Tags are attribution data, not access control. They are sent to AWS on every call, and org-distributed values are readable by every member of the organization. Do not put secrets or credentials in `requestMetadata`. ### Environment variable interpolation Like `apiKey`, the `awsRegion`, `awsProfile`, `bedrockBaseUrl`, and `requestMetadata` fields support `${VAR_NAME}` references in `settings.json`/`settings.local.json`, expanded per workspace at parse time. This lets a single centrally-distributed `customModels` config resolve to the correct value in each environment: ```json { "bedrock": { "awsRegion": "${AWS_REGION}" } } ``` If a referenced variable is not set, Droid fails fast with a clear error naming the missing variable instead of sending the literal `${VAR_NAME}` to AWS. In `requestMetadata`, keys are expanded alongside values, and an unresolved reference is only reported on the APIs that send the tags, so it never fails a model whose dialect ignores them. `awsAuthRefresh` and `awsCredentialExport` are **not** expanded by the settings parser. They run through a shell, so `${VAR}` in those commands is substituted by the shell at run time against the child process environment. ### Region resolution When `awsRegion` is omitted, Droid resolves the region using the AWS default chain, in order: 1. `AWS_REGION` 2. `AWS_DEFAULT_REGION` 3. The `region` set on the resolved AWS profile in `~/.aws/config` If none of these supply a region, Droid fails fast with a clear error rather than silently defaulting, so requests never target a region the workspace did not choose (important for VPC-endpoint region affinity). ## Understanding providers Factory supports three provider types that determine API compatibility: | Provider | API Format | Use For | Documentation | |----------|------------|---------|---------------| | `anthropic` | Anthropic Messages API (v1/messages) | Anthropic models on their official API or compatible proxies | [Anthropic Messages API](https://docs.claude.com/en/api/messages) | | `openai` | OpenAI Responses API | OpenAI models on their official API or compatible proxies. Required for the newest models like GPT-5 and GPT-5-Codex. | [OpenAI Responses API](https://platform.openai.com/docs/api-reference/responses) | | `generic-chat-completion-api` | OpenAI Chat Completions API | OpenRouter, Fireworks, Together AI, Ollama, vLLM, and most open-source providers | [OpenAI Chat Completions API](https://platform.openai.com/docs/api-reference/chat) | Factory is actively verifying Droid's performance on popular models, but we cannot guarantee that all custom models will work out of the box. Only Anthropic and OpenAI models accessed via their official APIs are fully tested and benchmarked. **Model Size Consideration**: Models below 30 billion parameters have shown significantly lower performance on agentic coding tasks. While these smaller models can be useful for experimentation and learning, they are generally not recommended for production coding work or complex software engineering tasks. --- ## Provider reference Every provider below works through the configuration above. Use `provider: "generic-chat-completion-api"` unless you are calling OpenAI's or Anthropic's official API. | Provider | `baseUrl` | Example `model` | | --- | --- | --- | | OpenAI (official) | `https://api.openai.com/v1` | `gpt-5.3-codex` (set `provider: "openai"`) | | Anthropic (official) | `https://api.anthropic.com` | `claude-sonnet-4-5-20250929` (set `provider: "anthropic"`) | | OpenRouter | `https://openrouter.ai/api/v1` | `openai/gpt-oss-20b` | | Fireworks AI | `https://api.fireworks.ai/inference/v1` | `accounts/fireworks/models/deepseek-v3p1-terminus` | | DeepInfra | `https://api.deepinfra.com/v1/openai` | `Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo` | | Groq | `https://api.groq.com/openai/v1` | `moonshotai/kimi-k2-instruct-0905` | | Baseten | `https://inference.baseten.co/v1` | your deployed model ID | | Hugging Face | `https://router.huggingface.co/v1` | `openai/gpt-oss-120b:fireworks-ai` | | Google Gemini | `https://generativelanguage.googleapis.com/v1beta/` | `gemini-2.5-pro` | ### Local models Run models on your own hardware with an OpenAI-compatible server, then point `baseUrl` at it: - **Ollama**: set `baseUrl` to `http://localhost:11434/v1` (no `apiKey` needed, though some builds want any placeholder value), and raise the context window to at least 32k (for example `OLLAMA_CONTEXT_LENGTH=32000 ollama serve`). Models with 30B+ parameters are recommended for agentic coding, and Ollama Cloud models work through the same endpoint. - **LM Studio**: set `baseUrl` to `http://localhost:1234/v1`, load an OpenAI-compatible model, and start its local server. ## Prompt caching The Droid CLI automatically uses prompt caching when available to reduce API costs: - **Official providers (`anthropic`, `openai`)**: Factory attempts to use prompt caching via the official APIs. Caching behavior follows each provider's implementation and requirements. - **Generic providers (`generic-chat-completion-api`)**: Prompt caching support varies by provider and cannot be guaranteed. Some providers may support caching, while others may not. ### Verifying prompt caching To check if prompt caching is working correctly with your custom model: 1. Run a conversation with your custom model 2. Use the `/cost` command in the Droid CLI to view cost breakdowns 3. Look for cache hit rates and savings in the output If you're not seeing expected caching savings, consult your provider's documentation about their prompt caching support and requirements. --- ## Using custom models Once configured, access your custom models in the CLI: 1. Use the `/model` command 2. Your custom models appear in a separate "Custom models" section below Factory-provided models 3. Select any model to start using it Custom models display with the name you set in `displayName`, making it easy to identify different providers and configurations. --- ## Troubleshooting - Check JSON syntax in `~/.factory/settings.json` (or `config.json` if using legacy format) - Settings changes are detected automatically via file watching - Verify all required fields are present - Provider must be exactly `anthropic`, `openai`, or `generic-chat-completion-api` - Check for typos and ensure proper capitalization - Verify your API key is valid and has available credits - Check that the API key has proper permissions - Confirm the base URL matches your provider's documentation - If using `apiKeyHelper`: run the command yourself and confirm it exits `0` and prints a valid token to stdout. A failed helper enters a short cooldown, and `apiKeyHelper` is only honored from org-managed settings. - Ensure your local server is running (e.g., `ollama serve`) - Verify the base URL is correct and includes `/v1/` suffix if required - Check that the model is pulled/available locally - Check your provider's rate limits and usage quotas - Monitor your usage through your provider's dashboard Browse Factory-managed models and multipliers. Restrict custom models, base URLs, and provider access at the org level. # Autonomy Level Choose Off, Low, Medium, or High to control what Droid can do without repeated confirmations. Autonomy Level sets the highest-risk work Droid can run without pausing for approval. It is separate from interaction mode: Normal Mode executes work, while Spec Mode plans before implementation. ## Choose a level Execute commands and MCP tools have a risk level (`low`, `medium`, or `high`). Droid runs them automatically when the risk is at or below your Autonomy Level, unless a denylist or sandbox check requires approval, or the command matches the blocklist. Blocklisted commands never run at any level and have no approval prompt. | Autonomy Level | What can run without approval | Examples | | -------------- | ------------------------------------------------------- | ---------------------------------------------------------------------- | | **Off** | Built-in read tools and allowlisted commands only | `Read`, `LS`, `ls`, `pwd`, `git status` | | **Low** | File edits plus low-risk commands and MCP tools | `Edit`, `Create`, `rg`, showing logs | | **Medium** | Everything from Low plus reversible workspace changes | `npm install`, `pip install`, `git commit`, `mv`, `cp`, build tooling | | **High** | High-risk actions unless safety checks require approval | `docker compose up`, `git push` if allowed, migrations, custom scripts | Droid still streams output and highlights file changes at every level. ## How approvals work Autonomy Level controls automatic approval, not which tools are available. Tool policy, MCP configuration, model support, and organization controls can still restrict tools. {/* sweep-allow: term-bullets */} - **Normal vs. Spec Mode**: In Normal Mode, Autonomy Level controls approvals. Spec Mode is read-only planning; after approval, Droid exits Spec Mode and returns to Normal Mode for implementation. - **File changes**: Low or higher lets Droid create, edit, and patch files without asking first. - **Commands and MCP tools**: Droid compares the tool risk level to your Autonomy Level. If the risk is higher, it asks before continuing. - **Allowlisted commands**: Commands in the allowlist can run without approval unless they also match the denylist or blocklist. - **Safety checks**: Denylisted dangerous commands still ask at High, including dangerous commands nested inside `$(...)` or backticks. Blocklisted commands are rejected outright at every level, with no approval prompt. [Sandbox](/autonomy-and-safety/sandbox) read, write, and network checks can also prompt separately. - **Allow always**: Choosing an "always allow" option raises the current Autonomy Level to the level required by that prompt. Sandbox "allow always" options instead persist the allowed path or domain. - **Spec approval**: When approving a Spec Mode plan, choose **Proceed with implementation** to continue with manual approvals, or choose an available Low, Medium, or High option for implementation. Organization Maximum Autonomy Level can hide higher options. ## Command allowlists, denylists, and blocklists Use `commandAllowlist`, `commandDenylist`, and `commandBlocklist` in [Settings](/droid-cli/settings) to encode command policy for your user profile, a project, a local project override, or a nested folder. - Allowlist entries are treated as low-risk for the matching scope. - Denylist entries always take precedence over allowlist entries. A denied command still runs if you explicitly approve it. - Blocklist entries can never run: there is no approval prompt, and the block holds even under full autonomy, auto-run, or `--skip-permissions-unsafe`. Use it for commands that must be hard-stopped regardless of approvals. - Commands not covered by any list fall back to the active Autonomy Level and command restrictions. - Organization-managed settings have the highest priority. Local and project settings can add defaults for a repo or machine, but they cannot weaken organization command policy or raise autonomy above the organization maximum. See [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control). The built-in denylist covers common destructive patterns such as filesystem wipes, disk formatting, shutdown commands, fork bombs, and broad permission or ownership changes. Add project-specific commands when your repo has additional dangerous scripts or deployment paths. Because the blocklist cannot be bypassed by approval, it is the strongest control available. droid resolves the actual program being invoked before matching, so a blocked command cannot be slipped through with a wrapper shell (for example `bash -c "…"`), an absolute path, quoting tricks, or command substitution. ## Change the level - Press to cycle `Off → Low → Medium → High → Off`. Organization policy can cap the highest available level. - Press to switch between Normal Mode and Spec Mode. - Set a default in `/settings` for future sessions. - Change Autonomy Level before implementation, from the [Spec Mode](/autonomy-and-safety/specification-mode) approval dialog, or any time after leaving Spec Mode. ## Where Autonomy Level applies {/* sweep-allow: term-bullets */} - **Interactive CLI**: `droid` uses the session's current Autonomy Level. `droid ""` starts the same interactive CLI with an initial prompt, so the first task uses your configured default. See the [CLI reference](/droid-cli/cli-reference). - **Factory App**: Factory App sessions use the same Normal, Spec, Mission, and Autonomy Level controls as CLI sessions. - **Droid Exec**: `droid exec` is read-only by default. Use `--auto low`, `--auto medium`, or `--auto high` for non-interactive runs that need edits, local development commands, or broader automation. See [Droid Exec](/droid-exec/overview). - **Custom Droids (Subagents)**: Task-launched subagents inherit the parent session's Autonomy Level by default, or use the **Subagent autonomy level** setting (`inherit`, `off`, `low`, `medium`, `high`) when configured, always clamped to the organization's Maximum Autonomy Level. In Spec Mode they are read-only. Organization and Droid tool policy can still restrict them. See [Custom Droids](/harness/subagents). - **Factory Missions**: Missions orchestration requires High autonomy or `--skip-permissions-unsafe` (unsafe: skips all permission checks; use only in isolated sandboxes), and admins can restrict who can start Missions. See [Factory Missions](/missions/overview). ## Enterprise Controls Enterprise admins can set organization-wide autonomy boundaries with organization-managed settings. See [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control). - **Default Autonomy Level** sets the starting level for new sessions. - **Maximum Autonomy Level** caps how high members can raise autonomy. If the maximum is Medium, High is unavailable in the CLI. These controls layer with command allowlists, command denylists, MCP restrictions, sandbox settings, and Missions access controls. ## Use it safely - Start new or high-stakes work with Off or Low until you trust the plan. - Match the minimum level to the work: use Low for file edits and generated reports, Medium when the run must install dependencies, build, test, or make local commits, and High for pushes, deployments, Task-launched subagents, Missions, or other orchestration. - Add defense in depth with [blocking hooks](/harness/hooks), command denylists, MCP restrictions, least-privilege credentials, and isolated runners. - For CI workflows, choose the lowest `droid exec --auto` level that allows the workflow to complete. See [Automated Code Review](/software-factory/code-review-ci) and [Droid Exec](/droid-exec/overview). - If you spot a suspect command, interrupt, provide guidance, and resume at the Autonomy Level that fits the remaining risk. Understand Normal, Spec, and Mission Mode and how they relate to autonomy. Enforce deterministic policy and validation at every tool call. # Interaction Modes Choose Normal, Spec, or Mission Mode, then set Autonomy Level for how much Droid can do without approval. Interaction Modes set the shape of a session. **Normal Mode** works directly, **Spec Mode** plans before implementation, and Mission Mode coordinates larger work with an orchestrator. Autonomy Level is separate. It controls which actions Droid can run without pausing for approval. Some CLI and IDE surfaces combine the labels as `Auto (Off)`, `Auto (Low)`, `Auto (Medium)`, or `Auto (High)`, which means Normal Mode plus the selected Autonomy Level. ## Choose a mode | Mode | Use it when | What Droid can do | How to enter | | ---- | ----------- | ----------------- | ------------ | | **Normal Mode** | You want Droid to answer, edit, run commands, or implement now. | Uses the enabled tools for the session. Approval prompts are governed by Autonomy Level, command policy, sandbox rules, and org controls. | Default mode. Select **Normal Mode** from the mode menu, or press from Spec Mode in the CLI. | | **Spec Mode** | You want Droid to investigate and propose a plan before any code changes. | Uses read-only planning behavior, then calls `ExitSpecMode` to ask for approval. | Select **Spec Mode**, press from Normal Mode in the CLI, or start with `--use-spec`. | | **Mission Mode** | You want Droid to coordinate a larger effort with worker agents and validation. | Runs a mission orchestrator session with Mission Control. Mission workers use their own configured model and autonomy settings. | Start from the [Missions](/missions/overview) workflow, or use `droid exec --mission --auto high`. | In the CLI, toggles between Normal Mode and Spec Mode. Mission Mode is not part of the cycle. ## Set Autonomy Level separately Autonomy Level applies in Normal Mode and after you approve a Spec Mode plan. | Autonomy Level | What can run without approval | | -------------- | ----------------------------- | | **Off** | Built-in read tools and allowlisted commands. Other actions ask first. | | **Low** | File edits plus low-risk commands and MCP tools. | | **Medium** | Everything in Low plus reversible workspace changes such as installs, builds, and local commits. | | **High** | High-risk actions, unless a blocklist, denylist, sandbox rule, or org control requires a stop or prompt. | Press in the CLI to cycle `Off → Low → Medium → High → Off`. Organization policy can cap the highest available level. See [Autonomy Level](/autonomy-and-safety/auto-run) for the full approval model. ## Use Spec Mode Spec Mode is for research and planning before implementation. Use it for architecture changes, migrations, security-sensitive work, or any task where you want to review the plan before Droid edits files. Select **Spec Mode** from the mode menu, press in the CLI, or start a session with `--use-spec`. Explain what should change, the constraints that matter, and how Droid should verify the work. Droid researches the repo, reads relevant files, and proposes a concrete implementation plan. Approve the plan to return to Normal Mode for implementation, choose a higher Autonomy Level for implementation, or keep iterating in Spec Mode. During Spec Mode, Droid should not edit files, change configuration, make commits, start services, or write to external systems. It can read files, search the repo, inspect linked artifacts, and ask clarifying questions. ### Make the request specific Include the outcome, constraints, verification steps, and relevant existing patterns: ```text Users need to reset passwords using email verification. The reset link should expire after 24 hours. Include rate limiting and tests for invalid or expired links. Follow the background job pattern used by report generation. ``` ### After the plan After Droid proposes a plan, choose one of the approval options: proceed with manual approvals, proceed with Low/Medium/High Autonomy Level, or keep iterating on the spec. Organization Maximum Autonomy Level can hide higher approval options. ## Use Mission Mode for orchestrated work Mission Mode is for multi-step work that benefits from orchestration, worker agents, validation, and progress tracking. Use Mission Mode when a task is too large for a single linear session, or when you want explicit validation milestones. Mission Mode is not a planning toggle. Starting a Mission upgrades the session into an orchestrator workflow, and Mission workers run with the model, reasoning, and autonomy settings configured for Missions. See [Factory Missions](/missions/overview). ## Change controls and defaults - Press to switch between Normal Mode and Spec Mode. - Press to cycle Autonomy Level. - Use `/settings` to change session defaults. - Use `/model` and the Spec Mode model setting when planning should use a different model. - In the Factory App, use the mode selector for **Normal Mode**, **Spec Mode**, or **Mission Mode**, and the Autonomy selector for `Auto Off`, `Auto Low`, `Auto Medium`, or `Auto High`. ```bash droid --use-spec droid --auto medium droid exec --use-spec "Plan the migration, then ask for approval" droid exec --auto high "Run the approved release checklist" droid exec --mission --auto high "Coordinate the migration" ``` Set defaults in [settings.json](/droid-cli/settings): ```json { "sessionDefaultSettings": { "interactionMode": "spec", "autonomyLevel": "low", "specModeModel": "", "specModeReasoningEffort": "high" } } ``` Use: - `sessionDefaultSettings.interactionMode`: `auto` for Normal Mode or `spec` for Spec Mode. - `sessionDefaultSettings.autonomyLevel`: `off`, `low`, `medium`, or `high`. - `sessionDefaultSettings.specModeModel`: optional planning model for Spec Mode. - `sessionDefaultSettings.specModeReasoningEffort`: optional reasoning effort for the Spec Mode model. - `sessionDefaultSettings.autonomyMode`: deprecated legacy field. Prefer `interactionMode` plus `autonomyLevel`. Enterprise administrators can set `maxAutonomyLevel` to cap available autonomy. See [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control). ## Save Spec Mode plans Spec Mode can save approved plans as Markdown. Open the CLI settings and enable **Save spec as Markdown**. - By default, plans are saved to `.factory/docs` inside the nearest project-level `.factory` directory. If none exists, the CLI falls back to `~/.factory/docs`. - Use **Spec save directory** to choose a project directory, your home directory, or a custom path. - Custom values support absolute paths, `~` expansion, `.factory/...` shortcuts, and relative paths from the current workspace. - Files are named `YYYY-MM-DD-slug.md`, with a counter appended when needed. Full settings reference. Run non-interactive sessions with `--auto`, `--use-spec`, and `--spec-model`. # Droid Shield Automatic secret detection to prevent accidental exposure of credentials in your git commits and pushes. ## What is Droid Shield? Droid Shield is a built-in security feature that automatically scans un-committed changes for potential secrets before committing and pushing them to remote. It is a safety net that prevents accidental exposure of sensitive credentials like API keys, tokens, and passwords in your version control history. Organization admins can enforce Droid Shield through [Enterprise Controls](/enterprise/hierarchical-settings-and-org-control). For the broader enterprise safety model, see [Agent Safety & Controls](/enterprise/llm-safety-and-agent-controls). ## How Droid Shield works When you use Droid to perform `git commit` or `git push` operations, Droid Shield automatically: 1. **Scans the diff** - Analyzes only the lines being added (not removed or unchanged) 2. **Detects secrets** - Uses pattern matching to identify potential credentials 3. **Blocks execution** - Stops the git operation if secrets are detected 4. **Reports findings** - Shows exactly where potential secrets were found Droid Shield only scans git operations performed through Droid. Manual git commands run outside of Droid are not affected. --- ## What Droid Shield detects Droid Shield scans for a wide range of credential patterns, including: Factory API keys, GitHub tokens, GitLab tokens, npm tokens, and API keys from (e.g. AWS, Google Cloud, Stripe, SendGrid) and more. (e.g. JWT, OAuth, session tokens), and URLs with embedded credentials. (e.g. SSH private keys, PGP keys, age secret keys, OpenSSH keys), and other cryptographic key formats. (e.g. Slack webhooks and tokens, Twilio credentials, Mailchimp keys, Square OAuth secrets, Azure storage keys). ### Detection algorithm Droid Shield uses smart pattern matching with randomness validation: {/* sweep-allow: term-bullets */} - **Pattern matching** - Identifies credentials by format - **Randomness check** - Validates that captured values look like actual secrets - **Context awareness** - Considers variable names and assignment patterns to reduce false positives --- ## Droid Shield 2.0: learned secret detection Droid Shield 2.0 augments the deterministic scanner with two fine-tuned classification models that add a semantic layer on top of pattern-based detection. Every commit and push still passes through the deterministic scanner first; the models then classify specific contexts the scanner flagged or missed. Droid Shield 2.0 is in Private Preview. To enable it for your organization, contact [support@factory.ai](mailto:support@factory.ai). ### The two classification models The models sit on opposite sides of the deterministic scanner and are optimized for distinct failure modes: Runs on changed lines the deterministic scanner did **not** fire on, but that still look secret-bearing based on surrounding context. Optimized for recall to catch real secrets the pattern set would miss. Produces a warning when it flags a possible missed secret. Runs on lines where the scanner **did** fire. It reviews the surrounding context with the detected secret masked out, and decides whether the scanner hit is a true positive (keep blocking) or a false alarm (downgrade to a warning). --- ## When Droid Shield activates Droid Shield automatically activates during these git operations: - **`git commit`** - Scans staged changes before creating the commit - **`git push`** - Scans commits that would be pushed to the remote If secrets are detected, the git operation is blocked to prevent credential exposure. You'll need to remove the secrets before proceeding. --- ## Managing Droid Shield settings ### In the CLI You can toggle Droid Shield on or off through the settings menu: 1. Run `droid` 2. Enter `/settings` 3. Toggle **"Droid Shield"** setting 4. Changes take effect immediately Droid Shield is **enabled by default** for your protection. We strongly recommend keeping it enabled. --- ## What to do if secrets are detected When Droid Shield detects potential secrets, you'll see an error message like: ```text Droid-Shield has detected potential secrets in 2 location(s) across files: src/config.ts, .env.example If you would like to override, you can either: 1. Perform the commit/push yourself manually 2. Disable Droid Shield by running /settings and toggling the "Droid Shield" option ``` ### Recommended actions Carefully examine the files and lines mentioned to identify what was detected. - Use environment variables instead of hardcoded credentials - Move secrets to secure credential stores - Add sensitive files to `.gitignore` - Use git filter-branch or BFG Repo-Cleaner if secrets were already committed Once secrets are removed, run the git command again through Droid. **Never disable Droid Shield just to bypass the check.** Exposed credentials can lead to security breaches, unauthorized access, and compliance violations. ### If you get a false positive Droid Shield uses conservative patterns to err on the side of caution. If you believe a detection is a false positive: 1. **Verify it's not a real secret** - Double-check that the value isn't sensitive 2. **Use a manual commit** - Perform the git operation yourself outside of Droid 3. **Report the pattern** - Contact [support@factory.ai](mailto:support@factory.ai) if you encounter recurring false positives When Droid Shield 2.0 is enabled, the downgrade model automatically reviews scanner hits and can clear obvious false alarms (placeholders, examples, test fixtures) without manual intervention. --- ## Best practices ### Use environment variables Store all secrets in environment variables or secure credential managers, never hardcode them in source files. ```javascript // Good - Using environment variable const apiKey = process.env.FACTORY_API_KEY; // Bad - Hardcoded secret const apiKey = "never-hardcode-secrets"; ``` ### Keep Droid Shield enabled Droid Shield provides an essential safety layer. Keep it enabled at all times, especially in team environments. ### Review before committing Even with Droid Shield, manually review your changes before committing to ensure no sensitive data is included. ### Educate your team Make sure all team members understand how Droid Shield works and why it's important to keep it enabled. --- ## Limitations **Droid Shield is a detection tool, not a guarantee.** While it catches many common secret patterns, the deterministic scanner alone cannot detect: - Custom secret formats not in the pattern database - Secrets that don't follow recognizable patterns - Obfuscated or encoded credentials - Business logic vulnerabilities or code security issues --- Learn about Factory's comprehensive security features and best practices. Configure Droid settings including Droid Shield preferences. Classification models, training data, and results. How Droid Shield fits into the broader enterprise safety model. Email the security team at security@factory.ai. Contact support@factory.ai about persistent false positive patterns. # Sandbox OS-level sandboxing isolates Droid from your filesystem and network using kernel-enforced policies. OS-level sandboxes let users set filesystem and network boundaries for Droid. All shell commands initiated by Droid run in a separate process that is limited to the filesystem and network boundaries configured by users and enforced at the OS kernel level. ## How OS-level isolation works The sandbox enforces its filesystem and network boundaries with the same operating-system primitives the OS uses to confine untrusted software, so a blocked read, write, or connection is denied by the kernel rather than relying on Droid to police itself: {/* sweep-allow: term-bullets */} - **macOS**: Seatbelt (the built-in macOS sandbox) profiles restrict filesystem and process access. - **Linux and WSL2**: [bubblewrap](https://github.com/containers/bubblewrap) with a seccomp filter provides the same confinement. - **Network (all platforms)**: outbound traffic is routed through an HTTP/SOCKS filtering proxy that permits only the allowed domains (Factory's own domains and WorkOS authentication domains are always allowed). Because the sandbox relies on these host primitives, it blocks rather than running without isolation when they are unavailable. ## Isolation modes The `sandbox.mode` setting selects *what* is placed inside the OS boundary. Both modes enforce the same default policies and configuration; they differ only in scope. | Aspect | `per-command` (default) | `whole-process` | | --------------------------------- | ----------------------------------------------- | ------------------------------------ | | **Runs inside the OS sandbox** | Each shell command and its child processes | The entire Droid process | | **MCP servers, hooks, subagents** | Wrapped or checked against your policy per call | Isolated inside the process boundary | | **Main Droid process** | Not isolated | Isolated | | **If isolation is unavailable** | The affected command or hook is blocked | Droid refuses to start | - **`per-command`** runs each Droid-initiated action through the sandbox individually. Shell commands (and their child processes) are confined at the OS level; other tools are mediated by policy checks before each call. The main Droid process itself is not isolated. - **`whole-process`** launches the entire Droid process inside the OS sandbox, so the main process and everything it spawns (MCP transports, subagents) is isolated too. Its own network requests (not just those from the Execute tool) are filtered against `allowedDomains`, with interactive domain prompts in TUI mode. If the sandbox cannot be established at startup (unsupported platform or a failed isolation check), Droid refuses to start rather than running unsandboxed. ## Enable and configure the sandbox Set `sandbox.enabled` to `true` in your settings to turn on the sandbox: ```jsonc { "sandbox": { "enabled": true, }, } ``` Once enabled, the default policies below apply immediately. No other configuration is required to start. Settings resolve across the hierarchy (org > project > user), so you can enable the sandbox for yourself (user settings), a repository (project settings), or everyone (organization settings). ### Default access policies | Resource | Default policy | Configurable via | | --------------- | ------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------- | | **File reads** | Allow all. Only explicit `denyRead` entries are blocked. | `sandbox.filesystem.denyRead` | | **File writes** | Deny all except **CWD** (current working directory). Additional paths can be allowed. `denyWrite` overrides `allowWrite`. | `sandbox.filesystem.allowWrite`, `sandbox.filesystem.denyWrite` | | **Network** | Deny all except Factory's own domains and WorkOS authentication domains (always allowed by default). Additional domains must be explicitly allowed. | `sandbox.network.allowedDomains` | ### Full settings reference ```jsonc { "sandbox": { "enabled": true, // Isolation scope: "per-command" (default) or "whole-process" "mode": "per-command", "filesystem": { // Additional writable paths beyond CWD (which is always writable) "allowWrite": ["/tmp/build-output", "~/.config"], // Deny writes to specific subpaths even if parent is in allowWrite "denyWrite": ["/tmp/build-output/cache/locks", "~/.config/secrets"], // Block reads to specific paths (everything else is readable) "denyRead": ["~/.aws/credentials", "~/.ssh/id_rsa"], }, "network": { // Only these domains are reachable (Factory's own domains and WorkOS authentication domains always included) "allowedDomains": ["github.com", "*.npmjs.org"], }, }, } ``` Settings merge across the hierarchy (org > project > user). `denyWrite`/`denyRead` use union merge: org denies cannot be removed downstream. ### Organization controls - Org-level `denyWrite`/`denyRead` settings cannot be overridden by user **Allow always**. - The violation prompt shows `(organization policy)` when the deny comes from org settings. - Admins can set the sandbox isolation mode (`per-command` or `whole-process`) org-wide from [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control). ## Coverage and enforcement With the sandbox enabled, a tool runs only if the sandbox can check everything it might do against your filesystem and network rules. Tools whose actions can't be checked are blocked. | Surface | Enforcement | | :------ | :---------- | | File tools | `Read`, `Edit`, `Create`, `LS`, `Grep`, `Glob`, and `ApplyPatch` are checked before every operation, enforcing `denyRead` for reads and `allowWrite` plus `denyWrite` for writes. | | Execute tool | Shell commands run inside the OS sandbox, with network traffic routed through a filtering proxy for domain-level control. | | FetchUrl and WebSearch | Network access is checked against `sandbox.network.allowedDomains`. | | MCP tools | Filesystem and network requests are checked; locally launched servers receive a minimal environment (ambient host environment variables are dropped, keeping only a safe operational allowlist plus keys set in `mcp.json`). | | Subagents (Task tool) | Delegated subagents inherit the parent sandbox policy. | | Hooks | Hook commands (`PreToolUse`, `PostToolUse`, and the rest) are wrapped by the sandbox and run with the same proxy and runtime environment as the Execute tool; if the sandbox cannot be established, the hook is blocked instead of running on the host. | ### When something is blocked **Interactive permission prompts (TUI mode):** - Sandbox violations interrupt the agent loop with a TUI prompt, even at Auto (High) autonomy. - Options: **Allow once**, **Allow always** (persists to settings), **Deny**. For `denyRead`/`denyWrite` violations, "Remove from deny list" replaces "Allow always". - Execute network violations show a real-time domain prompt with a 60s auto-deny timeout. **Non-interactive mode (`droid exec`):** - Sandbox violations are auto-denied without prompting: no hang, no user interaction required. - The agent receives a denial message and reports it in the output. **Allow-always persistence:** - Choosing **Allow always** saves the exception to your user settings (for example, adding a domain to `allowedDomains` or a path to `allowWrite`) so it won't prompt again. - Changes take effect immediately in the current session. **TUI indicators:** - `SANDBOX` status indicator in the footer when the sandbox is enabled. - "Sandbox Violation" prompt with violation details (path, domain, reason). ## Security limitations Sandboxing reduces the impact of a mistake or a prompt-injection attack, but it does not eliminate risk: {/* sweep-allow: term-bullets */} - **Allowed network egress can still leak data.** Every domain in `allowedDomains` is a channel through which data the agent can read might leave. Keep the allowlist as narrow as your workflow permits. - **A writable path can still be modified.** The sandbox limits *where* writes land, not *what* is written within the allowed paths, including your own project code under the working directory. - **Your settings files are protected.** Sandboxed tool processes cannot write to Droid's own configuration files (for example `settings.json`), so they cannot silently change sandbox policy. Treat the sandbox as defense in depth alongside [Autonomy Level](/autonomy-and-safety/auto-run) approvals and human review, not as a hard boundary around untrusted code. To evaluate code you do not trust, run Droid inside a container or a dedicated virtual machine. Where `sandbox.*` lives. How org policy merges with user settings. Broader security model for the Droid CLI. Approval tiers that complement sandboxing as defense in depth. # AGENTS.md Give Droid durable project instructions, commands, and guardrails with AGENTS.md. `AGENTS.md` gives Droid the project briefing it should carry into every session: how to install, run, test, edit, verify, and stay inside your repository's boundaries. Use it for guidance that should be loaded before Droid writes code. Keep it short, specific, and easy to verify. ## Create your first AGENTS.md Add `AGENTS.md` at the repository root. Start with one file, then add nested files only where a package, app, or service needs different rules. List the exact install, development, test, type check, lint, and build commands Droid should run. Add flags or package filters when broad commands are too expensive. Capture repository layout, generated-file rules, security boundaries, ownership boundaries, and required proof before Droid calls work finished. Project instructions belong in git so every teammate and automation session gets the same operating context. ```markdown title="AGENTS.md" # Repository guide ## Commands - Install dependencies: `pnpm install` - Start local server: `pnpm dev` - Run unit tests: `pnpm test` - Run type checks: `pnpm typecheck` - Run lint: `pnpm lint` ## Project layout - `app/` contains routes and server-rendered pages. - `components/` contains reusable UI components. - `lib/` contains shared business logic and integrations. - Tests live next to the code they cover as `*.test.ts` or `*.test.tsx`. ## Coding conventions - Prefer small, focused modules with explicit types. - Keep business logic in pure functions where possible. - Match existing component and file naming patterns. - Do not add a new dependency unless the existing stack cannot solve the problem. ## Verification Before calling work finished, run: 1. `pnpm test` 2. `pnpm typecheck` 3. `pnpm lint` If a command fails, fix the issue and rerun it. ``` Keep human onboarding, screenshots, and contributor background in `README.md`. Put Droid-specific commands, guardrails, and completion criteria in `AGENTS.md`. ## What to include | Section | What good looks like | | :------ | :------------------- | | Project overview | One short paragraph explaining the product, architecture, or package. | | Commands | Exact install, local run, test, lint, type check, build, and regeneration commands. | | Repository map | Directories Droid must understand before editing. | | Conventions | Project-specific patterns that are not obvious from code. | | Testing rules | Focused test commands, full gate commands, and when to add tests. | | Generated files | Which files are generated, what source to edit, and how to regenerate. | | Security rules | Secret handling, production boundaries, approval rules, and migration safety. | | PR expectations | Required evidence, screenshots, changelog notes, or review checks. | Prefer concrete rules that change how Droid works: ```markdown ## Testing - Run `pnpm test -- path/to/file.test.ts` for focused changes. - Run `pnpm test` before completing broad refactors. - Add a regression test for every bug fix when practical. ``` Avoid vague rules that cannot be checked: ```markdown ## Testing - Be careful. - Make sure everything works. ``` ## Choose the right instruction surface Use the narrowest durable surface for the job. | Surface | Use for | How Droid treats it | | :------ | :------ | :------------------ | | `AGENTS.md` and compatible names | Always-on coding instructions, commands, repository conventions, and safety rules. | Loaded as coding guidelines at startup and during dynamic discovery. | | `DESIGN.md`, `Design.md`, and `design.md` | Always-on design-system, UX, visual, and interaction guidance. | Loaded separately as design guidelines. | | `SKILL.md` | Reusable workflows, checklists, or domain expertise that should load only when relevant. | Discovered as a skill. The full body loads when invoked. | | `.factory/commands/*` | Simple user-invoked prompt or executable shortcuts. | Exposed as custom slash commands. | | `.factory/settings.json` | Droid preferences, defaults, and policy settings. | Parsed as configuration, not as prose instructions. | Do not paste long run books into `AGENTS.md`. Link to the source of truth, or move reusable procedures into [skills](/harness/skills) so the extra context loads only when needed. ## Discovery and precedence When a git root is present, Droid searches from the current working directory up to the git root. At each level, it checks the directory itself and these context directories: | Context directory | Common use | | :---------------- | :--------- | | `.factory/` | Factory-specific project instructions. | | `.agents/` | Compatibility with shared agent folder conventions. | | `.agent/` | Compatibility with singular agent folder conventions. | Droid also checks personal instruction directories in your home folder: | Personal directory | Common use | | :----------------- | :--------- | | `~/.factory/` | Your default Droid preferences across projects. | | `~/.agents/` | Compatible personal instructions. | | `~/.agent/` | Compatible personal instructions. | When multiple instruction files apply, use this mental model: - The current user request takes priority over standing instructions. - Nested project files refine root project files for a specific directory tree. - Project files should override personal defaults. - Personal files should describe preferences, not requirements that fight repository rules. ## Compatible filenames Use `AGENTS.md` for new guidance. Droid also reads compatible filenames so existing instruction files from other coding tools keep working. | Filename | Use it when | | :------- | :---------- | | `AGENTS.md` | Recommended for new project guidance. | | `agents.md` | Lowercase compatibility variant. | | `Agents.md` | Title-case compatibility variant. | | `CLAUDE.md` | Compatibility with existing Claude-oriented instructions. | | `Claude.md` | Title-case compatibility variant. | This is a compatibility set, not a setup checklist. Do not create every file. Duplicating rules across supported filenames spends context without adding signal. ## Nested instructions Nested files are useful when a directory tree has different commands, conventions, generated assets, or ownership rules. ```text repository/ AGENTS.md apps/ web/ AGENTS.md src/ page.tsx ``` If Droid starts at `repository/`, it loads the root guidance. When Droid later reads files under `apps/web/`, it can discover `apps/web/AGENTS.md` and apply the more specific guidance to that directory tree. Use nested files for real differences: - A package uses another package manager or test runner. - A service has generated code that must not be edited by hand. - A frontend area has design-system rules that do not apply elsewhere. - A deployment or migration path requires extra approval. Do not use nested files to repeat root rules. ## Context budget Instruction files count against the session context. Keep the top of each file dense with the rules Droid most needs. | Load phase | Limit | | :--------- | ----: | | Initial AGENTS-style guideline load | 80,000 characters | | Dynamic Read-path guideline discovery | 40,000 characters | These are caps, not targets. Smaller files are usually better. ## Root template ```markdown title="AGENTS.md" # Repository guide ## What this project is One or two sentences describing the product, service, or package. ## Commands - Install: `REPLACE_ME` - Run locally: `REPLACE_ME` - Run tests: `REPLACE_ME` - Run type checks: `REPLACE_ME` - Run lint: `REPLACE_ME` - Build: `REPLACE_ME` ## Project layout - `REPLACE_ME/` contains ... - `REPLACE_ME/` contains ... - `REPLACE_ME/` contains ... ## Conventions - Match existing patterns before introducing new ones. - Keep changes scoped to the requested task. - Prefer existing utilities and dependencies. - Add or update tests when changing behavior. ## Verification Before completing work, run the relevant focused checks and then the required project checks: 1. `REPLACE_ME` 2. `REPLACE_ME` 3. `REPLACE_ME` ## Safety - Do not commit secrets or credentials. - Do not run destructive commands without explicit approval. - Do not modify generated files by hand. Regenerate them with `REPLACE_ME`. ``` ## Nested template ```markdown title="apps/web/AGENTS.md" # Subproject guide This file applies when working inside this directory tree. Follow the root `AGENTS.md` unless this file gives a more specific rule. ## Local commands - Run this package's tests: `REPLACE_ME` - Run this package's lint: `REPLACE_ME` - Regenerate local artifacts: `REPLACE_ME` ## Local conventions - Use `REPLACE_ME` for ... - Do not use `REPLACE_ME`; this package uses `REPLACE_ME` instead. - Keep public exports in `REPLACE_ME`. ## Verification For changes in this directory tree, run: 1. `REPLACE_ME` 2. `REPLACE_ME` ``` ## What to avoid Do not use `AGENTS.md` as a dumping ground. Avoid: - Secrets, credentials, private tokens, or private host names. - Full copies of long docs that already live elsewhere. - Stale directory inventories that drift from the repository. - Personal preferences that conflict with project rules. - Temporary task notes that belong in an issue, PR, or spec. - Instructions to skip validation without a narrow, explicit reason. If guidance is useful for one workflow, make it a skill. If it is useful for one project area, put it in a nested `AGENTS.md`. ## Troubleshooting Make the rule concrete. Replace guidance like "test your work" with exact commands, paths, and completion criteria. Put the most important rules near the top of the file. Confirm that the file uses a supported filename and lives in the directory tree Droid inspected. If the session started higher in the repository, Droid may discover nested guidance only after it reads files in that directory tree. Rewrite personal guidance as defaults. Repository instructions and explicit user requests should be able to override personal preferences. Keep only durable, high-signal rules in `AGENTS.md`. Link to long reference docs, move reusable procedures into skills, and split subsystem guidance into nested files. Configure Droid behavior, model defaults, autonomy, and local preferences. The terminal surface that reads AGENTS.md context into every session. # Connectors Connect third-party apps so Droid can use their tools in sessions, with managed authentication and organization-level controls. Connectors give Droid managed access to third-party apps such as GitHub, Linear, Notion, Slack, Sentry, and Google Workspace. Connect an app once, then ask Droid to work with it in natural language. Droid discovers the relevant tools and calls them under your current [Autonomy Level](/autonomy-and-safety/auto-run). Use Connectors when the app appears in Factory's catalog and you want guided authentication with no server configuration. Use [MCP](/harness/mcp) when you need a custom tool, a self-hosted service, or direct control over the server and transport. ## Connect an app Open **Settings → Connectors** in the Factory App. Search the catalog or browse the featured and full integration lists. If your organization manages connector availability, only apps an Owner or Manager enabled appear. Select **Connect**. Factory opens the app's authorization flow in a new browser tab. Sign in and approve the access requested by the app. Return to your session and ask Droid to use the app. Droid discovers the connected tools when the task needs them. In an interactive session, ask Droid to connect an app: ```text Connect my Linear account so you can update the issue for this branch. ``` Droid returns an authorization link when the app is available. Open the link, finish the provider's authorization flow, then tell Droid to continue. Chat-based connection requires an interactive session because you must open an authorization link. Connect the app in the Factory App before running a non-interactive `droid exec` workflow. Connections belong to both your Factory user and the active organization. If you belong to more than one organization, connect the app separately in each organization where Droid needs access. ## Use a connected app Ask for the outcome. You do not need to know tool names or API schemas. ```text Summarize the open Linear issues assigned to me, then update the one for this branch with the implementation status. ``` ```text Read the latest Sentry errors for checkout, correlate them with this code, and propose a fix. ``` ```text Create a Notion page with the rollout plan from this repository and include links to the relevant GitHub pull requests. ``` Droid keeps connector tools in the deferred tool catalog. It loads the matching tool definitions only when a task needs them, which keeps unrelated app schemas out of the session context. Connected tools can also become available during a running session. Connector calls follow the same approval model as other external tools: | Tool classification | Default behavior | | :------------------ | :--------------- | | Read-only | Classified as low risk. It runs automatically when the session's Autonomy Level allows low-risk actions. | | Write, destructive, or unknown | Classified as high risk. Droid asks for confirmation unless the session's Autonomy Level allows high-risk actions. | Connector approvals apply to the current action. They are not stored as [persistent MCP permissions](/harness/mcp#persistent-tool-permissions). ## How connectors work ```mermaid %%{init: { "theme": "base", "themeVariables": { "fontFamily": "Geist Mono, monospace", "fontSize": "14px", "primaryColor": "#161413", "primaryBorderColor": "#342F2D", "primaryTextColor": "#FAFAFA", "secondaryColor": "#101010", "secondaryBorderColor": "#4D4947", "secondaryTextColor": "#D6D3D2", "tertiaryColor": "#020202", "tertiaryBorderColor": "#342F2D", "tertiaryTextColor": "#8A8380", "lineColor": "#4D4947", "textColor": "#D6D3D2", "mainBkg": "#161413", "nodeBorder": "#342F2D", "edgeLabelBackground": "#020202", "clusterBkg": "#0C0A0A", "clusterBorder": "#342F2D", "activationBkgColor": "#EE6018", "activationBorderColor": "#EE6018" } }}%% flowchart LR U[Engineer] O[App authorization] C[Factory connector service] D[Droid] A[Third-party app] U -->|Connect for this organization| O O -->|Authorize| C D -->|Call an approved tool| C C -->|API request| A A -->|Tool result| C C -->|Result| D ``` The connection and execution paths stay separate: | Stage | What happens | | :---- | :----------- | | Availability | Factory maintains a curated connector catalog. Enterprise organizations can choose which catalog entries their members may use. | | Connection | You finish the app's authorization flow. The resulting connection is scoped to your Factory user and active organization. | | Discovery | Droid retrieves tools for apps you connected and adds them to its deferred tool catalog. | | Execution | Droid sends approved tool calls through Factory's managed connector service. No connector process runs on your local machine. | | Disconnection | Factory removes your credential for that app in the active organization. Organization availability for other members does not change. | If an MCP server and a connected app expose the same capability, Droid removes overlapping tools from the session catalog where possible. Self-hosted and cloud instances remain separate when they may point to different systems. ## Connectors and MCP Connectors and MCP both extend Droid with external tools. They solve different setup and governance needs. | | Connectors | MCP | | :-- | :--------- | :-- | | Setup | Choose an app and finish its authorization flow. | Configure or install an MCP server. | | Catalog | Factory-curated third-party apps and tool sets. | Any compatible local or remote MCP server. | | Runtime | Factory-managed remote connector service. | Local `stdio` process or remote `http` / `sse` endpoint. | | Authentication | Guided per-user, per-organization connection. | App-provided authorization, headers, environment variables, or server-defined authentication. | | Configuration | No repository file or server lifecycle to manage. | User, folder, project, or organization `mcp.json`. | | Organization controls | Owners and Managers enable or disable each catalog app for the organization. | Administrators control server access with managed MCP policy. | | Best fit | Supported cloud apps that should work with minimal setup. | Custom tools, private services, self-hosted instances, and explicit server control. | You can use both in the same session. Choose the surface that represents the system you intend to access. For example, keep a self-hosted GitLab MCP server even if you also connect a cloud GitLab account. ## Manage your connections Open **Settings → Connectors** to review connected apps. - **Connect** starts a new authorization flow for that app. - **Disconnect** removes your connection for the active organization. It does not disable the connector for teammates. - **Suggest a connector** sends Factory the app name, website, and optional use case for review when the app is not in the catalog. Never put passwords, API keys, access tokens, or customer data in a connector suggestion. Use the app's authorization flow for credentials. ## Enterprise controls Enterprise organizations get an organization-specific connector catalog. Connectors start disabled when the organization first configures this control. An **Owner** or **Manager** chooses which apps members can connect and use. In the Factory App, open **Settings → Enterprise Controls → Connectors**. The table shows every connector in Factory's curated catalog and whether it is available to this organization. Turn on the apps the organization approves. Turn off an app to remove it from member settings and the tools available to Droid. Availability does not share an account or credential. Each member connects the enabled app with their own identity in **Settings → Connectors**. The server enforces connector availability. Disabling an app: - removes it from the organization's Connectors page; - blocks new authorization links; - removes its tools from the catalog available to Droid; and - rejects direct tool calls, including calls from an existing session. Disabling an app does not delete each member's saved connection. If an Owner or Manager enables it again, a member may still be connected. Members can use **Disconnect** to remove their own credential. Organizations without Enterprise connector controls use Factory's curated catalog and do not see organization-level availability switches. Owners and Managers can also recommend enabled connectors to members by work persona, so new members are guided to the right apps during onboarding. See [Personas](/enterprise/personas). ## Security and data handling | Control | Behavior | | :------ | :------- | | Identity scope | A connection is isolated by Factory user and active organization. It is not a shared organization credential. | | App permissions | The third-party app controls the permissions shown during authorization. Review them before approving access. | | Tool approval | Autonomy and confirmation checks run before Droid calls a connector tool. Unknown or non-read-only tools default to high risk. | | Server-side enforcement | Organization availability is checked before Factory creates an authorization link or runs a tool. A disabled app cannot be called by bypassing the UI. | | Local environment | Connector calls run through the managed remote service. They do not start a local process or read local files unless a separate Droid tool supplies that data. | | Connector audit event | Factory records the connector tool name for the connector invocation event, without adding tool arguments or result content to that event. | Only connect accounts and workspaces that Droid is allowed to access. The permissions granted by the third-party app remain the final boundary on what its tools can read or change. ## Troubleshooting Confirm that you selected the intended organization. In an Enterprise organization, ask an Owner or Manager to open **Enterprise Controls → Connectors** and enable the app. If Factory does not offer it, use **Suggest a connector** or configure an [MCP server](/harness/mcp). Allow pop-ups for the Factory App, then select **Connect** again to create a fresh link. Confirm that your account can authorize third-party access in the target app. Return to the session and ask Droid to retry. If the active session still has the earlier tool catalog, start a new session. Confirm that the session and connection use the same active organization. Connector tools follow the session's Autonomy Level. Read-only tools are low risk; write, destructive, and unclassified tools are high risk. Connector approvals do not persist as MCP tool permissions. Check whether an Owner or Manager disabled the app for the organization. If it is still enabled, disconnect and reconnect the app to refresh its authorization. Configure custom, private, and self-hosted tool servers. Control which connector actions run without confirmation. Understand organization settings, precedence, and governance. Manage Factory roles, membership, SSO, and Directory Sync. # Model Context Protocol (MCP) Connect your own tools with Model Context Protocol Model Context Protocol (MCP) servers extend Droid's capabilities with extra tools and context. You can manage them two ways: an interactive manager inside the TUI for browsing and setup, or `droid mcp` CLI commands for scripting and automation. Droid supports three transports: **stdio** (local processes), **http** (Streamable HTTP, the current MCP standard, which streams responses over SSE internally), and **sse** (the legacy standalone HTTP+SSE transport, for older servers). For a supported cloud app with Factory-managed setup and guided authentication, use [Connectors](/harness/connectors). Use MCP when you need a custom tool, a self-hosted service, or direct control over the server and transport. ## Quick start: add from the registry The fastest way to start is the built-in server registry. Type `/mcp` inside Droid and select **Add from Registry**. Browse the registry and choose one, such as `linear`, `sentry`, or `playwright`. For remote servers that need OAuth, follow the browser prompt. The server is then ready to use. Representative registry servers include: | Server | Description | | :----- | :---------- | | linear | Issue tracking and project management | | sentry | Error tracking and performance monitoring | | notion | Notes, docs, and project management | | figma | Generate code with Figma context | | stripe | Payment processing APIs | | supabase | Create and manage Supabase projects | | vercel | Manage projects and deployments | | playwright | End-to-end browser testing | The registry includes many more servers than shown here. The registry is the quickest path for popular servers. For custom servers or automation, use the CLI commands below. ## Manage servers interactively (`/mcp`) Type `/mcp` inside Droid to open the interactive manager. From there you can: {/* sweep-allow: term-bullets */} - **Browse servers** and see their connection status. - **View tools** that each connected server provides. - **Enable or disable** servers without removing them. - **Authenticate** OAuth-enabled servers via the browser. - **Clear auth** to remove stored credentials for a server. - **Add from registry** for one-click setup of popular servers. - **Remove** user-configured servers. Run `/mcp off` to disable every configurable server for the current session. Organization-managed servers are unaffected. ## Add servers from the CLI For scripting and automation, use `droid mcp add`. The transport flag determines how the rest of the arguments are parsed. ```bash droid mcp add --type ``` - `name` is a unique server identifier. - `--type` defaults to `stdio` when omitted. - `--env KEY=VALUE` sets environment variables (stdio only, repeatable). - `--header "KEY: VALUE"` sets HTTP headers (http/sse only, repeatable). - `--no-oauth` disables OAuth for a header- or API-key-authenticated remote server (stores `oauth: false`). HTTP servers are remote MCP endpoints, the recommended way to connect to cloud services. ```bash droid mcp add linear https://mcp.linear.app/mcp --type http ``` Pass authentication headers with `--header`, repeating the flag as needed: ```bash droid mcp add twelvelabs https://mcp.twelvelabs.io --type http \ --header "x-api-key: YOUR_API_KEY" ``` Use `--no-oauth` when a server authenticates by header or API key and you do not want Droid to attempt an OAuth flow: ```bash droid mcp add internal https://mcp.internal.example.com/mcp --type http \ --header "Authorization: Bearer YOUR_TOKEN" --no-oauth ``` `sse` is the legacy HTTP+SSE transport. Prefer `http` (which already streams over SSE internally) and reach for `sse` only when a server offers just the older standalone SSE endpoint. Arguments mirror HTTP; only `--type` changes. ```bash droid mcp add example-sse https://mcp.example.com/sse --type sse \ --header "Authorization: Bearer YOUR_TOKEN" ``` Stdio servers run as local processes, ideal for tools that need direct system access. Quote the command if it contains spaces, and pass environment variables with `--env`. ```bash droid mcp add airtable "npx -y airtable-mcp-server" \ --env AIRTABLE_API_KEY=your_key ``` `npx` examples install the latest published version of a package. For security-sensitive setups, pin an explicit version (for example `airtable-mcp-server@1.4.0`) so updates are deliberate and auditable. Many remote servers require OAuth. After adding one, run `/mcp` to complete the browser authentication flow. ## Manage servers from the CLI List every configured server with its connection and authentication status: ```bash droid mcp list ``` Each server reports its current status: **connected**, **connecting**, **needs authentication**, or **failed**. Servers that require OAuth show **needs authentication** until you finish the sign-in flow with `/mcp`. Remove a user-configured server: ```bash droid mcp remove ``` ### Persistent tool permissions When you approve an MCP tool, Droid can remember that approval so it persists across sessions. Each approval is bound to a stable fingerprint of the server's transport configuration (its stdio command and arguments, or its http/sse URL). If a previously trusted server name is later re-pointed at a different command or URL, the stored approval no longer applies and the tool must be approved again. Manage these approvals with `droid mcp permissions`: ```bash droid mcp permissions list droid mcp permissions revoke [tool] droid mcp permissions clear --confirm ``` - `list` shows all persistent permissions. - `revoke ` removes a server's approval, including all of its per-tool approvals. Add a `tool` argument to revoke a single tool. - `clear --confirm` removes every persistent permission. ## Configuration file MCP server configurations are stored in `mcp.json` files at three levels: | Level | Location | Purpose | | :---- | :------- | :------ | | **User** | `~/.factory/mcp.json` | Your personal servers, available in every project. | | **Folder** | `.factory/mcp.json` in an ancestor directory of the project | Servers shared across nested projects under a common parent. | | **Project** | `.factory/mcp.json` in the project root | Shared team servers, committed to the repo. | Organizations can also provide servers centrally and restrict which ones are allowed through managed settings (see [Enterprise: MCP policy](#enterprise-mcp-policy)). **Behaviors to know:** - Servers you add with `droid mcp add` or the registry always go to your **user** config. - **Project servers cannot be removed** with `droid mcp remove` or the `/mcp` manager. To remove them, edit `.factory/mcp.json` directly. - When you **enable or disable** a project-defined server, Droid writes a copy to your user config with the new state and leaves the project file untouched, so your teammates are unaffected. - When the same server name is defined at more than one level, Droid loads one definition for it. Organization-managed servers and the [MCP policy](#enterprise-mcp-policy) always take precedence. The `/mcp` manager shows which file each server comes from. OAuth tokens are stored globally in your system keyring (or a fallback file), not per project, so authenticating with a server in one project authenticates it everywhere that server is configured. Use the `/mcp` manager's **Clear Auth** action to remove stored credentials. Project-level `.factory/mcp.json` is committed to the repo. Never put secrets there: header auth tokens (such as `Authorization`), `oauth.clientSecret`, or API keys. Keep them in your user-level config (`~/.factory/mcp.json`), supply them through environment variables, and rely on Droid's keyring for OAuth tokens. Droid reloads automatically when an `mcp.json` file changes, so new servers are available immediately. ```json HTTP { "mcpServers": { "linear": { "type": "http", "url": "https://mcp.linear.app/mcp", "disabled": false } } } ``` ```json SSE { "mcpServers": { "example-sse": { "type": "sse", "url": "https://mcp.example.com/sse", "headers": { "Authorization": "Bearer YOUR_TOKEN" }, "disabled": false } } } ``` ```json stdio { "mcpServers": { "playwright": { "command": "npx", "args": ["-y", "@playwright/mcp@latest"], "disabled": false } } } ``` ## Schema reference Each server entry accepts these common fields: | Field | Type | Description | | :---- | :--- | :---------- | | `type` | `"stdio" \| "http" \| "sse"` | Transport. May be omitted for stdio servers, which default to `stdio`. | | `disabled` | `boolean` | Temporarily disable the server (default: `false`). | | `disabledTools` | `string[]` | Tool names to exclude from this server. Excluded tools are never loaded into context. | | `timeout` | `number` | Tool call timeout in milliseconds. Bounds each tool invocation, not the initial connection. Falls back to the built-in default when omitted. | | `connectTimeout` | `number` | Connection timeout in milliseconds for the initial server handshake. Falls back to the transport default when omitted: 10 seconds (10000ms) for `http`/`sse`, 30 seconds (30000ms) for `stdio`. | Transport-specific fields: - **stdio** servers use `command` (the executable), `args` (an array of arguments), and `env` (an object of environment variables). - **http** and **sse** servers use `url` (the endpoint), `headers` (an object of HTTP headers), and `oauth` (OAuth overrides, or `false` to disable OAuth entirely). ### Tool filtering A server can expose many tools, and you may not want all of them in every session. Use `disabledTools` to exclude specific tools persistently in `mcp.json`. Every tool the server reports is loaded except the listed names, and excluded tools are never registered with the model, so they do not consume context tokens. ```json title="mcp.json" { "mcpServers": { "my-server": { "type": "stdio", "command": "npx", "args": ["-y", "@some/mcp-server"], "disabledTools": ["tool_i_dont_need", "another_unused_tool"] } } } ``` Run `/mcp` to see the exact tool names a server exposes, then copy them into `disabledTools`. ## Secrets and variable expansion Droid expands `${NAME}` references in `mcp.json` against your current shell environment when it connects to a server. This keeps secrets out of the file itself so you can source them from a secret manager, a `.env` loader, or your shell profile. Only the `${NAME}` form is supported; there is no default-value syntax. Expansion applies to credential-bearing fields only: - `env` values for **stdio** servers. - `headers` values for **http** and **sse** servers. - `oauth.clientId` and `oauth.clientSecret` for **http** and **sse** servers. It does **not** apply to `command`, `args`, or `url`. ```json title="mcp.json" { "mcpServers": { "context7": { "type": "http", "url": "https://mcp.context7.com/mcp", "headers": { "CONTEXT7_API_KEY": "${CONTEXT7_API_KEY}" }, "disabled": false }, "airtable": { "type": "stdio", "command": "npx", "args": ["-y", "airtable-mcp-server"], "env": { "AIRTABLE_API_KEY": "${AIRTABLE_API_KEY}" } } } } ``` If a referenced variable is unset, the connection to that server fails with an error naming the missing variable. The raw `mcp.json` file is never rewritten with expanded values; expansion happens in memory at connection time, so secrets stay out of disk and version control. ## OAuth overrides For most remote servers, OAuth works with zero configuration: Droid discovers the authorization server, registers a client automatically via Dynamic Client Registration (DCR), and uses Factory's published client metadata when the server supports Client ID Metadata Documents (CIMD). **Prefer these defaults.** Only set `oauth` overrides when a provider requires a custom trust or compatibility policy. Set `oauth: false` to disable OAuth for a server entirely. The `oauth` object on an **http** or **sse** server supports: | Field | Type | Description | | :---- | :--- | :---------- | | `scopes` | `string[]` | OAuth scopes to request instead of the discovered defaults. | | `resource` | `string \| false` | OAuth resource indicator to send instead of the normalized MCP server URL; set to `false` to omit it. | | `authorizationServerIssuer` | `string` | Authorization server issuer URL. Required when `clientId` / `clientSecret` are set. | | `clientId` | `string` | Pre-registered OAuth client ID, skipping dynamic registration. | | `clientSecret` | `string` | Client secret for the pre-registered client. | | `clientMetadataUrl` | `string` | HTTPS URL of a custom Client ID Metadata Document (CIMD) to use as the public client identity. | | `tokenEndpointAuthMethod` | `"none" \| "client_secret_basic" \| "client_secret_post"` | Force a token endpoint authentication method instead of the discovered one. | | `callbackPort` | `number` | Fixed localhost port for the OAuth callback (1-65535). | **Constraints:** - `clientMetadataUrl` must be an HTTPS URL with a non-root path and no credentials, query string, fragment, or dot segments. - `clientMetadataUrl` is mutually exclusive with `clientId` / `clientSecret`: a metadata document **is** the client identity, so pre-registered credentials cannot be combined with it. - `clientMetadataUrl` describes a public client, so `tokenEndpointAuthMethod` must be `"none"` (or omitted) when it is set. - `clientId` / `clientSecret` require `authorizationServerIssuer` to be set. ```json title="mcp.json" { "mcpServers": { "internal-tools": { "type": "http", "url": "https://mcp.internal.example.com/mcp", "oauth": { "clientMetadataUrl": "https://auth.example.com/oauth/client-metadata.json" } } } } ``` ```json title="mcp.json" { "mcpServers": { "partner-api": { "type": "http", "url": "https://mcp.partner.example.com/mcp", "oauth": { "tokenEndpointAuthMethod": "none" } } } } ``` ## MCP timeouts Two independent per-server settings bound how long Droid waits on an MCP server, both in milliseconds: - `timeout` bounds each tool invocation. Long-running tools (large data exports, browser automations, model-backed servers) can exceed the built-in default and fail with a timeout error. - `connectTimeout` bounds the initial connection and initialization handshake. Slow-starting servers (heavy `npx` installs, servers that compile on first launch) can exceed the transport default: 10 seconds for `http`/`sse` servers, 30 seconds for `stdio` servers. ```json title="mcp.json" { "mcpServers": { "slow-server": { "type": "stdio", "command": "my-long-running-server", "connectTimeout": 60000, "timeout": 120000 } } } ``` A server's values override the built-in defaults; there is no global timeout setting. Increasing these timeouts only changes how long Droid waits. They do not extend any timeouts enforced by the MCP server itself or its upstream APIs. ## Per-droid server selection [Custom droids](/harness/subagents) can choose which configured MCP servers they may use through the `mcpServers` field in their frontmatter. This scopes a subagent to specific servers (for example `mcpServers: ["linear", "github"]`) instead of inheriting every server in the session. For finer-grained control, a droid's `tools` list can name exact registered MCP tool IDs. See [Selecting MCP servers](/harness/subagents#selecting-mcp-servers) for details. ## Enterprise: MCP policy Organizations can centrally control which MCP servers are allowed through the `mcpPolicy` setting in [org-managed settings](/enterprise/hierarchical-settings-and-org-control#mcp), so users can only connect to vetted servers. A server is allowed when any allowlist entry matches its hostname or stdio command or arguments, not its configured name. | Field | Type | Description | | :---- | :--- | :---------- | | `enabled` | `boolean` | Whether policy enforcement is active (default: `false`). When `false` or absent, no policy is enforced and configured servers are allowed. | | `allowlist` | `string[]` | Allowed hostname or stdio command matchers, applied only when the policy is enabled. HTTP/HTTPS URL entries are also accepted, but only their hostname is matched. Entries do not match the configured server name. When the policy is enabled with an empty or absent allowlist, all servers are blocked. | ```json title="settings.json" { "mcpPolicy": { "enabled": true, "allowlist": [ "https://mcp.linear.app", "https://mcp.sentry.dev", "*.mcp-gateway.example.com", "npx" ] } } ``` `mcpPolicy` is enforced through managed settings, and individual users cannot override it. Servers disallowed by policy remain in `mcp.json` and are still loaded into your configuration, but they are filtered out from running or connecting and do not appear as available in the `/mcp` manager. ### Remote hostname matching Remote `http` and `sse` servers are matched by hostname only. You can use a bare hostname or an HTTP/HTTPS URL; a URL is parsed and reduced to its hostname. The scheme, port, credentials, path, query string, and fragment do not restrict access. In a hostname, `*` matches zero or more characters, including dots, so it can span multiple subdomain levels. Other characters are literal, not regular expressions or extended glob syntax. Entries without wildcards continue to match both the hostname itself and its subdomains. Each row below is evaluated on its own, with no other allowlist entries: | Allowlist entry | Example allowed URLs | Example blocked URL | | :-------------- | :------------------- | :------------------ | | `example.com` | `https://example.com/mcp`, `https://nested.tools.example.com/mcp` | `https://notexample.com/mcp` | | `https://example.com/mcp` | `http://example.com:8080/other`, `https://api.example.com/other` | `https://example.com.attacker.example/mcp` | | `*.example.com` | `https://tools.example.com/mcp`, `https://nested.tools.example.com/mcp` | `https://example.com/mcp` | | `https://*.example.com/mcp/*` | `http://tools.example.com:8080/other` | `https://example.com/mcp` | | `tools*.example.com` | `https://tools.example.com/mcp`, `https://tools1.example.com/mcp` | `https://other-tools.example.com/mcp` | | `*.bücher.example` | `https://tools.bücher.example/mcp`, `https://tools.xn--bcher-kva.example/mcp` | `https://bücher.example/mcp` | | `bü*.example` | `https://bücher.example/mcp`, `https://xn--bcher-kva.example/mcp` | `https://bĺ.example/mcp` | A full URL is not an endpoint restriction. `https://example.com/mcp`, `https://example.com/`, and `example.com` allow the same hostnames, including their subdomains. They also allow HTTP, different ports, and paths other than `/mcp`. Enforce scheme-, port-, path-, or tenant-specific restrictions at your gateway or another network control. Use a valid HTTP/HTTPS URL when including a scheme, port, or path. For example, `http://localhost:3000/mcp` allows other ports and paths on that hostname. Scheme and port wildcards such as `*://example.com` and `http://localhost:*` are not supported. Bare entries with paths or ports, such as `example.com/` and `example.com:443`, are not hostname matches. Hostname matching is case-insensitive, trims surrounding whitespace in entries, and ignores a trailing DNS dot. Wildcard matching normalizes internationalized hostnames and compares their decoded form so an embedded wildcard keeps its position. For an internationalized hostname without a wildcard, use a full URL such as `https://bücher.example/mcp` or its ASCII hostname `xn--bcher-kva.example`, not the bare Unicode hostname `bücher.example`. Only a literal ASCII `*` enables wildcard matching; encoded stars such as `%2A` and full-width stars such as `*` are not wildcard syntax. A wildcard must match the entire hostname. For example, `*.example.com` does not allow `notexample.com` or `tools.example.com.attacker.example`, and cannot match text in a URL's credentials, path, or fragment. A bare `*` allows every remote hostname, but does not allow malformed URLs, URLs without a hostname, or stdio commands. These allowlist rules are separate from the picomatch glob patterns in [`mcpAutonomyUrlOverrides`](#enterprise-mcp-autonomy-url-overrides), which control tool risk levels rather than server access. ### Example: allow production and staging gateways Suppose each approved MCP service has its own subdomain under your production or staging gateway. Add both gateway patterns to your org-managed settings, replacing the example domains with domains your organization controls: ```json title="settings.json" { "mcpPolicy": { "enabled": true, "allowlist": [ "*.mcp-gateway.prod.example.com", "*.mcp-gateway.staging.example.com" ] } } ``` Users can then configure servers such as these in `~/.factory/mcp.json`, or a team can share them in the project's `.factory/mcp.json`: ```json title="mcp.json" { "mcpServers": { "ticketing": { "type": "http", "url": "https://atlassian.mcp-gateway.prod.example.com/mcp" }, "code-search": { "type": "http", "url": "https://sourcegraph.mcp-gateway.prod.example.com/mcp" }, "staging-tools": { "type": "sse", "url": "https://tools.mcp-gateway.staging.example.com/sse" } } } ``` All three servers pass the policy because of their URL hostnames, regardless of the names `ticketing`, `code-search`, and `staging-tools`. The allowlist does not add servers or authenticate them; users still need to configure them and complete any required authentication. With only the two gateway entries above: | Server URL | Policy result | Why | | :--------- | :------------ | :-- | | `https://docs.team.mcp-gateway.prod.example.com/mcp` | Allowed | `*` spans multiple subdomain levels. | | `http://atlassian.mcp-gateway.prod.example.com:8080/other` | Allowed | Scheme, port, and path do not restrict the hostname match. | | `https://mcp-gateway.prod.example.com/mcp` | Blocked | `*.` does not include the parent hostname. | | `https://tools.mcp-gateway.dev.example.com/mcp` | Blocked | Neither entry allows the development gateway. | | `https://tools.mcp-gateway.prod.example.com.attacker.example/mcp` | Blocked | The hostname does not end at the approved domain. | | `https://attacker.example/atlassian.mcp-gateway.prod.example.com/mcp` | Blocked | An approved hostname in the path does not count. | To also allow the parent gateway, replace `*.mcp-gateway.prod.example.com` with `mcp-gateway.prod.example.com`; a non-wildcard hostname allows both the parent and its subdomains. This policy does not grant blanket access to stdio servers. ### Local command matching For `stdio` servers, entries match substrings of the command or individual arguments after lowercasing and removing everything except ASCII letters and digits. Hostname wildcard rules do not apply to `stdio` servers. For example, with `"command": "npx"` and `"args": ["-y", "figma-mcp"]`: | Allowlist entry | Policy result | Why | | :-------------- | :------------ | :-- | | `npx` | Allowed | Matches the command, so it also allows other servers launched with `npx`. | | `figma-mcp` | Allowed | Matches an argument after normalization. | | `figma-*` | Allowed | Becomes `figma`, a substring of the normalized argument; `*` is removed, not expanded. | | `fig*mcp` | Blocked | Becomes `figmcp`, which is not a substring of the command or any argument. | | `*` | Blocked | Normalization leaves an empty matcher. | Choose matchers deliberately: allowing a launcher such as `npx` is broader than matching a package argument. Stdio matching is a substring check, not package identity or integrity verification. ## Enterprise: MCP autonomy URL overrides Administrators can assign a default [autonomy risk level](/autonomy-and-safety/auto-run) to remote MCP servers by URL with the `mcpAutonomyUrlOverrides` org-managed setting, controlling how much confirmation a matching server's tools require before Droid runs them. Each rule maps a URL pattern to a risk level: | Field | Type | Description | | :---- | :--- | :---------- | | `urlPattern` | `string` | Glob pattern ([picomatch](https://github.com/micromatch/picomatch) syntax) matched against the server's URL. | | `defaultLevel` | `"low" \| "medium" \| "high"` | Risk level for the matching server's tools, compared against the user's Autonomy Level to decide auto-run vs. confirm. | ```json { "mcpAutonomyUrlOverrides": [ { "urlPattern": "https://mcp.internal.example.com/**", "defaultLevel": "low" }, { "urlPattern": "https://*.partner.example.com/**", "defaultLevel": "medium" }, { "urlPattern": "https://**", "defaultLevel": "high" } ] } ``` **Matching and precedence:** {/* sweep-allow: term-bullets */} - **Remote servers only.** Rules match remote servers that have a URL (`http` and `sse` transports); local `stdio` servers are unaffected. - **First match wins.** Rules are checked in order, so list the most specific patterns first. - **Safety floor.** A rule sets a tool's risk *classification*, not an absolute prompt: `high`-classified tools still follow the normal [Autonomy Level](/autonomy-and-safety/auto-run) comparison. The floor is one-directional: `low` or `medium` cannot pull a tool that is not read-only (destructive, or missing safety metadata) below `high`, so you can relax confirmation for read-only tools but never auto-approve destructive ones. - **Fallback.** Servers with no matching rule use Droid's built-in tool risk classification (read-only hints and curated defaults). `mcpAutonomyUrlOverrides` is admin-managed (MDM): it is delivered through [org-managed settings](/enterprise/hierarchical-settings-and-org-control) and users cannot override or weaken it. Connect supported third-party apps with managed authentication. Scope which MCP servers a subagent may use through its frontmatter. Configure MCP servers, policy, and autonomy URL overrides at every scope. Control when MCP and connector tools require confirmation. # Custom droids (subagents) Define specialized subagents with their own system prompt, model, and tool policy that Droid delegates focused tasks to in a fresh context window. A custom droid is a reusable subagent defined in Markdown. Each droid carries its own system prompt, model preference, and tool policy, so Droid can hand off a focused task, such as code review, a security sweep, or research, without you re-typing instructions. Every invocation runs in a fresh context window through the **Task** tool. Droid also ships two built-in droids, `worker` and `explorer`, that you can use without defining anything. Custom droids extend that set with your own prompts, models, and tool restrictions. A subagent gives you: - **Context isolation**: it runs in a fresh context window, so the parent session stays focused and lean. - **Its own tooling and autonomy**: you can restrict it to read-only, edit-only, or a curated tool set, and it runs at its own autonomy level. - **Its own model**: it can inherit the parent's model or use a different one tuned for the task. - **A single return value**: it hands back one final message, which is not shown to you unless the parent summarizes it. Custom droids are enabled by default; you can toggle them off in Settings (`/settings`) under the Experimental section. Subagents run non-interactively. The `AskUser` tool is disabled for a subagent, and a subagent cannot spawn its own subagents (the `Task` tool is not available to it). If something is unclear or blocked, the subagent reports back instead of prompting. ## Custom droids vs skills Custom droids and [skills](/harness/skills) both package reusable work, but they solve different problems. | Reach for a custom droid when | Reach for a skill when | | :---------------------------- | :--------------------- | | You want work to run in a fresh context window. | You want a workflow to run inline in the current session. | | You need a different model than the parent session. | The current model is fine. | | You need a stricter, enforced tool policy. | You only need to document intended tools. | | The task is a complex checklist best encoded in a system prompt. | The task is a lightweight, reusable procedure. | A custom droid is a runtime tool boundary and a separate agent. A skill is a discoverable instruction set. See [Skills](/harness/skills) for skill mechanics, frontmatter, and slash invocation. ## Where they live Custom droids are `.md` files in either of two locations. The CLI scans the top level of each `droids/` folder, validates every definition, and exposes valid droids as `subagent_type` targets for the Task tool. - **Project droids** sit in `/.factory/droids/` and are shared with teammates through the repository. - **Personal droids** live in `~/.factory/droids/` and follow you across workspaces. When a project droid and a personal droid share the same name, the project definition wins. ## Create a droid Run `/droids`, then choose **Create a new Droid**. Choose project or personal storage. The CLI writes `.md` into the matching `droids/` directory and normalizes the filename to lowercase and hyphenated. Set a description, a system prompt (auto-generated or hand-written), an identifier, a model (or `inherit`), and a tool selection. Ask Droid to delegate to it, for example "Use the subagent `code-reviewer` on the staged diff." New or edited droid files are picked up on the next menu open or Task tool invocation. To scaffold a droid without writing the prompt by hand, use the built-in `GenerateDroid` tool. Describe what the droid should do (for example, "review pull requests for security issues in Node.js services") and Droid drafts a normalized `name`, a focused system prompt, and a sensible expanded description, then saves the `.md` file. `GenerateDroid` defaults to the `project` location; pass `location: personal` to save it to `~/.factory/droids/` instead. You can open and edit the generated file afterward. ## Configuration Each droid file is Markdown with YAML frontmatter followed by the system prompt body. ```markdown title=".factory/droids/code-reviewer.md" --- name: code-reviewer description: Focused reviewer that checks diffs for correctness risks model: inherit tools: read-only --- You are the team's senior reviewer. Examine the diff the parent agent shares and: - flag correctness, security, and migration risks - list targeted follow-up tasks if changes are required - confirm tests or manual validation needed before merge ``` You can also pass `tools` as an array and pin a specific model: ```markdown title=".factory/droids/deep-analyzer.md" --- name: deep-analyzer description: Thorough analysis with extended thinking model: claude-sonnet-4-5-20250929 reasoningEffort: high tools: ["Read", "Grep", "Glob", "WebSearch"] --- Perform deep analysis of the code or problem presented. ``` | Field | Notes | | :---- | :---- | | `name` | Required. Lowercase letters, digits, `-`, `_` (`^[a-z0-9-_]+$`). Drives the `subagent_type` value and filename. | | `description` | Optional but recommended. Shown in the `/droids` list. A description over 500 characters raises a validation warning. | | `model` | `inherit` (default) uses the parent session's model. Otherwise use a public model ID from [Models](/models), for example `claude-sonnet-4-5-20250929`. For BYOK custom models, use `custom:` plus the `model` field from your config (for example `custom:gpt-4o-mini`), not the display name. | | `reasoningEffort` | Optional. `low`, `medium`, or `high` for models that support it. Ignored when `model` is `inherit`, and must be compatible with the selected model. | | `tools` | Omit to allow all tools, use a category string (for example `read-only`), or pass an array of tool IDs. Tool IDs are case-sensitive. | | `mcpServers` | Optional. Array of [MCP server](/harness/mcp) names whose tools are exposed to the droid. See [Selecting MCP servers](#selecting-mcp-servers). | The body after the frontmatter is the system prompt and cannot be empty. `DroidValidator` reports errors (invalid names, unknown models, unknown or forbidden tools) and warnings (missing description, duplicate tools) when a file loads. Three tool-policy rules are enforced at load time: - `TodoWrite` and `Skill` are always included for every droid so it can track tasks and load skills. You do not list them, and they do not appear in the tool count. - `ExitSpecMode` and `GenerateDroid` cannot be enabled by a custom droid; listing either one is a validation error. - The literal value `tools: all` is rejected. Omit the `tools` field entirely to allow every tool. ### Tool categories Use a category name as the `tools` value (for example `tools: read-only`) or list individual tool IDs in an array. | Category | Tool IDs | Purpose | | :------- | :------- | :------ | | `read-only` | `Read`, `LS`, `Grep`, `Glob` | Analysis and file exploration | | `edit` | `Create`, `Edit`, `ApplyPatch` | Code generation and modification | | `execute` | `Execute` | Shell command execution | | `web` | `WebSearch`, `FetchUrl` | Internet research and content | | `mcp` | Dynamically populated | Model Context Protocol tools | Arrays must use valid IDs from this table or exact registered MCP tool IDs. Unknown IDs cause a validation error. When `Edit` is enabled with an OpenAI model, `ApplyPatch` is added automatically for compatibility. When `model` is `inherit`, both are enabled to cover providers that differ at runtime. ### Selecting MCP servers Use `mcpServers` to limit which [MCP servers](/harness/mcp) a droid can reach. The droid receives the tools from each listed server in addition to anything in `tools`; configured servers that are not listed are excluded. ```markdown title=".factory/droids/issue-researcher.md" --- name: issue-researcher description: Researches issues using repository context and tracker data model: inherit tools: ["Read", "Grep"] mcpServers: ["linear", "github"] --- Investigate the issue referenced in the prompt using the codebase and the selected MCP servers, then summarize findings and propose next steps. ``` - Server names must match entries in `~/.factory/mcp.json` or `.factory/mcp.json`. - Omitting `mcpServers` keeps the parent session's MCP tool availability. - Setting `mcpServers: []` excludes every MCP server, even globally configured ones. - For finer control, list exact registered MCP tool IDs in `tools` to allow specific tools rather than a whole server. - Servers blocked by an [enterprise MCP policy](/harness/mcp#enterprise-mcp-policy) stay unavailable even if listed here. ## Built-in droids Two droids ship with Droid and need no definition. Use them directly as `subagent_type` values. | Droid | Tools | Default complexity | Use for | | :---- | :---- | :----------------- | :------ | | `worker` | All tools | `medium` | General-purpose, multi-step tasks including edits and commands. | | `explorer` | Read-only | `light` | Fast codebase exploration, search, and structure questions. | When you invoke a built-in droid without setting `complexity`, it uses its default tier above. A few additional built-in droids (for example `scrutiny-feature-reviewer` and `user-testing-flow-validator`) are written to `~/.factory/droids/` for use within [Missions](/missions/overview) validation and are not meant for general delegation. ## Invoking custom droids Droid calls droids through the **Task** tool. It may delegate on its own, or you can ask directly: "Use the subagent `security-sweeper` on the files I changed." The Task tool accepts: - `subagent_type` (required): the droid name, for example `code-reviewer`, `worker`, or `explorer`. - `description` (required): a short label for the UI. - `prompt` (required): the full task. - `image_paths` (optional): local image file paths to attach to the subagent (paths only, never base64). - `complexity` (optional): `light`, `medium`, or `heavy`. When set, model selection follows your configured complexity-to-model routing in settings. - `run_in_background` (optional): when `true`, the task returns a `task_id` immediately and runs asynchronously. Retrieve the result with the `TaskOutput` tool. - `resume` (optional): a `task_id` from a previous invocation, to continue that task with its full context preserved. Run `/droids` to open the manager and confirm a droid's name before delegating to it. ### Foreground and background - **Foreground (default)**: the parent waits for the subagent to finish. The Task tool streams live progress (tool calls, results, and `TodoWrite` updates) as the subagent runs, then returns its final message. - **Background (`run_in_background: true`)**: the Task tool returns a `task_id` immediately and the subagent keeps running independently. Use it for genuinely independent work that can run in parallel. The parent is notified when it completes. To retrieve or manage a background subagent, the parent uses two companion tools: - **`TaskOutput`**: fetch a background task's output by `task_id`. Use `block=true` to wait for completion, or `block=false` to poll status without waiting. - **`TaskStop`**: stop a running background task by `task_id` (sends `SIGTERM`, then `SIGKILL` if needed). The parent fetches the result itself with `TaskOutput` rather than assuming you will be notified later. ### Resuming and parallel runs Pass `resume` with a prior `task_id` to send a follow-up turn to an existing subagent session. The subagent keeps its full prior context, and its autonomy level is re-aligned to the parent's current level for the new turn. This works for both foreground and background subagents. To run subagents concurrently, the parent issues multiple Task tool calls in the same turn, or launches several with `run_in_background: true` and collects the results with `TaskOutput`. ## Autonomy level Subagents run at an autonomy level, just like the main session (`Off`, `Low`, `Medium`, `High`; see [Autonomy Levels](/autonomy-and-safety/auto-run)). Control it with the **Subagent autonomy level** setting in `/settings` under **Subagents**. | Value | Behavior | | :---- | :------- | | `inherit` (default) | The subagent runs at the parent session's current autonomy level. | | `off` | Read tools and allowlisted commands only. | | `low` | File edits plus low-risk commands and MCP tools. | | `medium` | Adds reversible workspace changes (installs, local commits, builds). | | `high` | Adds high-risk actions unless safety checks require approval. | - The resolved level is always clamped to the organization's **Maximum Autonomy Level**, so an explicit setting can never exceed the enterprise cap. - When the parent session is in Spec Mode, subagents are restricted to read-only operations and low-risk shell commands; file edits and file creation are disabled. - On `resume`, a subagent's autonomy level is re-aligned to the parent's current level for the follow-up turn. ## Model selection Each subagent's model is resolved from its droid config plus the parent's complexity routing, in this order: 1. **Droid `model`**: a `model` pinned in the droid's frontmatter wins. Use a public model ID from [Models](/models), or `custom:` plus your BYOK `model` field. Set `model: inherit` (the default) to defer to the parent. 2. **Complexity-to-model routing**: when the droid's model is `inherit` and the parent passes a `complexity` tier, Droid maps that tier to a model using the routing you configure in `/settings` under **Subagents**. The **Light**, **Medium**, and **Heavy** task-model settings each map a tier to a specific model (with an optional reasoning effort) or to the Auto model router, or leave it on **Inherit** to use the spawning session's model. 3. **Parent fallback**: if no explicit routing applies, the subagent uses the parent session's active model and reasoning effort. 4. **Validation fallback**: if a droid pins a model that is not allowed (blocked by org policy, or a BYOK model that is not configured), Droid falls back to the parent's model rather than failing. The built-in `worker` and `explorer` droids use `model: inherit`, so they follow complexity-to-model routing based on their default tier (`medium` and `light`) unless the parent overrides `complexity`. ## Enterprise controls Administrators can govern subagent autonomy and models centrally through [organization-managed settings](/enterprise/hierarchical-settings-and-org-control). Org-level values win over user, project, and folder settings and cannot be weakened downstream. - **`subagentAutonomyLevel`** pins the autonomy level for all Task-launched subagents (`inherit`, `off`, `low`, `medium`, or `high`). - **`maxAutonomyLevel`** caps autonomy for every session and subagent, so a user or project setting can never exceed the org cap. - **`subagentModelSettings`** pins the complexity-to-model routing per tier (`lightModel`, `mediumModel`, `heavyModel`, each with an optional reasoning effort) so subagents run on approved models. - A droid that pins a model blocked by org policy falls back to the parent's model rather than running the disallowed model. - Droid tool policy and the [enterprise MCP policy](/harness/mcp#enterprise-mcp-policy) still apply to subagents; servers or tools blocked at the org level stay unavailable even if a droid lists them. See [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control) for the full managed-settings schema and precedence rules. ## Manage droids in /droids `/droids` opens a modal that lists each droid with its name, model, description preview, location badge (Project or Personal), and tools summary. From the menu you can: {/* sweep-allow: term-bullets */} - **Create a new Droid** through the guided wizard. - **View, Edit, or Delete** an existing droid. - **Import from Claude Code** to convert existing agents. - **Reload** to refresh the list after editing files on disk. The detail view shows the resolved tools and any selected MCP servers so you can confirm the configuration persisted. ## Import from Claude Code The droids menu can import agents created in Claude Code. Open `/droids`, start the import flow, and the CLI scans both `/.claude/agents/` (project scope) and `~/.claude/agents/` (personal scope). For each selected agent it: - maps the agent name, description, and instructions to the droid `name`, `description`, and system prompt body - maps the model family to a Factory model: `inherit` stays `inherit`, and `sonnet`, `haiku`, or `opus` map to the first available model in that family (unmatched names fall back to `inherit`) - maps tool names to Factory tools, warning on any that have no equivalent - saves each agent to the matching Factory location, so project agents become project droids and personal agents become personal droids Agents that already exist are pre-deselected. If an import reports invalid tools, edit the droid to remove unmapped tools or omit the `tools` field to allow all tools. ## Examples ### Code reviewer (project scope) ```markdown title=".factory/droids/code-reviewer.md" --- name: code-reviewer description: Reviews diffs for correctness, tests, and migration fallout model: inherit tools: ["Read", "LS", "Grep", "Glob"] --- You are the team's principal reviewer. Given the diff and context: - summarize the intent of the change - flag correctness risks, missing tests, or rollback hazards - call out migrations or data changes that need coordination ``` ### Security sweeper (personal scope) ```markdown title="~/.factory/droids/security-sweeper.md" --- name: security-sweeper description: Looks for insecure patterns in recently edited files model: inherit tools: ["Read", "Grep", "WebSearch"] --- Investigate the files referenced in the prompt for security issues: - identify injection, insecure transport, privilege escalation, or secrets exposure - suggest concrete mitigations - link to relevant CWE or internal standards when helpful ``` ### Fast explorer with a smaller model ```markdown title=".factory/droids/repo-explorer.md" --- name: repo-explorer description: Quickly maps unfamiliar code without making changes model: claude-haiku-4-5-20251001 tools: read-only --- Trace how the feature named in the prompt is implemented. Report the entry points, key modules, and data flow, then list the files worth reading next. ``` For model selection guidance across droids, see [Models](/models). Scope which MCP servers a subagent can reach at runtime. Control what a subagent can do without repeated approvals. # Skills Create reusable SKILL.md workflows that Droid can discover, invoke, and share across projects. Skills package reusable workflows for Droid. A skill is a directory that contains a `SKILL.md` entry point, optional supporting files, and frontmatter that tells Droid when the workflow applies. Use skills when instructions are too specific for always-on `AGENTS.md`, too structured for a one-off prompt, and useful enough to reuse. ## Create your first skill Put team-shared skills under `.factory/skills/` in your repository: ```bash mkdir -p .factory/skills/summarize-diff ``` Add required `name` and `description` frontmatter, then write the workflow Droid should follow. Start a new Droid session if the skill is not already visible. Ask naturally for a matching task, or run `/summarize-diff` when the skill is user-invocable. ```markdown title=".factory/skills/summarize-diff/SKILL.md" --- name: summarize-diff description: Summarize the staged git diff in 3-5 bullets. Use when the user asks for a summary of pending changes. --- # Summarize Diff ## Instructions 1. Run `git diff --staged`. 2. Summarize the changes in 3-5 bullets. 3. Call out migrations, risky areas, and suggested validation commands. ``` The skill entry point must be named `SKILL.md`. Do not use `skill.mdx` for Droid skills. ## When to use a skill Skills work best for repeatable procedures with a clear trigger and completion criteria: - Reviewing an API change against a team checklist. - Summarizing a staged diff in a consistent format. - Applying project-specific frontend conventions. - Running a validation workflow with known commands. - Producing a standard report, ticket, or handoff format. Use a different surface when the guidance has another shape: | Surface | Best for | | :------ | :------- | | `AGENTS.md` | Always-on repository instructions, commands, conventions, and safety rules. | | Skill | Reusable workflow Droid can choose or the user can invoke. | | Custom slash command | Simple user-invoked prompt or executable shortcut. | | Custom droid | Specialized subagent with its own prompt, model, and tool policy. | | MCP server | External tools and services Droid can call. | ## Skill anatomy Only `SKILL.md` is required. Keep supporting files beside it when they make the workflow clearer, safer, or easier to validate. ```text .factory/skills/review-api-change/ SKILL.md checklists.md schemas/ ship-check.schema.json scripts/ verify-api-change.sh ``` Good supporting files include: - Checklists for review, shipping, or validation. - Schemas that describe expected inputs or outputs. - Small support scripts the skill can call. - Reference files that point to existing modules, API surfaces, or run books. Supporting files are not loaded automatically. Mention them from `SKILL.md` so Droid knows when to read or run them. Do not copy production code, secrets, customer data, or private host names into a skill folder. Link to the source of truth instead. ## Discovery and loading Droid uses progressive disclosure so skill bodies stay out of context until needed: 1. **Discovery:** Droid finds `SKILL.md` files and reads each skill's `name` and `description`. 2. **Selection:** Droid compares the user request to the available skill descriptions. 3. **Invocation:** When a skill applies, Droid loads the full `SKILL.md` body and follows the workflow. The `description` is the routing surface. Write it for model matching: name the action, trigger, and boundary in one or two sentences. A good description says what the skill does and when to use it. Example: `Review API changes for backward compatibility. Use when a user edits public routes, schemas, or SDK-facing types.` ## Manage skills Run `/skills` to inspect the skills Droid discovered and the version that is active. The manager separates skills into **Effective**, **User**, **Project**, **Plugins**, **Built-in**, and **Import** tabs. The **Effective** tab groups every discovered skill by status: | Status | Meaning | | :----- | :------ | | Enabled | This is the active version Droid can invoke. | | Disabled | The skill is turned off by `SKILL.md` frontmatter or a settings ledger. | | Overridden | A higher-priority skill with the same name is active instead. | | Invalid | Droid found the skill, but its name, frontmatter, or prompt is invalid. | Select a skill and press to toggle it. The **User** tab writes the choice to `disabledSkills` in `~/.factory/settings.json`; the **Project** tab writes to `/.factory/settings.json`. If another scope or `enabled: false` frontmatter disabled the skill, the manager identifies that source instead of changing the wrong configuration. ```json title=".factory/settings.json" { "disabledSkills": ["deploy-production", "summarize-diff"] } ``` `disabledSkills` stores sanitized skill names. User and project arrays are combined, so a disable at either level wins. Removing a name re-enables the skill unless another settings level or its frontmatter still disables it. Names for skills that are not currently installed have no effect. ## Where skills live Droid can load skills from several scopes. A skill is any directory under a `skills/` folder that contains `SKILL.md`. | Scope | Location | Purpose | | :---- | :------- | :------ | | Project | `/.factory/skills//SKILL.md` | Team-shared workflows checked into a repository. | | Folder-specific | `//.factory/skills//SKILL.md` | Workflows that apply only after Droid inspects a deeper project area. | | Personal | `~/.factory/skills//SKILL.md` | Private workflows available across projects on your machine. | | Compatibility | `/.agents/skills/**/SKILL.md`, `/.agent/skills/**/SKILL.md` | Repository skills from compatible folder conventions. | | Personal compatibility | `~/.agents/skills/**/SKILL.md`, `~/.agent/skills/**/SKILL.md` | Personal skills from compatible folder conventions. | | Mission | `{missionDir}/skills/**/SKILL.md` | Skills scoped to a mission session. | | Plugin | `skills//SKILL.md` inside an installed plugin | Shared skills distributed with plugin packages. | | Built-in | Shipped with Droid | Product-provided skills available in supported sessions. | Droid searches skill folders recursively. The first directory that contains `SKILL.md` is treated as a skill directory. ### Resolution precedence When multiple sources provide the same sanitized skill name, Droid keeps one effective version and shows the others as overridden. The main user-facing precedence is: 1. Folder-specific and project skills. 2. Project plugin skills. 3. Personal skills. 4. User plugin skills. 5. Built-in skills. Mission skills and organization-managed or command-line settings can take precedence when those scopes are active. Open the **Effective** tab to see the winning source and the reason each other copy was overridden. Within one source bucket, duplicate names are invalid configuration. Rename or remove the duplicate, then run `/diagnostics` to confirm that the collision is resolved. ## Frontmatter reference ```yaml --- name: my-skill description: What this skill does and when to use it. Use when the user asks to ... allowed-tools: - Read - Grep - Glob enabled: true user-invocable: true disable-model-invocation: false license: MIT compatibility: droid version: 1.0.0 metadata: owner: platform-team --- ``` | Field | Required | Default | Description | | :---- | :------- | :------ | :---------- | | `name` | Yes | None | Skill identifier. Use lowercase letters, numbers, and hyphens. | | `description` | Yes | None | Short routing description. Include what the skill does and when to use it. | | `allowed-tools` | No | None | Declares the tools the skill is designed to use. It is metadata for review, packaging, and UI display; it does not add tools or sandbox standard skill invocation. | | `enabled` | No | `true` | Set to `false` to keep the skill on disk but disable it. | | `user-invocable` | No | `true` | Set to `false` to hide the skill from `/skill-name` slash invocation. | | `disable-model-invocation` | No | `false` | Set to `true` to prevent Droid from invoking the skill automatically. Users can still invoke it directly. | | `license` | No | None | Optional license metadata for shared skills. | | `compatibility` | No | None | Optional compatibility metadata for catalogs, plugins, or team tooling. | | `metadata` | No | None | Optional structured metadata for your own tooling. Do not put secrets here. | | `version` | No | None | Optional version string for shared or packaged skills. | Older skill files may contain a `tools` field. Treat `tools` as deprecated legacy metadata and use `allowed-tools` for tool restrictions. ## Control who invokes a skill By default, both you and Droid can invoke a valid enabled skill. | Configuration | User slash invocation | Droid invocation | Use when | | :------------ | :-------------------- | :--------------- | :------- | | Default | Yes | Yes | The skill is safe for direct and automatic use. | | `disable-model-invocation: true` | Yes | No | The skill should run only when the user explicitly asks for it. | | `user-invocable: false` | No | Yes | The skill is background guidance or support workflow users should not call directly. | | `enabled: false` | No | No | Keep the skill on disk but turn it off. | Use `disable-model-invocation: true` for manual workflows with side effects, such as deployment. Use `user-invocable: false` for background guidance where a slash command would not be meaningful. ## Declare intended tools `allowed-tools` records the tools the skill is designed to use. Droid stores it as skill metadata and shows it in skill surfaces, but standard skill invocation is not a runtime sandbox. A skill cannot add tools that are unavailable in the current session. ```yaml allowed-tools: - Read - Grep - Glob ``` Use `allowed-tools` when it makes the skill easier to review: - Read-only audit skills. - Workflows that should not edit files. - Workflows that should not call external systems. - Team-shared skills that need a narrow, reviewable capability set. Do not rely on `allowed-tools` as a security boundary for skills. To enforce tool access, use a [custom droid](/harness/subagents) with a tool policy or run the workflow in a restricted subagent. ## Skills as slash commands A valid enabled skill with `user-invocable` enabled appears as `/skill-name`. When you run the slash command, Droid loads the skill body and appends any text you typed after the command. ```text /summarize-diff focus on migration risk ``` If a custom slash command already uses the same name, the custom command keeps that slash command and the skill remains available for automatic selection by Droid. ## Best practices | Practice | Guidance | | :------- | :------- | | **Keep each skill narrow** | A good skill has one job. Prefer `review-api-change`, `summarize-diff`, or `prepare-ship-notes` over one large skill that tries to handle every engineering workflow. | | **Make the description precise** | Droid sees the description before it sees the full skill. Include the action, trigger, and boundary. Avoid generic names like `review` when the workflow is actually `review-api-change` or `review-security-diff`. | | **Write down success criteria** | Include the commands, checks, or evidence Droid should collect before calling the workflow done. If a skill can change files, say how to validate the change and what to do when validation fails. | | **Use supporting files deliberately** | Put checklists, schemas, examples, and small support scripts beside `SKILL.md` when they make the skill easier to execute. Keep source-of-truth product code in its normal project location. | | **Avoid secrets and private values** | Skills are often shared through repositories or plugins. Do not store API keys, private host names, customer data, or internal-only identifiers in `SKILL.md` or supporting files. | ## Troubleshooting Check that the file is named `SKILL.md`, that `name` and `description` are present, and that `enabled` is not `false`. If `user-invocable: false` is set, the skill is hidden from slash invocation by design. Confirm that `disable-model-invocation: true` is not set. Then tighten the description so it clearly says when to use the skill. If the skill lives deeper in a repository, Droid may discover it only after inspecting that project area. `allowed-tools` does not grant new tools. If the workflow needs editing, shell commands, or external API access, those tools still need to be available in the current session. Make skill names and descriptions more specific. Add NOT-for boundaries when two skills have similar triggers. Configure Droid behavior, model defaults, autonomy, and local preferences. Project-level instructions that complement skill-scoped workflows. # Custom Slash Commands Extend the CLI with reusable Markdown prompts or executable scripts that run from the chat input. Custom slash commands turn repeatable prompts or scripts into `/shortcuts` that you can run from chat. Droid scans `.factory/commands` folders and turns each registered file into a command. For new reusable workflows, prefer [Skills](/harness/skills); existing `.factory/commands` files keep working. ## Discovery and naming | Scope | Location | Purpose | | ----- | -------- | ------- | | **Workspace** | `/.factory/commands` | Project-specific commands shared with teammates. Overrides a personal command with the same slug. | | **Personal** | `~/.factory/commands` | Private or cross-project shortcuts available across sessions. | A file is registered only when it matches one of these rules: - It ends with `.md`. - Its first line starts with a `#!` shebang. Filenames are slugged: lowercase, spaces and non-URL characters become `-`, and the extension is dropped. `Code Review.md` becomes `/code-review`. Command files may be nested in subdirectories; Droid discovers them recursively. Invoke a command with `/command-name optional arguments`. Run `/commands` to open the Custom Commands manager for browsing, reloading, or importing. ## Markdown commands Markdown files render as a system notification that seeds Droid's next turn. Optional YAML frontmatter sets autocomplete metadata. | Frontmatter key | Purpose | | --------------- | ------- | | `description` | Overrides the generated summary shown in slash suggestions. | | `argument-hint` | Appends inline usage hints, such as `/code-review `. | Tool scoping is not available for custom commands. Use [Skills](/harness/skills) or [Custom Droids](/harness/subagents) for tool policy. `$ARGUMENTS` expands to everything typed after the command name. If you do not reference it, the body is sent unchanged. Positional placeholders like `$1` or `$2` are not supported in Markdown commands. Use `$ARGUMENTS` and parse the input inside the prompt. Markdown output is wrapped in a system notification so the next agent turn immediately sees the prompt. ## Executable commands Executable files must start with a valid shebang so the CLI can call the interpreter. ```bash #!/usr/bin/env bash set -euo pipefail echo "Preparing $1" npm install npm run lint echo "Ready to deploy $1" ``` - The executable receives positional arguments (`/deploy feature/login` sets `$1=feature/login`). - The script runs from the current working directory and inherits your environment. - Stdout and stderr (up to 64 KB) plus the script contents are posted back to the chat transcript for transparency. Failures still surface their logs. ## Managing commands {/* sweep-allow: term-bullets */} - **Edit or add** files directly in `.factory/commands`. In `/commands`, press `R` to reload, `I` to import, or `Esc` to close. - **Import** existing `.agents` or `.claude` commands: open `/commands`, press `I`, then use `Space` to select, `A` to toggle all, `Enter` to import, and `B` or `Esc` to return. Import scans `.agents/commands` and `.claude/commands` at the repo root and in `~`. It copies only `.md` files and skips files that already exist in `.factory/commands`. - **Remove** a command by deleting its file. Workspace commands take precedence, so deleting the repo version reveals the personal fallback if one exists. ## Examples ### Review checklist ```markdown --- description: Send a structured code review checklist argument-hint: --- Review `$ARGUMENTS` and respond with: 1. Summary of what changed and why it matters. 2. Correctness checks: tests, edge cases, and regressions. 3. Risks: security, performance, or migration concerns. 4. Follow-up TODOs with file paths and owners. ``` Run `/review feature/login-flow` to seed Droid with a consistent checklist. ### Deploy helper ```bash #!/usr/bin/env bash set -euo pipefail target=${1:-"src"} echo "Running lint and tests for $target" npm run lint -- "$target" npm test -- --runTestsByPath "$target" git status --short echo "Done" ``` Saved as `deploy.sh`, this shows up as `/deploy`. Pass a path (`/deploy src/widgets`) to constrain the checks and share the aggregated output in the thread. # Plugins Use and build shareable Droid packages for skills, commands, droids, output styles, hooks, and MCP servers. Plugins package reusable Droid extensions so a team can install the same skills, commands, droids, output styles, hooks, and MCP servers across projects. Use standalone `.factory/` configuration for personal or project-local experiments, then turn the pieces into a plugin when they should be shared, versioned, or governed. For org-approved catalogs and managed marketplace sources, see [Internal Plugin Marketplaces](/enterprise/internal-plugin-marketplaces). ## What plugins contain | Component | Location in a plugin | Loaded as | Use for | | :-------- | :------------------- | :-------- | :------ | | Skills | `skills//SKILL.md` | Model-invoked skill | Reusable procedures and domain knowledge. | | Slash commands | `commands/.md` or executable command files | `/name` command | User-invoked workflows. | | Droids | `droids/.md` | Task-callable subagent | Specialized agents with scoped prompts, tools, and models. | | Output styles | `output-styles/.md` | Settings picker option | Shared instructions for structuring interactive responses. | | Hooks | `hooks/hooks.json` plus scripts | Lifecycle hooks | Validation, policy, formatting, logging, and context injection. | | MCP servers | `mcp.json` | MCP tool configuration | External tools and data sources made available when the plugin is active. | Native Droid plugin layouts use a `.factory-plugin/` directory. Claude Code plugin layouts are also supported: `.claude-plugin/` is translated to `.factory-plugin/`, `agents/` is translated to `droids/`, and `.mcp.json` is translated to `mcp.json` when Droid copies the plugin into its cache. ## Manage plugins Use the interactive UI for browsing and one-off work: ```text /plugins ``` | Tab | Purpose | | :-- | :------ | | Available | Install plugins from registered marketplaces. Plugin IDs installed in any scope do not appear here. | | Installed | Enable or disable a plugin with Space where policy allows, and open its details or available actions. Org-scoped installs expose only their details. | | Marketplaces | Add or update a marketplace, and toggle its auto-update. Only user-added marketplaces can be deleted here. Remove org or project entries from the settings that declare them. | Use CLI commands for scripts and onboarding: ```bash # Marketplace management droid plugin marketplace add droid plugin marketplace list droid plugin marketplace update [name] droid plugin marketplace remove # Plugin management droid plugin install --scope user droid plugin install --scope project droid plugin list --scope user droid plugin update [plugin@marketplace] --scope project droid plugin uninstall --scope project ``` There is no `droid plugin enable` or `disable` command. Enablement is stored in the `enabledPlugins` setting, and organization-managed activation cannot be changed locally. Plugin IDs use `pluginName@marketplaceName`. Scoped npm package names are supported because Droid splits on the first `@` after the first character, so `@scope/plugin@marketplace` is valid. When Droid derives a marketplace name from a pinned source, it appends the pin to that name. `your-org/plugins#v1.2.0` registers as `plugins@v1.2.0`, and a plugin in it installs as `code-standards@plugins@v1.2.0`. A SHA pin uses the first 8 characters. Marketplaces declared in settings keep their `extraKnownMarketplaces` key instead. Run `droid plugin marketplace list` to read the registered name back. ## Build a plugin A minimal native plugin looks like this: ```text my-plugin/ ├── .factory-plugin/ │ └── plugin.json ├── commands/ │ └── hello.md ├── skills/ │ └── code-review/ │ └── SKILL.md ├── droids/ │ └── reviewer.md ├── output-styles/ │ └── review-notes.md ├── hooks/ │ ├── hooks.json │ └── check.sh ├── mcp.json └── README.md ``` Keep `commands/`, `skills/`, `droids/`, `output-styles/`, `hooks/`, and `mcp.json` at the plugin root. Do not put them inside `.factory-plugin/`; that directory is for plugin metadata. ### Plugin manifest Create `.factory-plugin/plugin.json`: ```json { "name": "my-plugin", "description": "A helpful plugin description", "version": "1.0.0", "author": { "name": "Your Team" }, "homepage": "https://github.com/your-org/my-plugin", "repository": "https://github.com/your-org/my-plugin", "license": "MIT", "keywords": ["review", "security"] } ``` | Field | Purpose | | :---- | :------ | | `name` | Source metadata for ecosystem compatibility. Installed identity comes from the marketplace entry. | | `description` | Source-level summary for anyone reading the plugin. | | `version` | Release metadata. Git-based installation still tracks the installed commit hash. | | `author`, `homepage`, `repository`, `license`, `keywords` | Optional source metadata for attribution and compatibility. | Droid does not require or read the fields inside `plugin.json`. It identifies an installed plugin from its marketplace entry name and registered marketplace name. A `.factory-plugin/` directory identifies the native layout; `.claude-plugin/`, `agents/`, or `.mcp.json` identifies a Claude Code layout for translation. The description and category shown while browsing come from the marketplace entry. ### Commands A Markdown command at `commands/review-pr.md` becomes `/review-pr`: ```markdown --- description: Review the current PR for issues --- Review the current pull request. Focus on: $ARGUMENTS ``` ### Skills A skill lives at `skills//SKILL.md`: ```markdown --- name: code-review description: Reviews code for correctness, safety, and maintainability. Use when reviewing code, checking PRs, or analyzing code quality. --- Check for logic errors, security issues, missing tests, and unclear ownership. Return specific, actionable findings. ``` ### Droids A droid lives at `droids/.md`: ```markdown --- name: reviewer description: Specialized code reviewer subagent model: inherit tools: ["Read", "Grep", "Glob"] --- You are a senior code reviewer. Report correctness, security, and test coverage issues with severity and evidence. ``` See [Custom Droids](/harness/subagents) for the full droid configuration surface. ### Output styles An output style lives at `output-styles/.md` and becomes an option under `/settings` → **Output style**: ```markdown title="output-styles/review-summary.md" --- name: Review Summary description: Format review results for the team --- Group findings by severity and include file references. ``` The frontmatter is optional. Output style bodies also support the plugin-root variables described under [Hooks](#hooks). Output styles apply only to interactive Droid CLI sessions. See [Output styles](/droid-cli/output-styles) for validation, trust, and fallback behavior. ### Hooks Plugin hooks live at `hooks/hooks.json` and may reference scripts in the plugin: ```json { "PostToolUse": [ { "matcher": "Create|Edit|ApplyPatch", "hooks": [ { "type": "command", "command": "${DROID_PLUGIN_ROOT}/hooks/check.sh", "timeout": 30 } ] } ] } ``` Droid expands `${DROID_PLUGIN_ROOT}`, `$DROID_PLUGIN_ROOT`, `${CLAUDE_PLUGIN_ROOT}`, and `$CLAUDE_PLUGIN_ROOT` to the installed plugin cache path when the plugin loads. See [Hooks](/harness/hooks#plugin-hooks) for hook events, input, and output behavior. ### MCP servers Plugin MCP servers use `mcp.json` at the plugin root: ```json { "mcpServers": { "my-api": { "command": "npx", "args": ["-y", "@example/mcp-server"], "env": { "API_KEY": "${MY_API_KEY}" } } } } ``` ## Test locally `marketplace add` takes a marketplace, not a plugin, so a plugin directory on its own is rejected with "This doesn't appear to be a valid marketplace." Wrap the plugin in a local marketplace and install from that: ```text title="local-marketplace.txt" my-marketplace/ ├── .factory-plugin/ │ └── marketplace.json # lists my-plugin with "source": "./my-plugin" └── my-plugin/ └── .factory-plugin/ └── plugin.json ``` ```bash droid plugin marketplace add ./my-marketplace droid plugin install my-plugin@my-marketplace --scope user ``` Before sharing, verify: - The plugin installs from a clean checkout. - Commands work with and without `$ARGUMENTS`. - Skills and droids have clear names and routing descriptions. - Output styles appear in `/settings`, and `/diagnostics` reports no style errors. - Hook scripts use absolute paths or plugin-root variables. - MCP servers do not embed secrets directly in `mcp.json`. - The README explains what the plugin does, when to use it, and prerequisites. ## Marketplaces A marketplace is a catalog of installable plugins. Droid reads `.factory-plugin/marketplace.json` first and falls back to `.claude-plugin/marketplace.json` for Claude Code compatibility. ```text your-marketplace/ ├── .factory-plugin/ │ └── marketplace.json ├── plugin-one/ │ └── .factory-plugin/ │ └── plugin.json └── plugin-two/ └── .factory-plugin/ └── plugin.json ``` ```json { "name": "your-marketplace", "description": "A collection of team plugins", "owner": { "name": "Platform Team" }, "plugins": [ { "name": "plugin-one", "description": "Description of plugin one", "source": "./plugin-one" } ] } ``` | Marketplace field | Required | Description | | :---------------- | :------- | :---------- | | `name` | Yes | Manifest metadata. The registered name is derived from the source or set by the `extraKnownMarketplaces` key. | | `description` | No | Optional manifest metadata. | | `owner` | No | Contact metadata. | | `plugins[].name` | Yes | Plugin identifier inside this marketplace. | | `plugins[].source` | Yes | Relative path string or source object. | | `plugins[].description`, `category`, `homepage`, `tags` | No | Browsing metadata. | ## Plugin sources Each marketplace plugin entry has a `source` field. | Source type | Shape | Use when | Pinning | | :---------- | :---- | :------- | :------ | | Relative path | `"./plugin-one"` | Plugin lives inside the marketplace repository. Absolute paths and paths that escape the marketplace directory are rejected. | Pin the marketplace source. | | `url` | `{ "source": "url", "url": "https://github.com/org/plugin.git" }` | Plugin has its own Git repository. | `ref` branch/tag or full 40-character `sha`. | | `git-subdir` | `{ "source": "git-subdir", "url": "https://github.com/org/repo", "path": "plugins/foo" }` | Plugin lives in a subdirectory of a larger repository. | `ref` branch/tag or full 40-character `sha`. | | `npm` | `{ "source": "npm", "package": "@org/plugin", "version": "^2.0.0" }` | Plugin is published to npm or a private npm-compatible registry. | `version` semver, range, or dist-tag. | The `github` source object registers a marketplace. For an external repository in `plugins[].source`, use `url`, or use `git-subdir` when the plugin occupies a subdirectory. `npm` is valid only inside a marketplace's `plugins[].source`. `droid plugin marketplace add npm:` is intentionally rejected. To distribute one npm plugin, create a small wrapper marketplace that lists the npm source. ### Add and pin marketplaces `droid plugin marketplace add` accepts `owner/repo` shorthand, a GitHub URL, any other Git URL, or a local path. Shorthand and plain `http(s)://github.com/owner/repo[.git]` URLs become a `github` source; other non-path inputs become a `url` source. Append `#` to follow a branch or tag, or `@<40-character-sha>` to pin an exact commit: ```bash droid plugin marketplace add Factory-AI/factory-plugins droid plugin marketplace add 'your-org/plugins#v1.2.0' droid plugin marketplace add https://github.com/Factory-AI/factory-plugins droid plugin marketplace add 'https://github.com/your-org/plugins#v1.2.0' droid plugin marketplace add https://github.com/your-org/plugins@1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7e8f9a0b droid plugin marketplace add ./your-local-marketplace ``` Local marketplace paths are resolved to absolute paths and cannot carry a ref or SHA pin. The Marketplaces tab in `/plugins` and the plugin settings in the Factory App accept the same forms. ### npm source details | Field | Required | Description | | :---- | :------- | :---------- | | `package` | Yes | npm package name. Scoped packages such as `@scope/name` are supported. | | `version` | No | Version, range, or dist-tag. Defaults to `latest`. URL, path, `file:`, `git+...`, and `npm:` alias specs are rejected. | | `registry` | No | HTTPS registry URL with no embedded credentials, query string, or fragment. | | `authTokenEnvVar` | No | Environment variable name containing the private registry token. Requires `registry` to have an effect. | Droid installs npm-source plugins in a per-plugin scratch directory with `npm install --ignore-scripts --no-save --no-audit --no-fund`, then copies the resolved package root into the plugin cache. Lifecycle scripts do not run, global npm configuration is not mutated, and the package must ship ready-to-use files. ## Team and enterprise distribution Register marketplaces and enable plugins through settings when you want teams to get them automatically: ```json { "extraKnownMarketplaces": { "your-org-internal-plugins": { "source": { "source": "github", "repo": "your-org/internal-plugins", "ref": "v1.2.0" } } }, "enabledPlugins": { "code-standards@your-org-internal-plugins": true } } ``` The installation scope follows where the setting is defined: org-managed settings install as org scope, user settings install as user scope, and project settings install as project scope. Use `strictKnownMarketplaces` in managed settings when an organization wants to restrict marketplace sources to an approved list. See [Internal Plugin Marketplaces](/enterprise/internal-plugin-marketplaces) for centralized governance. ## Versioning and updates | Source | Tracked version | Update behavior | | :----- | :-------------- | :-------------- | | Relative path inside a Git marketplace | Marketplace commit hash | Update the marketplace to move the plugin. | | `url` or `git-subdir` plugin source | Marketplace checkout commit hash only. The external commit is not recorded. | Updating the plugin clones the external source again at its configured `ref` or `sha`, or its default branch. | | `npm` plugin source | Resolved npm package version and metadata | Each plugin update resolves the `version` spec again. | | Plugin manifest `version` | Metadata only | Useful for humans and release notes, but Git-based installs are tracked by commit. | Automatic synchronization can skip an unchanged project context and a recently checked marketplace for up to six hours. An explicit `droid plugin marketplace update` requests an immediate update. ## Discover plugins Factory maintains an official marketplace at `Factory-AI/factory-plugins`: ```bash droid plugin marketplace add https://github.com/Factory-AI/factory-plugins ``` Common official plugins include: | Plugin | Purpose | | :----- | :------ | | [droid-control](/software-factory/droid-control) | Terminal, browser, and computer automation for demos, QA, and verification. | | droid-evolved | Skills for session navigation, writing, skill creation, design, frontend work, and browser automation. | | security-engineer | Security review, threat modeling, commit scanning, and vulnerability validation. | Droid can also install compatible Claude Code plugins. Claude Code layouts are translated during cache copy, not by mutating the source repository. ## Best practices | Practice | Why it matters | | :------- | :------------- | | Keep plugins focused | Small plugins are easier to review, compose, and retire. | | Document capabilities and boundaries | Users need to know when to use the plugin, prerequisites, and what data or tools it touches. | | Avoid hardcoded machine paths | Use plugin-root variables for plugin files and environment variables for secrets. | | Test on a clean machine | Plugin installs should not depend on uncommitted local files or global state. | | Ship prebuilt npm packages | `--ignore-scripts` means npm plugin sources cannot rely on postinstall builds. | | Govern high-trust plugins | Use managed settings and [Internal Plugin Marketplaces](/enterprise/internal-plugin-marketplaces) for approved enterprise distribution. | Define specialized subagents packaged inside a plugin. Govern approved plugin catalogs for your organization. Create response styles for users, projects, and plugins. Package reusable procedures inside a plugin. # Hooks Run deterministic shell commands at Droid lifecycle events for validation, policy, context injection, and automation. Hooks are shell commands that run at defined points in a Droid session. Use them for deterministic behavior that should happen every time, such as validating tool calls, formatting changed files, injecting local context, logging activity, or enforcing team policy. Hooks run automatically with your local environment and credentials. Review every hook command before registering it, use absolute script paths, and test in a safe environment first. ## Configuration Hooks live alongside settings at each scope: | Scope | File | Notes | | :---- | :--- | :---- | | User | `~/.factory/hooks.json` | Applies across your projects. | | Project | `.factory/hooks.json` | Commit to share with teammates. | | Enterprise | Org-managed settings | Loaded from Enterprise Controls and managed settings. | | Legacy | `.factory/hooks/hooks.json` | Still loads. The next save writes `.factory/hooks.json` and archives the old file as `hooks/hooks.migrated.json`. | If `hooks.json` is absent, Droid also reads hook declarations from the `hooks` key in the matching `settings.json`. Always use absolute paths in hook commands. Hooks execute from Droid's current working directory, which can differ from your repository root. Use `"$FACTORY_PROJECT_DIR"/path/to/script.sh` for project scripts or a full path such as `/usr/local/bin/script.sh` for global scripts. ### Structure Standalone `hooks.json` files are keyed directly by event name. Each event contains matcher groups, and each group contains one or more shell commands. ```json title="~/.factory/hooks.json" { "PreToolUse": [ { "matcher": "Execute", "commandRegex": "^git ", "hooks": [ { "type": "command", "command": "/usr/local/bin/audit-git-command.sh", "timeout": 30 } ] } ] } ``` When you declare hooks inside `settings.json` instead, wrap the same event map in a top-level `hooks` key: ```json title="~/.factory/settings.json" { "hooks": { "PreToolUse": [ { "matcher": "Execute", "commandRegex": "^git ", "hooks": [ { "type": "command", "command": "/usr/local/bin/audit-git-command.sh", "timeout": 30 } ] } ] } } ``` | Field | Required | Source-validated behavior | | :---- | :------- | :------------------------ | | `matcher` | No | Empty, omitted, or `*` matches everything. Exact strings match one tool or lifecycle matcher. Regex patterns are supported and are case-sensitive. | | `commandRegex` | No | Additional regex filter for Execute commands. It matches the actual shell command string when Droid has one. Invalid regex values are skipped. | | `hooks` | Yes | Array of hook commands for the matcher group. | | `type` | Yes | Currently only `"command"` is supported. | | `command` | Yes | Shell command executed with JSON hook input on stdin. | | `timeout` | No | Per-command timeout in seconds. Defaults to `60`. | Common tool matchers include `Execute`, `Read`, `Edit`, `Create`, `ApplyPatch`, `LS`, `Glob`, `Grep`, `Task`, `FetchUrl`, and `WebSearch`. MCP tools use the `mcp____` naming pattern, so `mcp__.*` matches all MCP tools. ## Quickstart This example logs every shell command Droid runs. In Droid, run: ```text /hooks ``` The manager opens with **User**, **Project**, **Plugins**, and **Effective** tabs. Select **User** to save the hook in `~/.factory/hooks.json`, or **Project** to save it in `.factory/hooks.json`. The **Plugins** and **Effective** tabs are read-only. Select **Add hook command**, choose the `PreToolUse` event, and set the matcher to `Execute`. Enter this command: ```bash jq -r '.tool_input.command' >> ~/.factory/bash-command-log.txt ``` Save the hook, ask Droid to run a simple command, then inspect the log: ```bash cat ~/.factory/bash-command-log.txt ``` The saved configuration looks like this: ```json title="~/.factory/hooks.json" { "PreToolUse": [ { "matcher": "Execute", "hooks": [ { "type": "command", "command": "jq -r '.tool_input.command' >> ~/.factory/bash-command-log.txt" } ] } ] } ``` If you chose the **Project** tab, the same unwrapped structure is saved to `.factory/hooks.json`. ## Hook events | Event | Runs when | Matcher or key fields | Common use | | :---- | :-------- | :-------------------- | :--------- | | `PreToolUse` | After Droid builds tool parameters and before the tool runs. | `tool_name`, `tool_input`; matcher usually targets a tool such as `Execute` or `Edit`. | Block risky operations, approve safe tools, rewrite tool input. | | `PostToolUse` | Immediately after a tool completes. | `tool_name`, `tool_input`, `tool_response`; same tool matchers as `PreToolUse`. | Format files, run validation, add feedback after an edit. | | `UserPromptSubmit` | Before Droid processes a submitted user prompt. | `prompt`, `has_images`. | Validate prompts or inject extra context. | | `Notification` | When Droid sends a notification. | `message`, `notification_type` (`permission_prompt`, `idle_prompt`, `auth_success`, `elicitation_dialog`). | Desktop alerts or compliance logging. | | `Stop` | When the main Droid is about to finish responding. | `stop_hook_active`, `tool_execution_count`, `elapsed_time`. | Require final checks or continue with follow-up instructions. | | `SubagentStop` | When a Task-launched sub-droid finishes. | `task_name`, `task_result`, `task_error`, `stop_hook_active`. | Validate subagent output or request more work. | | `PreCompact` | Before manual or automatic compaction. | `trigger` (`manual` or `auto`), `custom_instructions`, `message_count`, `estimated_tokens`. | Save context or add compaction guidance. | | `SessionStart` | When Droid starts, resumes, clears, or starts after compaction. | `source` (`startup`, `resume`, `clear`, `compact`), plus optional prior session IDs. | Load local context at session start. | | `SessionEnd` | When a Droid session ends. | `reason` (`clear`, `logout`, `prompt_input_exit`, `other`), `session_duration_ms`, `message_count`. | Cleanup, audit logs, session summaries. | ### Notification types | Type | Sent when | | :--- | :-------- | | `permission_prompt` | Droid is waiting for permission to run an action. | | `idle_prompt` | Droid is waiting for user input, including immediately after the user cancels a turn. | | `auth_success` | An authentication flow succeeds. | | `elicitation_dialog` | Droid needs structured input in an elicitation dialog. | Cancellation emits the informational `Notification` hook instead of `Stop`, so a hook cannot override the user's decision to stop the turn. ## Hook input Every hook receives JSON on stdin. Common fields are also exposed as environment variables when their values are strings, booleans, or numbers. ```typescript { session_id: string transcript_path: string cwd: string permission_mode: "off" | "spec" | "auto-low" | "auto-medium" | "auto-high" hook_event_name: string message_id?: string } ``` Tool hooks add tool-specific fields: ```json { "session_id": "abc123", "transcript_path": "/Users/.../.factory/projects/.../session.jsonl", "cwd": "/Users/me/project", "permission_mode": "off", "hook_event_name": "PreToolUse", "tool_name": "Create", "tool_input": { "file_path": "/Users/me/project/file.txt", "content": "file content" } } ``` `PostToolUse` includes `tool_response` as well. The exact shape of `tool_input` and `tool_response` depends on the tool. ## Hook output Hooks communicate through exit codes, stderr, stdout, and optional JSON emitted on stdout. | Output | Effect | | :----- | :----- | | Exit code `0` | Success. For `UserPromptSubmit` and `SessionStart`, stdout can add context. For other events, stdout is visible in transcript views. | | Exit code `2` | Blocking or corrective feedback. `PreToolUse` blocks the tool call, `PostToolUse` and `Stop` feed stderr back to Droid, and `UserPromptSubmit` blocks prompt processing. Other lifecycle events surface stderr to the user. | | Any other non-zero exit | Non-blocking error. Droid records stderr and continues where the event permits. | | JSON `continue: false` | Stops processing after the hook. `stopReason` can explain why to the user. | | JSON `suppressOutput: true` | Hides successful hook output from the main chat view while preserving it in the detailed transcript. | ### PreToolUse control Use `hookSpecificOutput.permissionDecision` for tool-call control: ```json { "hookSpecificOutput": { "hookEventName": "PreToolUse", "permissionDecision": "allow", "permissionDecisionReason": "Documentation reads are safe", "updatedInput": { "file_path": "/Users/me/project/README.md" } } } ``` | Decision | Effect | | :------- | :----- | | `allow` | Allows the tool call and can bypass the normal permission prompt. | | `deny` | Blocks the tool call and sends the reason to Droid. | | `ask` | Forces a user confirmation prompt. | `updatedInput` can modify tool parameters before execution. Use it carefully and only for fields you understand. ### PostToolUse, prompt, and stop control | Event | JSON fields | | :---- | :---------- | | `PostToolUse` | `decision: "block"` sends `reason` back to Droid after the tool ran. `hookSpecificOutput.additionalContext` adds more context. | | `UserPromptSubmit` | `decision: "block"` prevents the prompt from being processed and shows `reason` to the user. `additionalContext` is appended when not blocked. | | `Stop` and `SubagentStop` | `decision: "block"` prevents stopping. Include `reason` so Droid knows what to do next. | | `SessionStart` | `hookSpecificOutput.additionalContext` is appended to the new session context. | | `SessionEnd` | Cannot block session termination. Use it for cleanup and logging. | ## Plugin hooks Installed plugins can provide hooks. Droid loads hooks from `hooks/hooks.json` in the plugin root and merges them with user, project, and managed hooks when the plugin is enabled. ```text my-plugin/ ├── .factory-plugin/ │ └── plugin.json ├── hooks/ │ ├── hooks.json │ └── format.sh └── ... ``` ```json { "PostToolUse": [ { "matcher": "Create|Edit|ApplyPatch", "hooks": [ { "type": "command", "command": "${DROID_PLUGIN_ROOT}/hooks/format.sh", "timeout": 30 } ] } ] } ``` Plugin hook commands may use `${DROID_PLUGIN_ROOT}`, `$DROID_PLUGIN_ROOT`, `${CLAUDE_PLUGIN_ROOT}`, or `$CLAUDE_PLUGIN_ROOT`. Droid expands those variables to the installed plugin cache path when loading plugin hooks. Outside plugins, use `$FACTORY_PROJECT_DIR` or an absolute path. ## Org-managed hooks Organizations can define authoritative hooks in Enterprise Controls through managed settings. Org hooks always load unless hooks are globally disabled, and lower levels cannot remove them. | Field | Description | | :---- | :---------- | | `hooks` | Org-defined hook settings keyed by event name, using the same structure as `hooks.json`. | | `allowManagedHooksOnly` | When `true`, only org-managed hooks and hooks from org-enabled plugins load. User and project hooks are ignored. | ```json { "allowManagedHooksOnly": true, "hooks": { "PreToolUse": [ { "matcher": "Execute", "hooks": [ { "type": "command", "command": "/usr/local/bin/audit-command.sh", "timeout": 30 } ] } ] } } ``` See [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control) for how org, project, folder, and user settings merge. ## Examples ### Format changed TypeScript files ```json { "PostToolUse": [ { "matcher": "Create|Edit|ApplyPatch", "hooks": [ { "type": "command", "command": "python3 \"$FACTORY_PROJECT_DIR\"/.factory/hooks/format_changed_file.py" } ] } ] } ``` ### Block edits to sensitive files ```json { "PreToolUse": [ { "matcher": "Create|Edit|ApplyPatch", "hooks": [ { "type": "command", "command": "python3 \"$FACTORY_PROJECT_DIR\"/.factory/hooks/protect_paths.py" } ] } ] } ``` A blocking script exits `2` and writes the reason to stderr. ### Add context before each prompt ```json { "UserPromptSubmit": [ { "hooks": [ { "type": "command", "command": "python3 \"$FACTORY_PROJECT_DIR\"/.factory/hooks/add_ticket_context.py" } ] } ] } ``` ## Security considerations - Treat hook input as untrusted JSON. Validate and sanitize paths, prompts, and command strings. - Quote shell variables, for example `"$FACTORY_PROJECT_DIR"`, to avoid word splitting. - Block path traversal and sensitive paths such as `.env`, `.git/`, credentials, and deployment secrets. - Prefer small scripts checked into `.factory/hooks/` over long inline one-liners. - Test hooks manually before enabling them for a team or organization. - Remember that hooks inherit your local environment. Avoid commands that exfiltrate data or mutate systems unexpectedly. Droid snapshots hooks at startup and warns when hooks are modified externally. Review changes in the `/hooks` UI before relying on them in the current session. ## Debugging | Symptom | What to check | | :------ | :------------ | | Hook does not run | Confirm the event key, matcher casing, and that hooks are not disabled. Use `/hooks` to inspect the active configuration. | | Hook runs for the wrong tool | Matchers are case-sensitive and can be regexes. Use exact names like `Execute`, or a narrow regex. | | Script cannot be found | Use an absolute path or `"$FACTORY_PROJECT_DIR"/relative/path`. Do not rely on the shell's current directory. | | JSON parsing fails | Remember that hook input arrives on stdin. Test with sample JSON before registering the command. | | Hook hangs | Set a shorter `timeout`, inspect external network or process calls, and test the script outside Droid. | Run `droid --debug` to inspect hook matching and execution details. Fire hooks on Task-launched subagent lifecycle events. Combine hooks with autonomy levels for layered safety controls. # Factory for Enterprise Deploy, secure, govern, and observe Droid across cloud, hybrid, and fully airgapped environments. {/* cspell:ignore airgapped SDLC */} Factory helps enterprise teams run Droid across developer laptops, CI runners, VMs, Kubernetes clusters, and airgapped networks without losing central control. Use this section to plan where Droid runs, which models and tools it can use, how identity and policy are enforced, and how platform teams measure activity at scale. --- ## Enterprise playbook Start with the path that matches the decision you need to make. Choose cloud, hybrid, or airgapped deployment patterns, then configure proxies, certificates, mTLS, regions, and secure runtimes. Bound agent behavior with command policies, hooks, sandboxing, Droid Shield, and CLI security defaults. Design organization hierarchy, then manage identity, Enterprise Controls, model access, MCP servers, plugin registries, and service accounts. Export OpenTelemetry metrics, use the Analytics API, and map Droid activity to audit and compliance workflows. --- ## Core model Droid is a local agent runtime governed by managed settings and observable through product telemetry. {/* sweep-allow: term-bullets */} - **Deployment is flexible.** Droid runs where your code already lives, with [cloud, hybrid, airgapped, EU, proxy, and certificate patterns](/enterprise/network-and-deployment). - **Data paths are explicit.** Code and files stay local unless selected context is sent to a configured model or gateway, as described in [Data Flows & Privacy](/enterprise/privacy-and-data-flows). - **Controls are centralized.** [Organization Model](/enterprise/organization-model) defines the enterprise root and organization hierarchy, while [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control) define model access, command policies, MCP allowlists, hooks, sandbox policy, retention, and feature toggles. - **Safety is layered.** Deterministic controls, hooks, sandboxing, and Droid Shield reduce risk before model behavior matters, as covered in [Agent Safety & Controls](/enterprise/llm-safety-and-agent-controls). - **Measurement is structured.** Droid emits OpenTelemetry metrics and exposes hosted analytics for adoption, activity, and cost reporting, per [Telemetry & Analytics](/enterprise/telemetry). These controls let teams adopt Droid incrementally while keeping higher-risk workloads inside stricter environments. --- ## Trust Factory maintains a security and compliance program for enterprise procurement and regulated deployments: | Program | Scope | | --- | --- | | SOC 2 Type II | Security and availability | | ISO 27001 | Information security | | ISO 42001 | AI management systems | Find current reports, subprocessor lists, and security architecture details in the Trust Center. For audit mapping and regulated-deployment patterns, see [Compliance & Audit](/enterprise/compliance-audit-and-monitoring). # Deployment Patterns Choose cloud, hybrid, EU, and airgapped deployment patterns, then configure network access for Droid. Droids are designed to run anywhere: on laptops, in CI pipelines, on VMs and Kubernetes clusters, and in fully airgapped environments. Choose the deployment pattern that matches your data boundary, then combine Droid with proxies, custom CAs, mTLS, sandboxed containers, and managed settings. --- ## Choose a pattern Factory supports three canonical patterns. The difference between them is **how far data travels and where the boundary sits**. The runtime is the same in every case. You can mix patterns across teams and environments. Use this table to pick a starting point, then read the notes below for the trade-offs. | Pattern | Where Droid runs | Factory cloud involvement | LLM and telemetry traffic | | ------------------- | --------------------------------------------------------- | ---------------------------------------------------------------- | -------------------------------------------------------------------------- | | **Cloud-managed** | Developer laptops, CI/CD runners, optional devcontainers | Control plane, org metadata, web authentication, optional analytics | LLM traffic to your providers or gateways; OTEL metrics to your collectors | | **Hybrid** | Your VMs, containers, CI runners, and remote dev environments | Limited metadata only if you enable cloud features | Internal LLM gateways and OTEL/SIEM endpoints inside your network | | **Fully airgapped** | Isolated network with no outbound internet connectivity | None at runtime; artifacts imported via offline processes | On-prem or in-network model endpoints and collectors only | ### Cloud-managed Droid runs on developer machines and build infrastructure, while **Factory cloud** provides orchestration and optional analytics. LLM traffic can still be routed through **your own gateways and providers**; Factory does not need to broker model access. This pattern suits organizations that allow well-scoped cloud usage but want strong governance over models, keys, and telemetry. ```mermaid %%{init: {"theme": "base", "themeVariables": {"fontFamily": "Geist Mono, monospace", "fontSize": "13px", "primaryColor": "#161413", "primaryBorderColor": "#342F2D", "primaryTextColor": "#FAFAFA", "lineColor": "#4D4947", "textColor": "#D6D3D2"}}}%% flowchart LR d["Droid laptops / CI"] d -->|"LLM"| l["Your gateways"] d -->|"OTEL"| o["Collectors"] d -->|"orchestration"| f["Factory cloud"]:::muted classDef muted stroke-dasharray:4 3,stroke:#8A8380,color:#D6D3D2; ``` ### Hybrid Droid runs entirely within your infrastructure (on your VMs, containers, CI runners, and remote dev environments) while you may still use Factory cloud selectively for coordination. LLM traffic goes through your gateways or cloud providers under your accounts, and OTEL telemetry stays in your observability stack. Factory cloud sees only limited metadata (for example, org and project identifiers) if you enable cloud features. This pattern is common where **network segmentation** and **central governance** are mandatory. ```mermaid %%{init: {"theme": "base", "themeVariables": {"fontFamily": "Geist Mono, monospace", "fontSize": "13px", "primaryColor": "#161413", "primaryBorderColor": "#342F2D", "primaryTextColor": "#FAFAFA", "lineColor": "#4D4947", "textColor": "#D6D3D2"}}}%% flowchart LR d["Droid in your infra"] d -->|"LLM"| l["Internal gateway"] d -->|"OTEL"| o["Internal SIEM"] d -.->|"optional metadata"| f["Factory cloud"]:::muted classDef muted stroke-dasharray:4 3,stroke:#8A8380,color:#D6D3D2; ``` ### Fully airgapped Droid runs in an isolated network with no outbound internet connectivity. Models are served from on-prem or in-network endpoints, and OTEL collectors live entirely inside the airgap. Factory cloud is not reachable at runtime; binaries and configuration are imported through your own artifact repositories or offline processes. This is the default pattern for national security, defense, and other highly classified workloads. ```mermaid %%{init: {"theme": "base", "themeVariables": {"fontFamily": "Geist Mono, monospace", "fontSize": "13px", "primaryColor": "#161413", "primaryBorderColor": "#342F2D", "primaryTextColor": "#FAFAFA", "lineColor": "#4D4947", "textColor": "#D6D3D2"}}}%% flowchart LR d["Droid in isolated network"] d -->|"LLM"| l["On-prem models"] d -->|"OTEL"| o["In-network collectors"] a["Offline artifacts"] --> d ``` In every pattern, tighten the boundary with the same levers: restrict outbound hosts to the minimum set of domains (cloud-managed deployments reach `*.factory.ai`), route all LLM traffic through a central gateway you monitor, and apply the HTTPS proxy and custom-CA settings below. Hybrid and airgapped runtimes inherit your existing firewall, VPN, and Kubernetes network controls. ## Proxies, custom CAs, and mTLS Enterprise networks frequently require HTTP(S) proxies, organization-specific certificate authorities, and mutual TLS. ### HTTP(S) proxy support Droid respects standard proxy environment variables: ```bash export HTTPS_PROXY="https://proxy.example.com:8080" export HTTP_PROXY="http://proxy.example.com:8080" # Bypass proxy for specific hosts export NO_PROXY="localhost,127.0.0.1,internal.example.com,.corp.example.com" ``` Use these to route traffic from Droid to LLM gateways and any Factory cloud endpoints through your corporate proxy. ### Custom certificate authorities If your organization uses custom CAs for HTTPS inspection or internal endpoints, configure the runtime environment so Droid trusts those CAs (for example, via `NODE_EXTRA_CA_CERTS` or OS-level trust stores). ### Mutual TLS (mTLS) For environments that require client certificates when calling gateways or internal APIs, configure your containers, VMs, or runners with the appropriate certificate, key, and passphrase. These settings are usually handled at the HTTP client or proxy layer that Droid uses. --- ## Running in secure containers and VMs Running Droid inside hardened containers and VMs is one of the most effective ways to **bound the blast radius** of any agent mistakes or misconfigurations. | Runtime | Recommended boundary | | --- | --- | | Devcontainers | Lock down filesystem mounts and outbound network rules, and reserve higher-autonomy runs for these containers rather than the host. | | Isolated VMs | Use dedicated VMs for production-adjacent work such as migration tooling, with OS policies that limit which repos, secrets, and networks they reach. | | CI/CD pipelines | Run Droid in ephemeral jobs with short-lived credentials and minimal privileges, paired with hooks and Droid Shield. | Run Droid with higher autonomy levels only inside a sandboxed container or VM, never directly on a developer host or a machine with standing production credentials. See [Agent Safety & Controls](/enterprise/llm-safety-and-agent-controls) for how autonomy, allow/deny lists, and Droid Shield interact with these environments. ## EU deployment and data residency Factory operates a dedicated EU deployment, fully isolated from the Global (US) deployment, with its own backend, database, and LLM inference endpoints in Europe. Each Factory client (CLI, desktop app, and web app) ships as a single build that works in both regions; after login, the client detects the organization's region and routes all subsequent requests to the correct deployment. **Enterprise Feature**: The EU deployment is available to enterprise customers with European data-residency requirements. Contact Sales to provision an EU organization for your team. ### Infrastructure The EU API backend is hosted within the European Economic Area and reachable at `api.eu.factory.ai`; the EU web app is served from `app.eu.factory.ai`. Customers should allowlist `*.factory.ai` hostnames. Outbound LLM traffic from the EU backend targets EU-region provider endpoints only. ### Data residency - **Stored in the EU:** session content (prompts, assistant messages, tool calls and results, Droid messages, Git AI notes, and Git AI checkpoints) lives in a managed database instance in Europe, provisioned separately from the global database with its own credentials, backups, and access controls. Raw request and response bodies are never written to logs or object storage. - **Stored in the global deployment (US):** organization records (a small pointer document with the org's region tag), user profiles, and billing data. These contain no user-generated content and are required for cross-region authentication and account management. ### Inference All LLM requests from EU organizations are dispatched to EU-region endpoints only; prompts and conversation history do not transit US infrastructure. Models or providers not available in the EU are hidden in the model picker and rejected at the server. See [Models](/models) for the current catalog. ### Region enforcement Identity is global (WorkOS-based SSO), but each organization is pinned to a region at creation time. Region matching is enforced at multiple layers so traffic and data can never be served by the wrong backend. - **Server-side.** Session data access and LLM inference are both gated by region-aware checks. Requests for an organization whose region does not match the deployment fail closed, and requests for ineligible models or providers are rejected before any traffic is dispatched. EU traffic cannot reach US-only data or providers, and Global traffic cannot reach EU-only routes. - **Client-side.** After login, each client reads the org's region and routes all subsequent API requests to the matching backend. Each canonical web hostname (`app.factory.ai`, `app.eu.factory.ai`) is statically backed by exactly one regional backend; users who land on the wrong region are shown a redirect page (not signed out). The desktop app targets the matching API host on every request, and the CLI pins the region to the local token cache on first login, re-deriving it on logout and re-login. ## Configuration surfaces Network and deployment configuration is expressed through: Proxies, gateways, OTEL endpoints, custom certificates. Managed settings for model access, sandbox network rules, IP restrictions, feature gates, and defaults. For large organizations, a central configuration service can distribute a standard `.factory` bundle to all environments. For MDM-deployed and airgapped fleets, drop a settings file at a hardcoded platform path so org policy is in effect before any user authenticates. See [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control#system-managed-settings-file). That same reference covers the full settings hierarchy and merge behavior. Where code, prompts, and telemetry travel. Enforce proxy, custom-CA, and network policies through managed settings. # Airgapped Deployment Run Droid inside an isolated network with no outbound internet connectivity, using your own model endpoints, settings, and collectors. Droid ships in a dedicated airgap build for networks that have no route to the public internet. The build runs the full agent harness locally and contacts only endpoints you configure. This page describes how Airgap Mode behaves, the network contract it follows, and what to configure across a fleet. For choosing between cloud-managed, hybrid, and airgapped patterns, see [Deployment Patterns](/enterprise/network-and-deployment). Airgap builds are an enterprise capability and are not distributed through the public download channels. Contact Sales or your Factory representative for access, then import the binaries through your own artifact repositories. ## How Airgap Mode behaves Airgap builds run in Airgap Mode automatically. In this mode, Droid: - **Skips sign-in.** There is no login flow and no Factory account requirement; identity stays local to the machine. - **Offers your models only.** The model picker lists the custom models from your settings and nothing else. Factory-managed models and the Factory Router are unavailable; all inference goes to the endpoints you configure through [BYOK](/model-independence/byok). - **Skips Factory-bound work.** Update checks, release-notes fetches, session sync, crash reporting, and Factory telemetry are all disabled rather than left to time out. - **Reads policy from your sources.** Org policy comes from the [system-managed settings file](/enterprise/hierarchical-settings-and-org-control#system-managed-settings-file) or a managed-settings endpoint you host inside the network. - **Exports telemetry to your collector only.** With no collector configured, the telemetry pipeline does not run. See [Airgapped deployments in Telemetry & Analytics](/enterprise/telemetry#airgapped-deployments). An airgapped installation needs at least one custom model. Without one, Droid refuses to start a turn and asks you to configure a custom model in `settings.json`. ## Network contract In airgapped operation, Droid makes no requests to Factory-operated services during interactive and headless (`droid exec`) use: no Factory API, LLM proxy, telemetry ingest, crash reporting, update checks, release-notes fetches, documentation lookups, or login services. The outbound traffic in an airgapped session is the traffic your configuration creates. ### Destinations Droid uses | Destination | Where the address comes from | | -------------------------- | ------------------------------------------------------------------ | | Model endpoints | `customModels[].baseUrl` in your settings | | MCP servers | Your user or org MCP configuration | | Git remotes | The repositories you work in | | Managed-settings endpoint | An org configuration URL you host (optional) | | OTEL collector | [`telemetry.endpoint`](/enterprise/telemetry) in managed settings | | Plugin marketplaces | Marketplace sources in your org or user settings | | Commands the agent runs | The agent's Execute tool, bounded by your sandbox network policy | ### Enforcement layers Airgap Mode is a reliability contract, not an enforcement mechanism. The agent can run arbitrary commands, and an in-network gateway looks the same as a public API to the harness, so isolation must come from your infrastructure: | Layer | Owner | Guarantee | | ---------------------- | ----------------------------------------- | ---------------------------------------------------------------------------------- | | Network isolation | You | The boundary. Holds regardless of how any software inside it behaves. | | Sandbox network policy | You configure it, Factory ships it | Kernel-enforced egress rules for commands the agent runs. | | Airgap Mode | Factory | Cooperative. Droid skips Factory-bound work and stays reliable offline. | Keep your firewall and sandbox rules scoped to the destinations in the table above. Droid does not need any other egress to operate. ## Unavailable in Airgap Mode Features that depend on Factory cloud are unavailable in airgapped deployments: - Factory account, billing, and usage-limit management - Session sharing and cross-device session sync - Slack integration - Agent readiness reports - Bug-report upload - Built-in web search and URL-fetch tools - In-product release notes - Automatic updates: import new builds through your artifact process instead - Factory-hosted analytics: use [OTEL export](/enterprise/telemetry) to your own stack Everything local works as in connected deployments: the full agent loop and tools, missions, specification mode, hooks, MCP servers, and installed plugins and skills, with all model traffic going to your endpoints. ## Configuring an airgapped fleet Drop org policy at the hardcoded platform path so it applies before any user session, with no Factory connectivity required. See [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control#system-managed-settings-file). Ship `customModels` entries in managed settings. Org-distributed models use a keyless endpoint or [`apiKeyHelper`](/model-independence/byok#dynamic-credentials-with-apikeyhelper) rather than static keys. Pin `telemetry.endpoint` to your in-network collector in managed settings. See [Telemetry & Analytics](/enterprise/telemetry). Standard proxy variables and custom CA trust apply to in-network endpoints. See [Deployment Patterns](/enterprise/network-and-deployment). Choose between cloud, hybrid, and airgapped patterns. Configure the custom models an airgapped install requires. Export metrics to your own collector. Distribute policy through the system-managed settings file. # Data Flows & Privacy Understand where code, prompts, model traffic, and telemetry move across cloud, hybrid, and airgapped Droid deployments. High-security organizations need precise answers to **what data goes where, when, and under whose control**. Factory's answer is simple: **data boundaries are determined by the models, gateways, and deployment pattern you choose**, and Droid is configurable to respect those boundaries. --- ## Overview: three main data flows When you run Droid, there are three primary ways data can move: 1. **Code and files**: local reads and writes on your filesystem and git repositories. 2. **LLM traffic**: prompts and context sent to model providers or LLM gateways. 3. **Telemetry**: OpenTelemetry metrics by default, with optional message content trace spans sent only to a customer-configured collector. How far each flow travels depends on whether you are in a **cloud-managed**, **hybrid**, or **fully airgapped** deployment. | Deployment pattern | Data boundary | | --- | --- | | Cloud-managed | Droid runs on laptops and CI, talking to Factory's cloud for orchestration and, optionally, analytics. LLM requests go to your chosen model providers or LLM gateways. | | Hybrid | Droid runs entirely inside your infra. LLM and OTEL traffic terminate in your network; Factory cloud may only see minimal metadata if you enable it. | | Fully airgapped | Droid, models, and collectors all live inside an isolated network. No runtime dependency on Factory cloud; **no traffic leaves the environment**. | ## Code and file access Droid is a filesystem-native agent, and this holds in every deployment pattern: - It reads your code, configuration, and test files **directly from the local filesystem** at the moment they're needed. - It uses LLMs to analyze existing code and generate new code, then applies patches on disk, with git as the source of truth where available. - It does **not** upload or index your codebase into a remote datastore; there is no static or "cold" copy of your repository stored in Factory cloud. - The **agent loop and runtime execute entirely on the machine** where Droid runs (developer workstation, CI runner, VM, or container), and **file reads and writes remain local** to that environment; any file contents sent off-machine are sent as part of LLM requests (prompts and context) to the model endpoints you configure, following the same LLM request pipeline described below. - Hooks run locally inside that agent loop unless you configure them to call an external service. Droid Shield's default source-backed protection scans staged git diffs before `git commit` and `git push` operations. You control which files and directories Droid can see through standard OS permissions and your repository layout. See [Agent Safety & Controls](/enterprise/llm-safety-and-agent-controls) for additional file-level protections. --- ## LLM requests and model-specific guarantees When Droid needs model output, it sends prompts and context to **your configured models and LLM gateways**. By default, Droid can target **enterprise-grade endpoints** for providers like Azure OpenAI, AWS Bedrock, Google Vertex AI, OpenAI, Anthropic, and Gemini, using contracts that support zero data retention and enterprise privacy controls. In these configurations, Factory routes traffic **directly** to the provider's official APIs; we do not proxy this traffic through third-party services or store prompts and responses in Factory cloud. If you instead configure your own endpoints - LLM gateways, self-hosted models, or generic HTTP APIs - the privacy guarantees are entirely determined by those systems and your agreements with them; Factory does not add additional protections on top. ### Model and gateway options OpenAI, Anthropic, Google, and others via their enterprise offerings. AWS Bedrock, GCP Vertex, Azure OpenAI, using your cloud accounts. Models served inside your network or airgapped environment via HTTP/gRPC gateways. Central gateways that normalize traffic, add authentication, enforce rate limits, and log usage. ### How Droid interacts with models - Org and project policies decide **which models and gateways are allowed**, and whether users can bring their own keys. - In high-security settings, orgs commonly: - Prefer the direct enterprise endpoints described above to get first-party zero-retention guarantees. - Disable ad-hoc user-supplied keys and generic internet endpoints. - Treat any LLM gateways or self-hosted models as in-scope security systems, subject to the same reviews and monitoring as other critical services. See [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control) for configuration details. --- ## Telemetry and analytics Telemetry is how you understand _what_ Droid is doing and _where_. Factory uses OpenTelemetry metrics for customer-owned observability and the hosted Analytics API for aggregated cloud analytics. Customers can optionally export message content as trace spans to their own collector. ### OTEL as the source of truth - Droid exports **OTLP metrics** by default to the endpoint pinned in org-managed settings (`telemetry.endpoint`) or configured per machine with `OTEL_TELEMETRY_ENDPOINT`. You can also [opt in to message content trace spans](/enterprise/telemetry/privacy#message-content-logging), which are sent only to your collector. - **Spans never reach Factory's collector**, in any configuration. Droid builds a trace pipeline only where you have configured a collector of your own, so with none set no span is recorded at all. - Typical destinations include OTEL collectors feeding **Prometheus, Datadog, New Relic, Splunk**, and similar systems. - Customer metrics export runs as a fan-out alongside Factory's own metrics export in connected deployments. In [airgapped deployments](/enterprise/telemetry#airgapped-deployments) the fan-out has one leg: your collector, with no Factory-bound export built at all. Airgapped deployments do not use Factory-hosted analytics. - Organizations that cannot collect per-individual analytics can pin [`telemetry.granularity: aggregate`](/enterprise/telemetry/privacy#data-granularity), which strips user identifiers from every datapoint at the source on both legs. ### Optional Factory cloud analytics In cloud-managed deployments, you can opt into Factory's own analytics dashboards, which may: - Aggregate anonymized usage metrics to show adoption, model usage, and cost trends. - Surface per-org and per-team insights for platform and leadership teams. These analytics are optional; org administrators decide whether to enable them. For telemetry setup and the hosted Analytics API, see [Telemetry & Analytics](/enterprise/telemetry); for the exported metric names and schema, see the [Telemetry Data Reference](/enterprise/telemetry/data-reference); for how this data maps to audit and compliance workflows, see [Compliance & Audit](/enterprise/compliance-audit-and-monitoring). ## Data retention and residency ### Factory-hosted components In cloud-managed mode, Factory may store limited operational logs and metrics for: - Authentication and administrative actions (for example, org configuration changes). - Service health and debugging. Retention and residency for these logs are documented in the Trust Center and can be tuned per customer engagement. ### Customer-hosted components For hybrid and airgapped deployments: Handled entirely by your providers and gateways; Factory does not see it. Stored in your observability stack; retention is governed by your own policies. Never leave your environment unless you explicitly send them to external services via hooks or gateways. In fully airgapped environments, Factory never receives any runtime data; you are responsible for all retention and residency decisions. See [Airgapped Deployment](/enterprise/airgapped-deployment) for the runtime behavior and network contract. --- ## Governance controls in practice To make these guarantees enforceable rather than aspirational, Factory exposes **governance levers** at the org and project levels: - Model and gateway allow/deny lists. - Policies for whether user-supplied BYOK keys are allowed. - Global and project-specific hooks for DLP, redaction, and approval workflows. - OTEL endpoint configuration (including requiring on-prem collectors). - Maximum autonomy level and other safety controls, especially in less-trusted environments. These levers are implemented through [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control). Proxies, custom CAs, and regional data boundaries. Export OTEL metrics and optional message content traces. # Agent Safety & Controls Control Droid with command lists, hooks, sandboxing, IP restrictions, and Droid Shield secret scanning in high-security environments. Factory treats LLMs as powerful but untrusted components. Droid's safety model combines **deterministic controls** that do not depend on model behavior with **Droid Shield** secret scanning for git commit and push operations. Even a frontier model operating with high autonomy stays inside the boundaries you set. --- ## Two layers of safety - Command allow / deny / block lists - Programmable hooks (9 lifecycle events) - Built-in sandboxing (filesystem + network) - Network egress allowlists - Secret and credential scanning - `git commit` and `git push` protection - Learned classification models Private Preview - Fail-closed behavior when a diff is too large to scan - Hook-based DLP integration for custom checks Build enterprise security primarily on deterministic controls. They are configured through [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control), so hard policy controls such as model access, command blocklists, sandbox rules, and managed hooks apply consistently across laptops, CI, VMs, and airgapped environments. ## Command controls Droid evaluates every shell command it proposes against three lists, all defined in managed settings and accumulated across levels (org entries can never be removed by projects or users): Patterns that are always allowed to run without additional approval. Patterns that always require explicit confirmation. A denylisted command can still run if the user approves it. Patterns that can **never** run. Unlike the denylist, a blocked command has no approval path: the block holds even under full autonomy, auto-run, or `--skip-permissions-unsafe`. The **blocklist is the hard stop.** Before matching, Droid resolves the actual program being invoked, so a blocked command cannot be bypassed with a wrapper shell, an absolute path, quoting tricks, or command substitution. Use it for organization-mandated prohibitions; use the denylist for commands that need a human in the loop, and the allowlist to pre-clear safe, common commands. ```json { "commandAllowlist": ["npm *", "pnpm *", "make *"], "commandDenylist": ["sudo *", "rm -rf *"], "commandBlocklist": ["mkfs", "shutdown", "curl"] } ``` Command risk is also emitted via OTEL, so security teams can monitor how often high-risk commands are proposed or attempted. --- ## Hooks Hooks are Droid's programmable enforcement and observability interface: they run your own code at defined points in the agent loop. There are **nine hook events**: | Event | Fires when | | --- | --- | | `PreToolUse` | Before a tool runs (including shell commands, file edits, and git operations). | | `PostToolUse` | After a tool completes. | | `UserPromptSubmit` | When the user submits a prompt, before it reaches the model. | | `Notification` | When Droid emits a notification. | | `Stop` | When the main agent finishes responding. | | `SubagentStop` | When a subagent finishes. | | `PreCompact` | Before the context is compacted. | | `SessionStart` | When a session starts. | | `SessionEnd` | When a session ends. | There are no separate "pre-git", "pre-command", or "post-edit" hooks. Those are `PreToolUse` and `PostToolUse` hooks scoped with a **matcher** that selects which tools (or command patterns) the hook applies to. A hook controls the agent through its output. `PreToolUse` hooks can return a `permissionDecision` of `allow`, `deny`, or `ask`. More broadly, hook output can set `decision` (`block` or `approve`), `continue`, `suppressOutput`, a `systemMessage`, and structured `hookSpecificOutput`. This lets you, for example, block direct `git push` operations and redirect developers to internal PR tooling, or forward prompts and file snippets to a DLP or CASB API and deny operations that violate policy. Set `allowManagedHooksOnly` in org settings to ensure only org-managed hooks run; project- and user-defined hooks are then ignored. See the [Hooks Guide](/harness/hooks) for the full event payloads, matcher syntax, and output schema. --- ## Sandboxing Droid includes a built-in **sandbox** that isolates command execution and file access, so you can grant more autonomy without exposing the host. Configure it under the `sandbox` setting: Turn sandboxing on. Sandbox scope: `per-command` (isolate each command) or `whole-process` (run the whole session sandboxed). Path access lists: `allowRead`, `allowWrite`, `denyRead`, `denyWrite`. Use deny lists to keep secrets such as `~/.ssh` and `~/.aws` out of reach. Network access: `allowedDomains` plus `allowUnixSockets`, `allowAllUnixSockets`, `allowLocalBinding`, and `httpProxyPort` / `socksProxyPort` for routing egress through a proxy. ```json { "sandbox": { "enabled": true, "mode": "per-command", "filesystem": { "denyRead": ["~/.ssh", "~/.aws"], "denyWrite": ["/etc"] }, "network": { "allowedDomains": ["api.github.com", "registry.npmjs.org"] } } } ``` Sandboxing is a first-class product feature, not just a recommendation to run Droid inside Docker. You can still combine it with external devcontainers or VMs for defense in depth, and tag OTEL sessions by environment (for example, `environment.type=local|ci|sandbox`) for environment-specific dashboards and alerts. --- ## Network controls The built-in sandbox controls command-level network egress, while Enterprise Controls can also restrict which client IPs may access Factory APIs and applications. Domains sandboxed commands may reach during a session. IPv4 addresses or CIDR ranges allowed to access Factory web, CLI, and API key surfaces. Must contain at least one entry when set. Use sandbox network rules and corporate proxies for outbound restrictions. Use `networkPolicy.allowedIps` to reduce where authenticated Factory requests can originate. ## Droid Shield: git secret scanning **Droid Shield** (`enableDroidShield`) scans staged diffs before `git commit` and `git push`. It blocks the command when it detects likely secrets, and it fails closed when a diff is too large to scan safely. For detection behavior, false positives, and recovery steps, see [Droid Shield](/autonomy-and-safety/droid-shield). Use hooks when you need custom DLP, approval workflows, prompt inspection, or calls to internal security systems. Org admins can enforce Droid Shield and managed hooks at the org or project level; users cannot disable controls that are locked by those settings. --- ## Steering as a complement Deterministic controls are the foundation; **LLM steering** reduces how often dangerous actions are even proposed. Org and project settings can define rules and instructions (security guidelines and coding standards applied to every request), standardized commands and workflows (for example, `/security-review`), and context enrichment through allowlisted MCP servers so models work from accurate information. Because these are instructions, not enforcement, they complement rather than replace the hard boundaries above. For the full schema behind every control on this page, see [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control). Configure the managed settings schema for command lists, sandbox, and hooks. Deep dive into secret scanning for git commit and push operations. # Security How the Droid CLI protects code and data with local execution, command approval, sandboxing, and enterprise controls. ## Security-first design The Droid CLI runs where your code already lives. Shell commands and file edits execute locally, while prompts and selected context follow the model and gateway configuration set by you or your organization. Use this page for user-facing CLI protections. For organization-wide enforcement, see [Enterprise Controls](/enterprise/hierarchical-settings-and-org-control) and [Agent Safety & Controls](/enterprise/llm-safety-and-agent-controls). --- ## Key security features OAuth login with encrypted token storage. Tokens auto-rotate every 30 days and are stored with OS-level file permissions. Risky operations require explicit approval unless your settings allow them. Organizations can enforce command allowlists, denylists, and blocklists. Droid reads and edits files on the machine where it runs. It does not require uploading or indexing a static copy of your repository. Org-managed settings control models, MCP servers, hooks, plugins, sandboxing, retention, and feature access. Optional isolation for filesystem and network access when running commands. Command policies, hooks, and Droid Shield bound agent behavior before model output matters. --- ## Security best practices Always review suggested code and commands before approval. You control what Droid can access and execute. {/* sweep-allow: term-bullets */} - **Review before approving.** Verify proposed commands and file changes, especially package installs or system-file edits, operations involving sensitive data or credentials, network requests to external services, and file operations outside your project directory. - **Use isolated environments.** Run Droid in containers or VMs for untrusted code repositories, external APIs or web services, experimental or potentially risky operations, and shared development environments. - **Manage permissions carefully.** Block commands that should never run, use "ask" for medium-risk operations that need oversight, "allow" only low-risk commands you trust completely, and review permissions regularly in the Settings menu. - **Protect sensitive data.** Never include secrets in prompts: use environment variables for API keys and tokens, store credentials in secure credential managers, exclude sensitive files from Droid's working directory, and use `FACTORY_API_KEY` for CI/CD and service-account automation. --- ## Built-in protections The Droid CLI includes multiple layers of security: Droid file-edit tools are scoped to the active project directory and subdirectories. Risky operations require explicit user confirmation. Managed allowlists, denylists, and blocklists constrain shell execution. Secret scanning can block `git commit` and `git push` when staged diffs contain likely credentials. Sandbox rules and enterprise IP restrictions reduce where commands and Factory requests can go. Each conversation maintains separate, secure context. --- ## Enterprise security SAML/OIDC single sign-on, Directory Sync, roles, and service accounts. Enterprise Controls for model access, BYOK policy, MCP allowlists, hooks, plugins, sandboxing, and retention. Current reports, subprocessor lists, and security materials are available in the Factory Trust Center. OpenTelemetry metrics, hosted analytics, audit events, and workflow logs for enterprise monitoring. --- Report security vulnerabilities through our responsible disclosure program. Contact [security@factory.ai](mailto:security@factory.ai) for details. Email the security team at security@factory.ai. Compliance documents and certifications. # Identity & Access Manage who can run Droid with SSO, Directory Sync, roles, service accounts, and Factory API keys. Identity and access management controls **who can run Droid**, in which environments, and under what policies. This page covers the identity model and roles, Single Sign-On and SCIM directory sync, and service accounts for automation. --- ## Identity model Every Droid run is associated with several dimensions of identity: | Identity dimension | What it controls | | --- | --- | | User or machine identity | Human developers authenticate via SSO (SAML/OIDC), inheriting their directory groups and roles. Automation (CI/CD, scheduled jobs, scripts) runs under **service accounts** with their own API keys. | | Org / project / folder | The active repo and `.factory/` folders determine the **org, folder, and project context**. Policies at these levels decide which models, tools, and integrations Droid may use. | | Runtime environment | Whether Droid runs on a laptop, in a CI runner, or in a sandboxed VM is captured as environment attributes. Policies can treat these differently, such as allowing higher autonomy only in CI or sandboxes. | | Session metadata | Each Droid session records metadata such as session ID, CLI version, and git branch, which is available for audit and OTEL telemetry. | Identity and environment information feed two systems: **policy evaluation** (managed settings use identity and context to decide which configuration applies; see [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control)) and **telemetry and audit** (OTEL metrics and audit events carry attributes such as `user.id`, `organization.id`, and `session.id`; see [Compliance & Audit](/enterprise/compliance-audit-and-monitoring)). --- ## Roles Factory organizations have three roles. Role information flows from your IdP (via SSO claims or SCIM groups) into Factory. | Role | Capabilities | | --- | --- | | **Owner** | Full control of the organization, including billing, deletion, all settings, members, and API keys. | | **Manager** | Manage members, organization settings, API keys, and service accounts. Manage org-level `.factory` policy (models, command allow/deny lists, telemetry, autonomy ceilings). | | **User** | Standard member. Run Droid locally, in IDEs, in CI, or via team scripts, and customize personal preferences in `~/.factory/`, but cannot change any setting locked at the org, folder, or project level. | The managed settings engine enforces what each role can effectively change. See [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control) for the exact precedence rules. ## Domains, SSO, and SCIM Factory uses **WorkOS** for enterprise identity, supporting domain verification, SSO (SAML/OIDC), and Directory Sync (SCIM). Contact your Factory representative or [support@factory.ai](mailto:support@factory.ai) to start setup; you will receive a secure setup link to configure your identity provider. WorkOS handles the technical configuration. The setup wizard generates the SAML/OIDC parameters and SCIM endpoint for your specific provider, so you do not configure ACS URLs, certificates, or attribute mappings by hand. For provider-specific screenshots and steps, follow the WorkOS admin guide linked in your setup flow. ### Domain verification (required first) Domain verification proves your organization owns specific email domains and is a **prerequisite for SSO**. You add each domain, WorkOS issues a TXT record, and you add it to your DNS; verification typically completes within 1-24 hours depending on propagation. Once a domain is verified: {/* sweep-allow: term-bullets */} - **Automatic claiming** - existing and new users with that email domain join your organization. - **Identity governance** - you can require SSO for domain users and apply org security policies (MFA, session controls, IP restrictions) consistently. - **SSO readiness** - you can enable SSO for the verified domains. Verify your primary domain first, then add subsidiary, legacy, or alias domains. Each domain is verified separately but shares the same organizational policies. SSO cannot be enabled without verified domains. All users authenticating via SSO must have email addresses from verified domains. ### Single sign-on (SAML / OIDC) With SSO enabled, users sign in through your IdP instead of a password: 1. The user chooses "Sign in with SSO" and is redirected to your IdP. 2. They authenticate with corporate credentials. 3. The IdP returns a SAML assertion or OIDC token to Factory. 4. The user is authenticated and, if **Just-In-Time (JIT) provisioning** is enabled, created on first login with a profile populated from the IdP claims. WorkOS supports all major IdPs (Okta, Microsoft Entra ID, Google Workspace, OneLogin, Ping, JumpCloud) plus generic SAML 2.0 and OIDC. Organizations that do not use SSO can use passwordless magic-link sign-in, which is also useful for external collaborators. ### Directory sync (SCIM) SSO controls **how users authenticate**; SCIM controls **which users and groups exist** in Factory. With Directory Sync enabled through your IdP: - Adding a user to a synced directory group **provisions** them in Factory. - Updating directory attributes **syncs** to Factory. - Removing a user from the directory **deactivates** their access (soft delete - work history is preserved). - Group membership changes propagate automatically, so RBAC stays defined in your IdP. Sync only the groups that need Factory access (for example, `factory-*` groups), and keep the SCIM token secret - store it only in your IdP's application configuration. #### Group-to-role mapping Map your directory groups to Factory's three roles. WorkOS role slugs resolve as follows: `owner` maps to **Owner**; `admin` or `manager` maps to **Manager**; every other slug maps to **User**. | Example IdP group | Factory role | | --- | --- | | `factory-owners` | Owner | | `factory-managers` (or `factory-admins`) | Manager | | `factory-developers`, `factory-contractors`, any other group | User | Use descriptive, prefix-based group names (`factory-*`) so IdP configuration stays maintainable and auditable alongside your other enterprise apps. ### Data priority and conflict resolution When a user exists from multiple sources (SCIM, SSO JIT, manual invite, API): 1. **Directory Sync data wins** - SCIM attributes overwrite other sources. 2. **Email-based matching** - users are matched by email, case-insensitively (`John.Doe@company.com` equals `john.doe@company.com`). 3. **Custom data is preserved** - Factory-specific fields not present in your directory are retained. 4. **Soft deletes only** - removed users are deactivated, not deleted. When both SSO JIT and SCIM are enabled, choose one primary method to avoid duplicate-user conflicts: either disable JIT and require directory provisioning (directory-first), or use JIT for creation and the directory for updates (SSO-first). ### Troubleshooting Verify ACS / redirect URLs match what Factory provided and confirm certificates have not expired or rotated. Check group memberships and the group-to-role mapping, and ensure the intended groups are included in SAML assertions or ID tokens. Confirm SCIM is enabled in both Factory and your IdP, and check your IdP's SCIM logs for invalid token, URL, or schema errors. Keep a small pilot group for initial rollout, and manage all role changes and access reviews in your IdP to reuse existing governance. ## Service accounts Service accounts are non-human identities that automate work on behalf of your organization: CI/CD pipelines, scripts, Slack automations, or Droid Computers. Use one when a workload should belong to your organization instead of a specific teammate. A service account authenticates with a Factory API key and runs as its own principal, so sessions, computers, billing, and audit events are attributed to it rather than a human user. | Step | What to configure | | --- | --- | | Create a service account | Add a stable identity from the **API Keys** section of organization Settings. Give it a clear name and description. Only **Owner** and **Manager** roles can create or edit service accounts. | | Generate API keys | Create keys for scripts, CI jobs, and other automated workloads. A key value is shown **only once** at creation - copy it and store it securely. | | Attach Droid Computers | Create Droid Computers owned by the service account so long-running work continues under the same identity instead of a teammate's account. | | Configure Git access | Use your organization's **GitHub App** installation for GitHub. For GitLab, add a service-account **GitLab token (PAT)** in Settings. For a self-managed server, see [Self-Managed Source Control](/enterprise/self-managed-scm). | A service account must be **active** to authenticate. Marking one inactive or deleting it causes new requests with its keys to fail, so move or restart dependent workloads first. Names are immutable identifiers - create a new account if you need a different name. ### Service account key hygiene | Practice | Why it matters | | --- | --- | | Use short-lived keys when possible | Set expiration dates for keys used in temporary automation to limit blast radius if a key is exposed. | | Rotate keys regularly | Create a new key, update the workload, then revoke the old key. | | Delete when finished | Deleting a service account revokes its access and cleans up owned Factory resources where possible. | --- ## Devices, environments, and workspace trust Because Droid is a CLI, it can run on developer laptops, remote dev servers, CI runners, and hardened VMs or devcontainers. Enterprises typically combine Droid with endpoint management and environment-aware policy: ### Endpoint and workspace controls | Control | Recommended use | | --- | --- | | Endpoint and MDM controls | Use Jamf, Intune, or other MDM solutions to control where Droid binaries can be installed, which users can run them, and which configuration files they can read. Common patterns: run only under managed user accounts, restrict configuration directories to corporate-managed volumes, and enforce OS-level disk encryption and screen lock. | | Workspace trust | Treat Droid as trusted only in known repositories. Pin Droid to specific paths or repos, require elevated approval or sandboxed environments for untrusted code, and use project-level `.factory/` folders to mark which repos are "Droid-ready." | | Environment-aware policies | The same developer may run Droid on a laptop, in CI, or in an isolated container. Policies can allow higher autonomy or more powerful tools **only** inside devcontainers or CI runners, restrict network access on laptops, and tag OTEL telemetry with environment attributes for environment-specific alerting. | Define which roles can change which policy levels in the settings hierarchy. Map identity and access events to audit trails and compliance workflows. # Organization Model Understand how enterprise accounts, organization hierarchy, identity, policy, billing, and data residency fit together in Factory. Factory's organization model lets an enterprise expand Droid across many teams without turning every team into a separate commercial or identity tenant. One enterprise root owns the contract, identity connection, billing relationship, and top-level administration. Organizations under that root provide operating boundaries where teams can manage membership, local admins, settings, credentials, usage limits, and regional posture. Use this page to design the operating model before creating organizations under the enterprise root. For the step-by-step setup guide, see [Organizations](/enterprise/structuring-and-managing-organizations). For the management surface, see [Admin Console](/enterprise/enterprise-admin-console). ## Model at a glance The enterprise root is the top of the Factory hierarchy. It owns the enterprise relationship and can govern the organizations below it. Each organization is an operating boundary inside that relationship. ```mermaid %%{init: {"theme": "base", "themeVariables": {"fontFamily": "Geist Mono, monospace", "fontSize": "13px", "primaryColor": "#161413", "primaryBorderColor": "#342F2D", "primaryTextColor": "#FAFAFA", "lineColor": "#4D4947", "textColor": "#D6D3D2"}}}%% flowchart TD Root["Enterprise root
contract, SSO, billing, top-level governance"] Platform["Platform services
shared administrators and defaults"] Delivery["Product delivery
repositories and integrations"] Regulated["Regulated workloads
stricter controls and regional deployment"] Sandbox["Innovation sandbox
broader experimentation"] Client["Client delivery
separate credentials and service accounts"] Root --> Platform Root --> Delivery Root --> Client Delivery --> Regulated Delivery --> Sandbox ``` ## Core concepts | Concept | What it means | | --- | --- | | Enterprise root | The top-level organization for the enterprise. It owns the enterprise contract, SSO connection, optional Directory Sync connection, billing account, and enterprise administration. | | Organization | An operating boundary inside the enterprise hierarchy. It can have its own members, roles, local admins, controls, usage limits, service accounts, and regional settings where root policy allows. | | Organization path | A slash-delimited label such as `Product Delivery/Regulated Workloads`. The path is human-readable. Factory also keeps a stable internal ID so renaming a team does not erase its identity. | | Active organization | The org context a user chooses when working in the app, desktop app, or CLI. New sessions are tied to this selected org for product context and membership behavior. | | Root attribution | Usage, billing, and commercial attribution roll up to the enterprise root even when work happens in another organization. | ## What rolls up to the root The root keeps the enterprise relationship centralized: - one enterprise contract and subscription; - one billing account and invoice; - one SSO connection; - one optional Directory Sync connection; - enterprise-level visibility and administration; - root-managed controls where the enterprise chooses to enforce them. Organizations under the enterprise root should not be configured as separate SSO tenants or separate billing accounts. They are hierarchy and governance boundaries inside the same enterprise deployment. ## What can differ by organization Organizations provide separation where teams need different operating rules. | Area | Can differ by organization | | --- | --- | | Membership and roles | Users can be assigned to one or more organizations with Owner, Manager, or User roles. | | Local administration | Organization owners and managers can administer their org where root controls allow. | | Enterprise Controls | Model policy, autonomy ceilings, MCP policy, command rules, sandboxing, retention, and related settings can vary by org unless restricted by root management. | | Usage limits | Global and per-user limits can be set locally where allowed. Root admins can copy or apply root limits to selected organizations. | | Region and data residency | An organization can be assigned to a regional deployment at creation. The default is Global. | | Service accounts and credentials | Automation identities should be scoped to the organization that owns the workflow. | | Integrations | Integrations are configured at the organization level, so teams can isolate repository, service account, and workflow access. | ## Design patterns Create a separate organization when a group needs a real boundary, not only a label. The strongest signals are different admins, different membership, different credentials, different data residency, different model policy, or different spend controls. ### Regional model Use regional organizations when residency, local administration, or deployment routing differs by geography. ```text Acme Corp ├── Global Platform ├── Europe │ └── Regulated Infrastructure └── North America ``` Declare the region when creating the organization. Mixing regions in a hierarchy is supported when the enterprise intentionally wants different residency boundaries for different teams. ### Regulated workload model Use a nested organization for high-sensitivity work that needs stricter model access, lower autonomy, tighter command policy, or BYOK / customer-hosted inference. ```text Acme Corp └── Product Delivery ├── Standard Workloads └── Regulated Workloads ``` This lets the regulated group inherit enterprise identity and billing while keeping stricter local controls. ### Project or client delivery model Consulting, agency, and client delivery teams often need a separate boundary per project or client environment. ```text Acme Consulting └── Professional Services ├── Project A │ ├── Client Environment │ └── Restricted Work └── Project B ``` Use this pattern when credentials, repositories, or client-specific service accounts should not be shared across projects. ### Department-only model, discouraged by default Avoid mirroring the HR department chart as the default hierarchy. A department branch is useful only when the department itself needs separate administration, membership, credentials, data residency, model policy, or spend controls. ```text Acme Corp ├── IT │ ├── Engineering │ └── Operations └── Finance ``` Use this model when those departments share one enterprise contract and SSO connection but truly need different owners, members, service accounts, and policy defaults. If the only difference is reporting, keep the organization structure simpler and use usage exports or analytics instead. ## Naming and hierarchy guidance Keep the hierarchy understandable for Finance, IT, security, and team admins: - prefer two or three levels unless deeper structure is required; - use stable names that match how the enterprise already refers to teams; - avoid creating organizations only for reporting if membership and policy are identical; - keep sibling names unique and easy to search; - choose paths that make usage exports and access reviews readable. ## Planning worksheet | Question | Decision to make | | --- | --- | | Who owns the enterprise root? | Identify the enterprise owners who can manage root policy, billing, and identity. | | Which groups need boundaries? | List teams that need distinct membership, admins, controls, credentials, regions, or limits. | | Which controls are root-managed? | Decide which Enterprise Controls, usage limits, and membership controls should be read-only for local organization admins. | | How will membership be managed? | Choose Directory Sync group mapping, UI-based assignment, or a migration period using both. | | Which regions are required? | Decide whether each organization uses Global or a regional deployment. | | Where should automation live? | Scope service accounts, integrations, CI, and other machine workflows to the owning organization. | Create organizations, configure membership, and manage user access. Manage directory, organizations, billing, and identity from the enterprise surface. Configure model, tool, autonomy, retention, and safety policies. Configure SSO, Directory Sync, roles, service accounts, and API keys. # Structuring and managing organizations Create organizations, assign users, configure Directory Sync mappings, manage usage limits, and plan migration under one enterprise root. Factory organizations let an enterprise divide one account into teams, business units, regions, client groups, environments, or departments that need true operating boundaries. Each organization can have its own members, roles, local administration, usage limits, service accounts, and controls while still rolling up to one enterprise contract, one invoice, one identity provider connection, and one enterprise root. Use this page when you are ready to create, configure, and operate the hierarchy in the Admin Console. For the conceptual model and design patterns, see [Organization Model](/enterprise/organization-model). ## Overview All usage, billing, and commercial attribution rolls up to the enterprise root. Organizations provide hierarchy, membership, role assignment, local administration, policy segmentation, and reporting structure, but the root remains the commercial owner. Use organizations when you want to: - organize teams or operating boundaries under one enterprise root; - give different teams different admins or settings; - keep one enterprise invoice attributed to the root organization; - manage users through one enterprise SSO connection, with Directory Sync or UI-based membership assignment; - prepare for users who work across multiple teams; - apply different region, residency, usage, or policy boundaries to different groups. ## Prerequisites Before creating organizations, confirm: - one root Factory organization is identified as the enterprise root; - SSO is configured for the enterprise; - the membership management approach is selected: Directory Sync for IdP-managed membership, UI-based assignment for manually managed membership, or both during migration; - if using Directory Sync, Admin Console access is available and IdP groups exist or can be created for each organization role; - an enterprise subscription or commit is in place so usage can roll up to one enterprise invoice; - the planned hierarchy and initial owners are known; - the region and data residency strategy is known for the root and each organization. Directory Sync is recommended for scaled IdP-managed membership, but it is not required. Customers can also assign organization membership in the UI. Invite new users through the enterprise root first, then assign them to the appropriate organization. ## Core concepts ### Enterprise root The enterprise root is the top-level Factory organization. It owns the enterprise contract, SSO connection, optional Directory Sync connection, billing account, top-level membership, and top-level administration. For existing enterprise customers, the existing organization usually becomes the enterprise root. This lets the customer migrate into an organization hierarchy without replacing the existing org or disrupting existing users, sessions, billing, or identity configuration. ### Organization An organization is an operating boundary inside the enterprise hierarchy. Organizations can be nested. ```text title="organization-hierarchy.txt" Acme Corp ├── Platform Services ├── Product Delivery │ └── Regulated Workloads └── Client Delivery ``` Factory can support many hierarchy shapes, but the hierarchy should be intentionally designed for clarity, administration, and long-term maintainability. ### Organization path Each organization has a slash-delimited path, such as: ```text Product Delivery/Regulated Workloads ``` The path is the human-readable hierarchy and membership label. The organization also has a stable internal ID, so renaming a team does not lose its identity. Billing and commercial attribution still roll up to the root. ### Org membership and primary org selection Users can be assigned to multiple organizations. When a user has access to more than one org, Factory provides an org selector so the user can choose which org context to use. When a user enters Factory for the first time, or has not selected an org yet, Factory chooses the starting org in this order: 1. The root org wins when the user has root access. 2. If the user does not have root access, assigned organizations are sorted by hierarchy and name, then the first is used. 3. If the user chooses a different accessible org in the selector, that selected org becomes the active org context. 4. For multiple roles on the same org, the highest role wins: Owner, then Manager, then User. Sessions are tied to the selected or primary org for their lifetime for product context and membership behavior. Local sessions on a user's machine can still be accessed from any org context on that machine. Usage and billing attribution still roll up to the root. ### Switching organizations Users can switch between organizations where they have access. - In the web or desktop app, open settings, go to **General**, then select an organization under **Organization**. - In the CLI, run `/settings`, open **Preferences**, then choose **Active Organization**. ## What is shared across the enterprise Organizations can be independent where teams need separation, but they intentionally share enterprise-level infrastructure. | Area | Enterprise-wide behavior | | --- | --- | | Billing | One billing account and invoice for all orgs. | | Contract and subscription | Managed on the enterprise root. | | Attribution | All usage, billing, and commercial attribution goes to the root. | | SSO | One WorkOS organization and SSO connection. | | Directory Sync | Optional. When used, one directory connection serves the enterprise. | | Identity provider setup | One IdP integration with groups mapped to Factory roles. | | Root-managed controls | Root controls can apply enterprise controls, usage limits, and membership restrictions to selected organizations. | Organizations should not be configured as separate SSO tenants or separate billing accounts. ## What can differ by organization | Area | Organization behavior | | --- | --- | | Members | Users can belong to multiple organizations. One primary assignment is recommended when possible. | | Roles | User, Manager, and Owner roles apply per assigned organization. | | Usage | Product context can be tied to the selected org, while commercial attribution rolls up to the root. | | Usage limits | Local usage-limit settings can exist where allowed. Root global limits can be copied during creation and later applied to selected organizations. | | Enterprise Controls | Settings can be customized locally unless the enterprise root restricts organization changes. | | Membership controls | Membership can be managed locally unless the enterprise root restricts local membership changes. | | Region / data residency | An organization region can be declared at creation. The default is Global. | | Service accounts | Automation identities are scoped to the organization where they operate. | | API keys | Legacy API keys remain org-scoped while they exist. Service accounts are the preferred path. | | Admin responsibilities | Organization owners and managers can administer their organization according to their role and any root-managed restrictions. | ## Roles and permissions Factory uses three role levels at each level of the hierarchy. | Role | Typical permissions | | --- | --- | | Owner | Manage settings, members, billing-visible views, and high-privilege admin actions. | | Manager | Manage day-to-day team settings and members where allowed. | | User | Use Factory products in the assigned organization. | An enterprise owner can manage the enterprise root and create or administer organizations according to enterprise policy. An organization owner or manager can administer the organizations where they have that role, except for settings, policies, membership controls, and services where the root has declared jurisdiction. If multiple roles apply to the same org, the most privileged role wins: Owner, then Manager, then User. ## Identity, membership, and directory sync Organizations use one enterprise SSO connection configured from the Admin Console. Membership can be assigned in two ways: 1. Directory Sync group mapping, recommended for scaled IdP-managed membership. 2. UI-based membership assignment, useful for manual setup, pilots, and customers not ready to manage organization membership through Directory Sync. New users must be invited through the enterprise root before they can be assigned to an organization. ### How directory sync works with organizations Directory Sync manages membership and role assignment from the Admin Console. It does not replace organization creation. | Question | Answer | | --- | --- | | Do I need to enable a separate "set up organization via Directory Sync" configuration? | No. Use the Admin Console to enable SSO and Directory Sync for the enterprise, then map IdP groups to Factory role targets for the root and any organizations. | | Can Directory Sync create organizations? | No. Create organizations in the Admin Console first. Directory Sync then assigns users to the role targets Factory exposes for those organizations. | | How does Directory Sync pick up an organization created in the Admin Console? | After the organization exists, Factory creates role targets for that organization. Map IdP groups to those role targets from the Admin Console identity flow, then run or wait for Directory Sync reconciliation. | | What if a role target is not visible yet? | Reopen the Admin Console identity setup surface or contact Factory support. Do not create a separate SSO tenant for the organization. | ### Directory sync setup When using Directory Sync: 1. Create the organization in **Admin Console → Organizations**. 2. Create or identify IdP groups for the organization's Owner, Manager, and User roles. 3. Open **Admin Console → Security & Identity**, then map each IdP group to the corresponding Factory role target. 4. Add users to the appropriate IdP groups. 5. Run or wait for Directory Sync reconciliation. 6. Confirm users can access the expected org contexts. Example mappings: | IdP group | Factory access | | --- | --- | | `Acme Factory Owners` | Owner of `Acme Corp` | | `Product Delivery Factory Users` | User of `Product Delivery` | | `Regulated Workloads Factory Managers` | Manager of `Product Delivery/Regulated Workloads` | | `Client Delivery Factory Users` | User of `Client Delivery` | During migration, users who need both root and organization access must be in both a root-mapped group and an organization-mapped group. ### UI-based membership setup When using UI-based membership: 1. Invite the user into the enterprise root if they are not already a member. 2. Add or update the user's organization role from the membership UI. 3. Confirm the user can select the intended org context. 4. Keep high-privilege Owner assignments small and reviewed. Use UI-based membership for pilots, manual administration, or customers that do not want Directory Sync as the sole membership path. ## Creating an organization Enterprise admins can create organizations in the Admin Console after Factory enables organization creation for the enterprise. ### UI creation flow 1. Open the Admin Console as an enterprise account owner. 2. Go to **Organizations**. 3. Choose the parent org for the new organization. 4. Enter the organization name. 5. Choose the region, or use the Global default. 6. Decide whether to copy root usage limits, Enterprise Controls, and membership controls. 7. Decide whether organization admins can make local changes to membership, usage limits, and Enterprise Controls. 8. Create the organization. The user who creates the organization is automatically added as an owner of that organization. The creator should confirm they can access the new organization and that any other intended Owner or Manager assignments are in place. ## Usage and billing The enterprise receives one invoice. All usage and billing attribution is attributed to the enterprise root, even when users work in organization contexts. Example usage summary: ```text Factory Standard Credits acme-corp: 48,750 acme-corp/it: 82,140 acme-corp/it/engineering: 36,425 acme-corp/finance: 19,880 Subtotal: 187,195 ``` Usage within `acme-corp` does not include usage within `acme-corp/it`. Each row represents usage in that org context, while the subtotal is the enterprise total. Enterprise admins can review these numbers across the root organization and its sub-organizations in the Admin Console: open **Organizations** and select **View usage** to group consumption by organization, filter to a single organization, and export the view as CSV. See [Usage analytics for sub-organizations](/enterprise/enterprise-admin-console#usage-analytics-for-sub-organizations). Recommended guidance: - choose the correct organization when starting new work; - expect invoice and commercial attribution to remain at the root; - keep long-running work in its original org unless there is a clear reason to recreate it; - use paths that Finance, IT, security, and team admins recognize; - review the first invoice or usage export after enabling organizations. ### Usage limits Usage limits control how much usage an org or user can consume within the enterprise subscription. | Level | Meaning | | --- | --- | | Global level | A default usage limit policy for the organization. | | Per-user level | A user-specific usage limit that can differ from the default. | During organization creation, the enterprise admin can copy the root global usage limit into the new organization. The enterprise admin can also restrict the organization from updating its own usage limits. After creation, when the enterprise root changes its global usage limit, the enterprise admin can choose which organizations should receive that update. Root changes are not automatically synced to every organization. Before setting or tightening a limit, review the organization's actual consumption with **View usage** in the Admin Console so the limit reflects real usage patterns. ## Region deployments and data residency Organizations should follow the enterprise's data residency and deployment strategy. At creation, the admin can declare the organization's region. If no region is declared, the organization defaults to Global. Region planning matters for: - where session data is stored; - which deployment serves the org; - which integrations and network policies apply; - how support, debugging, and audit workflows are routed; - whether users can move work between orgs without residency concerns. | Choice | Behavior | | --- | --- | | Default to Global | The organization uses the Global region. | | Declare a specific region | The organization uses the selected deployment region for its own data and requests. | | Mix regions within a hierarchy | Supported as an intentional setup choice when different teams have different residency needs. | If a customer wants independent data residency for a specific organization, declare that region during creation and document the intended data boundary. ## Service accounts, API keys, and automation access Automation access should be scoped to the organization where the automation operates. ### Service accounts Service accounts are the preferred model for non-human automation. They should be created and assigned according to the team or workflow they represent. | Service account | Suggested scope | | --- | --- | | `product-delivery-ci` | `Product Delivery` | | `regulated-reporting` | `Product Delivery/Regulated Workloads` | | `platform-automation` | root org or `Platform Services` | Guidance: - create service accounts in the organization that owns the workflow; - avoid using a root-level service account for team-specific automation unless the workflow truly operates enterprise-wide; - review service account ownership when moving a workflow between organizations; - include service accounts in access reviews alongside human owners and managers. ### API keys API keys are legacy and expected to be replaced by service accounts. While API keys still exist, treat them as org-scoped credentials. Guidance: - create API keys only in the organization that owns the integration; - do not reuse a root API key for organization-specific workflows; - prefer service accounts for new automation; - track any remaining API keys during migration so they can be replaced later; - rotate or remove API keys when a team, integration, or organization changes ownership. ### Automation ownership worksheet | Workflow | Owning organization | Human owner | Credential type | Notes | | --- | --- | --- | --- | --- | | CI session creation | `Product Delivery` | | Service account | | | Regulated reporting | `Product Delivery/Regulated Workloads` | | Service account | | | Legacy integration | | | API key | Replace with service account. | ## Enterprise Controls and membership controls Each organization can have local controls such as managed settings, security policy, usage limits, and membership administration where root restrictions allow it. During organization creation, the enterprise admin can copy root Enterprise Controls into the new organization. The enterprise admin can also restrict the organization from making its own Enterprise Control changes, usage-limit changes, or membership-control changes. After creation, when the enterprise root changes Enterprise Controls, global usage limits, or membership controls, the enterprise admin can choose which organizations should receive that update. Enterprise Controls include: - model allowlists and blocklists; - usage limits; - membership controls; - custom model access and user model policies; - max autonomy level and session defaults; - command allowlists and denylists; - MCP, network, and sandbox policy; - cloud session sync and Droid Shield; - member visibility; - API key creation; - managed computer availability; - session retention. Organization admins can customize local settings and membership behavior when the enterprise admin has not restricted local changes. When local changes are restricted, root-managed fields should be shown as read-only with the root identified as the managing org. ### Copying root settings Copying root settings is useful when: - the organization should begin with the same model policy; - command controls should match the root baseline; - membership controls should match the root operating model; - integration defaults should be the same; - the team wants a baseline before local customization. Starting from defaults is useful when: - the organization is a pilot or sandbox; - the team has a different risk profile; - the root has legacy settings that should not carry forward; - the customer wants to rebuild policy intentionally. Settings copied at creation are a starting point. If the enterprise root later changes global usage limits, Enterprise Controls, or membership controls, the enterprise admin chooses which organizations should receive that update. Root changes are not automatically synced to every organization. Enterprise admins choose which organizations should receive later root updates. ### Enterprise controls review checklist Which controls must apply enterprise-wide? Which controls should differ by risk profile, region, or workflow? Which usage limits should be copied from the root? Which organizations should be restricted from changing their own usage limits? Which membership controls should be managed by the root? Which organizations can manage custom models or BYOK? Which organizations can use MCP servers? Which organizations can create service accounts or legacy API keys? ## Recommended rollout ### Step 1: design the hierarchy Start with the boundaries that must differ, not the department chart. Create an organization when a group needs distinct membership, administration, policy, region, credentials, integrations, or usage limits. | Organization path | Owners | Managers | Users source | Notes | | --- | --- | --- | --- | --- | | `Platform Services` | Platform leadership | Platform leads | Directory Sync group | Shared services, defaults, and administrative ownership. | | `Product Delivery` | Engineering leadership | Team leads | Directory Sync group or UI assignment | Standard product work with shared repositories and integrations. | | `Product Delivery/Regulated Workloads` | Security or compliance owner | Regulated workload leads | Directory Sync group | Stricter controls, regional deployment, BYOK, or customer-hosted inference. | | `Client Delivery` | Services leadership | Engagement leads | Directory Sync group or UI assignment | Separate client credentials, repositories, service accounts, or automations. | Use department-based organizations only after this boundary review. A department should become an organization when it needs different owners, members, credentials, residency, policy, or spend controls, not only because it appears as a department in the HR system. ### Step 2: confirm enterprise prerequisites Confirm the root organization, SSO, membership approach, billing expectation, region plan, service accounts, and legacy API key ownership before enabling broad access. ### Step 3: create organizations Create each organization under its intended parent. Organization names must be unique among siblings. For each organization, decide whether it should copy root settings, allow local changes, declare a region, and own any service accounts, API keys, integrations, or automations. ### Step 4: configure roles and membership Configure User, Manager, and Owner access through Admin Console Directory Sync mappings, UI-based assignment, or both during migration. ### Step 5: pilot with a small group Before broad rollout, assign users to two or three organizations and confirm: - users can choose the correct org when starting work; - usage and billing attribution remain at the root; - selected org context appears as expected in product surfaces; - role gates, membership controls, and Enterprise Controls behave as expected; - region and residency expectations are met; - service accounts and legacy API keys are scoped to the correct org. ### Step 6: expand rollout After pilot validation, create the remaining organizations, map remaining IdP groups or assign users in the UI, re-scope credentials where needed, notify users how to select the right org, and monitor root attribution during the first billing period. ## Migration behavior Enabling organizations does not require replacing the existing organization. The existing org becomes the enterprise root. Migration behavior: - existing users can continue working in the root org; - users who should move into organizations can be assigned to both the root org and their target organization during migration; - with Directory Sync, the user must be in both a root-mapped group and an organization-mapped group; - with UI-based membership, the user must first be invited to or retained in the enterprise root, then assigned to the target organization; - existing sessions keep their existing org context; - new sessions can use organization context for membership, administration, and policy behavior; - all usage and billing attribution remains at the root. Recommended migration strategy: 1. Confirm the existing org is the enterprise root. 2. Create the target hierarchy under the root. 3. Copy root global usage limits, Enterprise Controls, and membership controls where they are a useful starting point. 4. Decide which organizations should be restricted from changing their own limits, controls, and membership settings. 5. Configure membership through Admin Console Directory Sync mappings or UI-based assignment. 6. Add migrating users to both the root org and their target organization for backward compatibility. 7. Ask users to switch into the appropriate organization when they are ready to create new work there. 8. Pilot a small group and confirm org switching, root attribution, membership controls, and Enterprise Controls. 9. Re-scope API keys and service accounts to the intended owning org where needed. 10. Expand membership gradually and monitor root-attributed usage plus selected org context during the first billing period. 11. If desired, remove users from the enterprise root after they have fully moved into their intended organizations. ## FAQ Yes. Users can belong to multiple organizations. To limit confusion and administrative overhead, assign each user to one primary organization when possible. No. Directory Sync is recommended for scaled IdP-managed membership, but customers can also assign organization membership in the UI. No. New users must be invited through the enterprise root. After the user exists in the enterprise root, they can be assigned to the appropriate organization. Yes. The root org is still a usable org. Usage and billing attribution goes to the root even when users work in another organization context. No. The enterprise receives one invoice attributed to the root organization. No. Organizations use one enterprise SSO connection. Yes. An organization's region can be declared at creation time. If no region is declared, it defaults to Global. Service accounts should live where the workflow operates. Use a root-level service account only for enterprise-wide workflows. Use an organization service account for team-specific automation, integrations, and CI workflows. Existing sessions continue to use their original org context. New sessions should be created in the intended organization context. Usage and billing attribution still goes to the root. Yes, when root management is enabled for a control area. Organization admins can make local changes only where the enterprise root allows them. No. Copying settings during organization creation does not automatically sync future root changes to every organization. Later root changes affect only the organizations the enterprise admin selects for that update. Design the enterprise root, organization hierarchy, and operating boundaries. Use the Admin Console to manage directory, organizations, billing, and identity. Configure SSO, Directory Sync, roles, service accounts, and API keys. Manage settings, model access, safety policy, and usage controls. # Admin Console Manage enterprise-wide directory, organization hierarchy, billing, SSO, domain verification, and Directory Sync from one Factory surface. The Enterprise Admin Console is the central command center for managing Factory at enterprise scale. It gives enterprise account owners a dedicated platform surface to manage identity, directory, organization structure, billing, and enterprise-level governance across the customer's Factory deployment. The Admin Console is intentionally separate from regular team settings and is not scoped to the currently selected organization. Team owners and managers can administer their own organizations, while enterprise owners retain control over identity, billing, hierarchy, and cross-organization membership visibility. ## Who can access the admin console The Admin Console is reserved for enterprise account owners who can administer the enterprise account rather than only a local organization. | User type | Access | | --- | --- | | Enterprise account Owner | Can access the Admin Console. | | Organization Owner | Can access local organization settings, but not the enterprise Admin Console unless they are also a root Owner. | | Manager | Can access team settings where their role allows it, but not enterprise-only Admin Console pages. | | User | Does not see team settings or Admin Console surfaces. | Direct navigation to restricted Admin Console pages shows restricted access. This gives customers a clear governance boundary: local ownership does not automatically grant enterprise administration. ## Admin console sections The Admin Console includes four enterprise-level sections. | Section | What it manages | | --- | --- | | Directory | Cross-organization membership visibility and membership management. | | Organizations | Organization creation, hierarchy, organization-level controls, usage limits, and sub-organization usage analytics. | | Billing | Enterprise account billing and commercial administration. | | Security & Identity | SSO, domain verification, and Directory Sync setup. | ## Directory Directory shows members across the enterprise account and its organizations. Admins can search the directory, filter by organization, review member status, see whether memberships are UI-managed or Directory Sync-managed, and manage memberships when allowed. Use Directory to answer: - who has access to Factory across the enterprise; - which organizations a user belongs to; - which users are managed by Directory Sync versus manual assignment; - which memberships need review during rollout or migration. When Directory Sync controls a membership, make the change in the identity provider. When a membership is UI-managed and the admin has permission, the Admin Console can update it directly. ## Organizations Organizations is where enterprise admins create and manage the organization hierarchy. Admins can search organizations, create a new organization under a selected parent, choose Global or a regional deployment, optionally copy enterprise-level controls, optionally copy the enterprise global usage limit, view directory members, manage Enterprise Controls, manage global usage limits for a specific organization, and review credit consumption across the root organization and its sub-organizations with **View usage**. Use Organizations for: - teams, business units, regions, subsidiaries, client groups, or departments with distinct operating boundaries; - per-organization policy and usage boundaries; - region and data residency selection at organization creation; - service account and automation ownership review; - controlled expansion from a pilot team to the broader enterprise. For detailed setup steps, see [Organizations](/enterprise/structuring-and-managing-organizations). ## Billing Billing gives enterprise account owners the commercial administration surface for the enterprise account. Billing stays with the enterprise root, reinforcing one commercial relationship even as usage expands across many organizations. Use Billing to keep finance aligned with: - the enterprise account and subscription; - centralized invoice ownership; - usage and commercial attribution that rolls up to the root; - expansion across teams without disconnected billing accounts. ## Security and identity Security & Identity gives enterprise admins direct access to identity configuration from the Admin Console. Admins can configure Single Sign-On, domain verification, Directory Sync, and IdP group mappings through the identity portal. Security & Identity supports: - SAML or OIDC-based SSO; - domain verification; - SCIM-based Directory Sync for automated provisioning; - IdP group mappings to Factory organization role targets; - one enterprise identity workflow for the root and its organizations. For identity model details, see [Identity & Access](/enterprise/identity-and-access). ## Enterprise Controls for organizations Enterprise admins can open Enterprise Controls for a specific organization from the Organizations page. This lets them tune policies per team without changing every other organization. Use per-organization Enterprise Controls to: - segment policy by team, region, or business unit; - give regulated teams stricter controls; - avoid one-size-fits-all governance across the enterprise; - keep centralized oversight while allowing local policy variation. Enterprise Controls include model policies, custom model access, autonomy ceilings, command allowlists and denylists, [connector availability](/harness/connectors#enterprise-controls), MCP policy, sandboxing, cloud session sync, Droid Shield, API key creation, managed computer availability, and session retention. ## Usage limits for organizations Enterprise admins can set or remove a global monthly credit limit per user for a specific organization. This gives admins a practical way to manage adoption and spend boundaries at the team level. Use organization usage limits to: - manage cost exposure during expansion; - give teams clear usage boundaries; - support controlled pilots and staged rollout phases; - separate global organization limits from individual user limits. Root admins can copy root usage limits to a new organization during creation. Later root changes apply only to the organizations the enterprise admin selects. ## Usage analytics for sub-organizations Enterprise admins can review actual credit consumption across the root organization and its sub-organizations from the Organizations page. Select **View usage** in the page header to open the sub-organization usage view. The usage view supports: - grouping consumption by organization, model, user, user type, source, or product; - filtering by organization or user type, with Human and Service Account options; - custom date ranges, with all data displayed in UTC; - CSV export of the current view for finance and capacity reviews. When a grouping has more than 50 values, the chart and the CSV export show the top 50 and roll the remainder into an Other series. Use sub-organization usage analytics to: - confirm that usage attribution matches the hierarchy you designed; - spot high-consumption organizations before setting or adjusting usage limits; - give finance a per-team consumption breakdown without manual exports for each sub-organization. The usage view is read-only and available to the same enterprise admins who can open the Organizations page. ## Recommended workflow 1. Confirm enterprise root ownership. 2. Design the hierarchy in [Organization Model](/enterprise/organization-model). 3. Create organizations in **Organizations**. 4. Configure SSO and domain verification in **Security & Identity**. 5. Enable Directory Sync if membership should be managed through IdP groups. 6. Map IdP groups to Factory organization role targets in the identity portal. 7. Review membership in **Directory**. 8. Configure Enterprise Controls and usage limits for each organization. 9. Review consumption across the root organization and its sub-organizations in **Organizations** with **View usage**. 10. Confirm billing and usage attribution in **Billing**. ## FAQ Team admins can manage local team settings where their role and root policy allow it. Enterprise administration is reserved for enterprise root Owners. Yes. Security & Identity supports SSO, domain verification, and SCIM-based Directory Sync. Directory also indicates when memberships are managed by Directory Sync. Yes. Enterprise admins can manage Enterprise Controls for individual organizations, which supports team-specific policy requirements. Yes. Enterprise admins can set global monthly credit limits per user for specific organizations, and can review actual credit consumption across the root organization and its sub-organizations from the Organizations page with **View usage**. Understand the enterprise root, organization hierarchy, and governance boundaries. Create organizations, configure users, and manage Directory Sync mappings. Configure SSO, domain verification, Directory Sync, and roles. Manage hard controls, session defaults, model policies, and safety settings. # Self-Managed Source Control Connect Factory to a self-hosted GitHub Enterprise Server or GitLab instance with one organization-level OAuth app. Factory connects to a self-managed GitHub Enterprise Server or GitLab instance through an OAuth application that you register on your own server. An admin registers the app once, enters its credentials in Factory, and authorizes it; from then on your organization's repositories are available in sessions like any other repository. Connecting a self-managed GitHub Enterprise Server or GitLab instance requires the Enterprise plan. ## Prerequisites | Requirement | Detail | | --- | --- | | Factory role | **Owner** or **Manager**. Configuring an integration is a privileged action. | | Server access | Permission to register an OAuth application (GitHub Enterprise Server) or an instance-wide application (GitLab). | | Reachability | The instance must resolve to a public IP address. Factory validates and pins the resolved address to prevent server-side request forgery, so private-network-only hosts cannot be connected. If your instance sits behind a proxy or firewall, ask [support@factory.ai](mailto:support@factory.ai) for the Factory egress addresses to allowlist. | Both integrations authorize **once per organization**. A single admin authorization is stored at the org level and used for every Factory operation against that server, so individual members never authorize the server themselves. Use a dedicated service account as the connecting admin rather than a personal employee account: it keeps org-wide access off one person's account lifecycle and lets you bound which repositories are exposed. ## Connect GitHub Enterprise Server Factory requests the `repo` and `read:org` OAuth scopes and imports only **organization-level** repositories. Personal repositories owned by the authorizing account are never imported. In your GitHub Enterprise Server instance, go to **Settings → Developer settings → OAuth Apps → New OAuth App**. Register it under an organization you own rather than a personal account. - **Application name:** `Factory Integration` - **Homepage URL:** `https://app.factory.ai` - **Authorization callback URL:** `https://api.factory.ai/api/integrations/redirect/github-es/callback` If you run Factory on a dedicated control plane, substitute your own Factory API domain in the callback URL. Register the app, then copy the **Client ID** and generate a **Client Secret**. Store the secret in your secret manager before leaving the page; GitHub shows it once. Open Settings → Integrations, choose **GitHub Enterprise**, and enter your server domain (for example `https://github.your-company.com`), Client ID, and Client Secret. Factory normalizes and validates the domain on save. Start the connection. Factory redirects you to your instance to approve the OAuth app, then imports the repositories the authorizing account can reach. Enable the ones you want available in Factory from the repository list. ## Connect GitLab Self-Hosted GitLab authorization runs as a dedicated instance user rather than the admin's own account, so repository access does not follow an individual's account lifecycle. In your GitLab instance, go to **Admin Area → Applications → Add new application**. - **Name:** `Factory Integration` - **Redirect URI:** `https://api.factory.ai/api/integrations/redirect/gitlab-sh/callback` - **Trusted:** yes - **Confidential:** yes - **Scopes:** `api`, `read_api`, `read_user`, `read_repository`, `write_repository`, `read_observability`, `write_observability` Save the application and note the **Application ID** and **Secret**. Factory requests a conservative subset of these scopes at authorization time (`api`, `read_user`, `read_repository`, `write_repository`) so the flow works on older GitLab versions. Go to **Admin Area → Users → New User** and create a user named `Factory Droid` with the username `factory-droid` and any email address you control. Then sign in as, or impersonate, that user before continuing. Open Settings → Integrations, click **Connect** next to GitLab, and toggle **Self-Hosted**. Enter your instance domain (use the externally reachable domain if you run split internal and external DNS), the Application ID, and the Secret, then save. Click **Manage GitLab Self-Hosted Permissions**. A pop-up authorizes the integration against your instance. You must still be signed in as, or impersonating, the `factory-droid` user at this point, or the authorization is recorded against the wrong account. Your repositories then appear in the repository list, where you enable the ones Factory should use. ## Token and credential lifecycle | Event | What happens | | --- | --- | | Expiring tokens | Factory refreshes them automatically. No admin action is needed. | | Non-expiring tokens | The authorization is long-lived and persists until it is revoked on your server. | | Revoked authorization | The integration disconnects for the whole organization. An admin has to reconnect it. | | Rotated client secret | Treated the same as a revocation. Update the secret in Factory and reauthorize. | Because the authorization is org-wide, deleting or deactivating the connecting account on your server disconnects every member at once. This is the main reason to authorize as a service account. ## Troubleshooting Confirm the repository is owned by an organization the authorizing account can see. Personal repositories are never imported, and repositories outside the connected account's visibility are excluded. On GitLab, check that the `factory-droid` user has access to the project. The org-level authorization has expired or was revoked on your server. Reconnect the integration from Settings → Integrations. The domain must resolve to a public IP address. Verify public DNS resolution from outside your network, and contact [support@factory.ai](mailto:support@factory.ai) for the Factory egress addresses if the instance is behind a proxy. The pop-up authorized whichever GitLab session was active. Sign in as or impersonate `factory-droid`, then run **Manage GitLab Self-Hosted Permissions** again. Integration configuration and authorization changes are recorded as `integration_change` events in the [audit log](/enterprise/audit-log). Roles, SSO, and the service accounts that should own org-wide integrations. Proxy, certificate, and network requirements for reaching your instance. Run Droid reviews on pull and merge requests in your CI. Track integration configuration changes as audit events. # Enterprise Controls & Managed Settings Manage Droid defaults and hard policy controls for models, tools, safety, network access, retention, and telemetry. Factory's enterprise story is built on a **single, predictable settings hierarchy**. Orgs express hard policy controls and shared defaults in managed settings, then let projects and users customize only where the policy allows it. Enterprise Controls govern models, tools, safety policies, network restrictions, retention, plugin marketplaces, service accounts, and feature defaults across laptops, CI, VMs, and airgapped environments. This page is the org-admin reference for the settings **hierarchy**, the **precedence** rules that decide who wins, the **merge semantics** for each data type, and the authoritative **org-managed settings schema**. For personal preferences a user manages in the Factory App, see [App settings](/factory-app/settings). For the CLI settings.json reference, see [Settings](/droid-cli/settings). For per-feature how-tos, see [Custom Models (BYOK)](/model-independence/byok), [MCP](/harness/mcp), [Sandbox](/autonomy-and-safety/sandbox), [Hooks](/harness/hooks), [Plugins](/harness/plugins), [Output styles](/droid-cli/output-styles), and [Autonomy Levels](/autonomy-and-safety/auto-run). ## The settings levels Settings are authored in `.factory/` folders, using the **same schema** at every level. What changes between levels is precedence. | Level | Where it lives | Who owns it | | :---- | :------------- | :---------- | | **Org** | Managed-settings endpoint, an org `.factory/` bundle, or a [system-managed `settings.json`](#system-managed-settings-file) | Org administrators | | **Folder** | `/...//.factory/` | Repo maintainers (monorepo subtrees) | | **Project** | `/.factory/` | Repo maintainers | | **User** | `~/.factory/` | Individual developers | Each `.factory/` folder can contain: - `settings.json`: general settings (models, safety, preferences, telemetry). - `hooks.json`: hook definitions. - `mcp.json`: MCP server configurations. - `droids/`, `commands/`, `skills/`, `output-styles/`: droid, command, skill, and output style definitions. At resolution time the four authored levels are placed into a longer **build order** that also includes a command-line overlay, remote dynamic config, and hardcoded defaults. The full order, highest priority first, is: | Priority | Level | Source | | :------- | :---- | :----- | | 1 | **Org** (and org plugins) | Remote managed settings or system-managed `settings.json` | | 2 | **Runtime** | The `--settings` overlay passed on the command line | | 3 | **Folder** | Nested `.factory/` directories in a monorepo subtree | | 4 | **Project** | `/.factory/` | | 5 | **User** | `~/.factory/` | | 6 | **Dynamic** | Remote dynamic config | | 7 | **BuiltIn** | Hardcoded Droid defaults | For **hard controls**, the highest level present wins, so **Org is authoritative**: lower levels can extend a policy where the schema allows accumulation, but they can never weaken or remove it. For **session defaults**, the order inverts and the most-local level wins. The next section explains why. --- ## How precedence works: hard controls vs session defaults This is the single most important distinction on this page. Managed settings resolve through **two precedence systems** at once, and they run in opposite directions. For the user-facing perspective on how personal preferences interact with these controls, see [App settings](/factory-app/settings). | | **Hard controls** | **Session defaults** | | :--- | :--- | :--- | | **Who wins** | Highest level present. Org (priority 1) is authoritative. | Most-local level. Runtime > Folder > Project > User > Org > Dynamic > BuiltIn. | | **What lower levels can do** | Extend where the schema allows accumulation (union arrays, add new locked keys); never weaken, remove, or re-enable. | Override the value chosen by a less-local level, but only within the ceilings set by hard controls. | | **Fields** | `modelPolicy`, `mcpPolicy`, `commandBlocklist` (and allow/deny lists), `sandbox`, managed `hooks`, `strictEnabledPlugins`, `strictKnownMarketplaces`, `subagentModelSettings`, plus scalars like `maxAutonomyLevel`, `subagentAutonomyLevel`, `cloudSessionSync`, `wikiCloudSync`, `voiceDictationEnabled`, `sessionRetentionDays`. | The keys inside `sessionDefaultSettings`: `model`, `reasoningEffort`, `interactionMode`, `autonomyLevel`, `autonomyMode`, `specModeModel`, `specModeReasoningEffort`. | Session defaults rank levels so that **lower (more-local) wins**: Runtime is most authoritative, then Folder, Project, and User, with Org, Dynamic, and BuiltIn acting as fallbacks. So a user who sets a preferred `model` overrides the org's default `model`, and a `--settings` runtime overlay overrides even the user. These choices are still **capped by hard controls**: the org `modelPolicy` allowlist decides which models exist at all, and `maxAutonomyLevel` caps the effective `autonomyLevel` no matter who set it. In short: orgs decide the **boundaries** (hard controls, top-down), and users pick their **preferred defaults inside those boundaries** (session defaults, bottom-up). --- ## Merge semantics Within each precedence system, the actual merge depends on the data type. Factory uses four merge modes so hard policy controls stay intact while defaults can still adapt to local context. | Merge mode | How it works | Examples | Best for | | :--------- | :----------- | :------- | :------- | | Simple values, first wins | For scalar hard-control values (strings, numbers, booleans), the first level that sets the value wins. Lower levels cannot change or remove it. | `maxAutonomyLevel`, `cloudSessionSync`, `wikiCloudSync`, `voiceDictationEnabled`, `sessionRetentionDays` | Organization-wide ceilings, feature gates, and retention decisions. | | Arrays, union | Array fields accumulate across levels. Org entries are always present; project, folder, and user levels can add more without removing or weakening higher-level entries. | Command allow lists and deny lists, enabled hooks | Policies like always-denied commands or always-enabled hooks that teams may still extend. | | Named arrays, locked entries | Entries merge by their identifier. Lower levels can add new identifiers but cannot modify or remove entries defined at a higher level. | `customModels`, keyed by `id` | Centrally managed catalogs that projects and users may extend. | | Objects, locked keys | Keys defined at a higher level are locked. Lower levels can add new keys but cannot modify or delete existing ones. | MCP server definitions | Centralized control over critical configuration, with room for project or user additions. | For example, if the org defines a `customModels` entry with the ID `custom:approved-model`, projects and users can add entries with other IDs, but they cannot change or remove the org-defined entry. The scalar keys inside `sessionDefaultSettings` are the exception to "first wins": the **most-local** level wins, not the first. See [How precedence works](#how-precedence-works-hard-controls-vs-session-defaults). --- ## System-managed settings file For deployments where org policy must be applied **before any user signs in**, including managed laptops, CI runners that aren't authenticated to Factory, airgapped installs, or restricted-network installs, IT/MDM administrators can drop a `settings.json` at a hardcoded, well-known path on each machine: | OS | Path | | ------------ | ----------------------------------------------------- | | macOS | `/Library/Application Support/Factory/settings.json` | | Linux & WSL | `/etc/factory/settings.json` | | Windows | `C:\Program Files\Factory\settings.json` | The file uses the same [org-managed settings schema](#org-managed-settings-schema) as the rest of this page. When this file is present it is the **authoritative source for org-level settings** on that machine. It short-circuits any API fetch or developer-mode environment overrides for org settings. The `factoryTier` signal is still queried from the Factory API on a best-effort basis (failures are non-fatal) so enterprise-gated behaviors keep working for managed deployments. If the file is absent or unreadable, Droid falls back to the existing org-settings flow (Factory App / managed-settings endpoint). If the file is present but malformed or fails schema validation, org settings resolve to an empty policy and the failure is logged. A broken admin deployment is **not** silently bypassed. This is the recommended way to ship base policy with an MDM image or installer; users cannot override or weaken it. ## Org-managed settings schema Below is the full schema for org-managed settings. These are configured by org admins through the Enterprise Controls in the Factory App. Every property is optional; omit any field you do not need to set. ### Session defaults Default session preferences the org ships to all members. Every key here is a [session default](#how-precedence-works-hard-controls-vs-session-defaults): the most-local level wins, capped by hard controls. Users can override these with their own preferences (see [App settings](/factory-app/settings)); the values here act as the org-wide fallback. **Properties** Default model identifier. See [Models](/models) for the available IDs. How much thinking the model does before responding. One of `none`, `dynamic`, `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`. Droid interaction mode. One of `auto`, `spec`, `agi`. Default autonomy level for the session. One of `off`, `low`, `medium`, `high`. Capped by `maxAutonomyLevel`. Model to use when spec mode is active. See [Models](/models). Reasoning effort override for spec mode. Same values as `reasoningEffort`. **Deprecated.** Use `interactionMode` and `autonomyLevel` instead. ### Autonomy, safety, and commands Maximum autonomy level any user or project can set. One of `off`, `low`, `medium`, `high`. A hard ceiling that caps the `autonomyLevel` session default. When `true`, removes the built-in Web Search tool (`WebSearch`) from the CLI tool catalog, tool discovery, and Script tool access. Defaults to `false`; omitting it or setting it to `false` preserves availability subject to existing controls. Droid reads this field only from org-managed settings. It ignores user, project, and folder values, so lower-level settings can neither disable Web Search nor re-enable it against org policy. In Enterprise Controls, turn on **Built-in Network Tools → Disable Web Search** to set `webSearchDisabled` to `true`. This controls tool availability. `builtInToolAutonomyOverrides` controls approval requirements for available tools. Web Fetch (`FetchUrl`) is unchanged and remains subject to its separate autonomy, network, and airgap controls. Per-tool [autonomy level](/autonomy-and-safety/auto-run) required to run a built-in network tool without a confirmation prompt. Map of built-in tool name to `low`, `medium`, or `high`. A tool auto-runs only when the session's Autonomy Level is at or above its configured requirement; otherwise Droid asks for approval. Tools without an entry default to `low`. Enforced from org-managed settings only, so user and project settings cannot weaken it. Configured under **Built-in Network Tools** in Enterprise Controls. Autonomy level required to run the built-in Web Search tool. One of `low`, `medium`, `high`. Defaults to `low`. Autonomy level required to run the built-in Web Fetch tool. One of `low`, `medium`, `high`. Defaults to `low`. Autonomy level applied to subagents spawned by the Task tool. One of `inherit`, `off`, `low`, `medium`, `high`; defaults to `inherit`, which falls back to the parent session's autonomy level. An explicit level is clamped to `maxAutonomyLevel` so it can never exceed the enterprise cap. Complexity-to-model routing for [subagents](/harness/subagents). Applies when a subagent's droid uses `model: inherit` and the parent passes a complexity tier. Model used for `light` complexity tasks. A model identifier, or omit to inherit the spawning session's model. Reasoning effort for the light tier. Same values as `reasoningEffort`. Model used for `medium` complexity tasks. A model identifier, or omit to inherit the spawning session's model. Reasoning effort for the medium tier. Same values as `reasoningEffort`. Model used for `heavy` complexity tasks. A model identifier, or omit to inherit the spawning session's model. Reasoning effort for the heavy tier. Same values as `reasoningEffort`. Enable Droid Shield safety checks. Shell command patterns that are always allowed (accumulated across levels). Shell command patterns that always require confirmation (accumulated across levels). A denylisted command can still run if the user explicitly approves it; use `commandBlocklist` for a hard block. Shell command patterns that can never run (accumulated across levels). Unlike the denylist, blocked commands have no approval path and cannot be bypassed with a wrapper shell, absolute path, or quoting tricks. For the full matching and bypass-resistance behavior, see [Agent Safety & Controls](/enterprise/llm-safety-and-agent-controls). When `true`, only org-managed hooks and hooks from org-enabled plugins are loaded; hooks defined at the user and project levels are dropped. Defaults to `false`, where hooks from every level participate. See [Org-managed hooks](/harness/hooks#org-managed-hooks). Org-managed hook definitions, keyed by [hook event](/harness/hooks#hook-events). Org-defined hooks are always loaded and cannot be removed by lower levels. For event semantics and authoring, see [Hooks](/harness/hooks). **Properties** One key per hook event: `PreToolUse`, `PostToolUse`, `Notification`, `UserPromptSubmit`, `Stop`, `SubagentStop`, `PreCompact`, `SessionStart`, `SessionEnd`. Each value is an array of matcher groups. Unknown event keys are tolerated but warned about at load time. **Matcher group properties** Tool name pattern to match (only applies to `PreToolUse` and `PostToolUse`). Optional regular expression matched against the command. Commands to run when the matcher matches. **Hook command properties** Must be `"command"`. Shell command to execute. Per-command timeout in seconds. Disable all hooks for the session. Show hook output in the session view. Built-in sandboxing for command execution and file access. For sandbox behavior and modes, see [Sandbox](/autonomy-and-safety/sandbox). **Properties** Whether sandboxing is enabled. Sandbox scope. One of `per-command` or `whole-process`. Filesystem access lists: `allowRead`, `allowWrite`, `denyRead`, `denyWrite` (each a string array of paths). Network access: `allowedDomains`, `allowUnixSockets`, `allowAllUnixSockets`, `allowLocalBinding`, and `httpProxyPort` / `socksProxyPort` for proxying egress. ### Models and BYOK For BYOK and gateway setup, see [Custom Models (BYOK)](/model-independence/byok). Org-provisioned custom model definitions. **Properties** Model name sent to the provider API. Unique identifier for this custom model entry. Must start with `custom:` (e.g. `custom:my-internal-model-v1`). Display order index. Base URL of the model API endpoint. API key for authenticating with the provider. Optional -- omit for keyless endpoints. Supports `${VAR_NAME}` environment variable references. **Static API keys are blocked for org-managed models.** The managed-settings endpoint rejects any custom model with a static `apiKey` value. Use a `${VAR_NAME}` environment variable reference instead. Authentication mode for HTTP Anthropic models. Set to `bearer` to send the configured credential as `Authorization: Bearer` instead of `x-api-key`. Bearer mode requires `provider: "anthropic"` and isn't supported for Bedrock models. See [Anthropic bearer authentication](/model-independence/byok#anthropic-bearer-authentication). Provider type. One of `anthropic`, `openai`, `generic-chat-completion-api`, `bedrock-converse`, `factory`, `google`, `xai`, `voyage`. Human-readable name shown in the model picker. Maximum context window size in tokens. Enable extended thinking / chain-of-thought. Maximum tokens for the thinking phase. Maximum output tokens. Additional HTTP headers to send with each request. Map of header name to value. For org-managed models, only approved non-secret headers are visible to non-managers; new non-approved header names are rejected. Additional provider-specific arguments passed to the API. Set to `true` if the model does not support image inputs. AWS Bedrock routing options for this model. See [AWS Bedrock](/model-independence/byok#aws-bedrock) for the full field list. **Properties** AWS region for Bedrock requests. Supports `${VAR_NAME}` interpolation, so one org-distributed entry resolves per environment. AWS profile used to resolve credentials and, as a last resort, a region. Supports `${VAR_NAME}` interpolation. Explicit Bedrock endpoint override, for example a VPC endpoint. A full URL or a `${VAR_NAME}` reference that expands to one. Key-value tags attached to every Bedrock inference call for this model, as a map of name to string value, so Droid usage is attributable in your Bedrock model-invocation logs and cost reports. Keys and values support `${VAR_NAME}` interpolation. Configured per model rather than org-wide, and readable by every member, so it is attribution data and must not carry secrets. See [Request metadata](/model-independence/byok#request-metadata). Org-level model access control. Allowlists use higher-wins (a higher level's allowlist is the ceiling); `blockedModelIds` is a union across levels. **Properties** Model IDs that are explicitly allowed. See [Models](/models) for the available IDs. Model IDs that are explicitly blocked. Combined across levels: once blocked at any level, always blocked. Whether users may add their own custom models (user BYOK). Set to `false` to disable user-supplied keys entirely. Base URLs that custom models are allowed to connect to. Use this to force all custom-model traffic through an approved LLM gateway. Allow all current and future Factory-hosted models except those in `blockedModelIds`. Blocking an individual model keeps that exception disabled while new models are allowed automatically. Whether fast model variants may be used. Model IDs that require an explicit org opt-in before they can be used. Whether the Factory Router may be used with bring-your-own-key credentials. Per-user model overrides. Map of user ID to policy object. **Properties** Model IDs this user is allowed to use. Model IDs this user is blocked from using. ### MCP MCP access is allowlist-only. For server configuration, see [MCP](/harness/mcp). Org-level MCP server access control. **Properties** Whether the MCP allowlist is enforced. Defaults to `false`. Allowed hostname or stdio command/argument matchers, not configured MCP server names. When enforcement is enabled, a server must match at least one entry; an empty or absent allowlist blocks all servers, even if a project or user configures them. There is no MCP blocklist. Remote `http` and `sse` servers are matched by hostname only. `*` matches zero or more characters, including dots: `example.com` includes the parent and its subdomains, while `*.example.com` includes only subdomains, at any depth. HTTP/HTTPS URL entries are reduced to their hostname; scheme, port, credentials, path, query string, and fragment do not restrict access. For example, `https://tools*.example.com/mcp` also allows `http://tools1.example.com:8080/other`. Stdio command and argument matching is unchanged and does not use hostname wildcard rules. See [MCP policy matching rules](/harness/mcp#enterprise-mcp-policy) and the [production and staging gateway example](/harness/mcp#example-allow-production-and-staging-gateways). Per-server autonomy overrides. Map of MCP server name to an object with an optional `defaultLevel` (`low`, `medium`, `high`) and an optional `tools` map of tool name to level. Autonomy overrides matched by URL. Each entry has a `urlPattern` and a `defaultLevel` (`low`, `medium`, `high`). ### Plugins and marketplaces For setup and rollout, see [Internal Plugin Marketplaces](/enterprise/internal-plugin-marketplaces) and [Plugins](/harness/plugins). Map of plugin name in `plugin@marketplace` format to `true` (enabled) or `false` (disabled). Plugins set to `true` are automatically installed on CLI startup once their marketplace is available. By default, entries merge additively across org, project, and user settings. When `true`, the `enabledPlugins` maps at this level and higher levels become the exhaustive plugin allowlist. Enables and manual installs at lower levels are rejected unless the plugin is listed with value `true`, and Factory's default plugins are disabled unless listed. Existing plugins remain installed and reactivate if the policy is removed. When omitted at every level, `enabledPlugins` remains additive. An explicit `false` at a higher level disables lower-level strict policies. Older CLI versions that do not recognize this setting do not enforce it. Additional plugin marketplaces. Map of marketplace name to a source object. Marketplaces listed here are automatically cloned and registered on CLI startup. **Properties** Marketplace source, one of the following shapes, discriminated by the `source` field: **GitHub source** Must be `"github"`. GitHub repository in `owner/repo` format. Optional Git branch or tag to track (e.g. `"main"`, `"v1.2.0"`). Optional full 40-character commit SHA. When set, the marketplace is pinned to this exact commit. **URL source** Must be `"url"`. Git repository URL of the marketplace (e.g. `https://gitlab.com/company/plugins.git`). Optional Git branch or tag to track. Optional full 40-character commit SHA to pin to. **Local source** Must be `"local"`. Local filesystem path to the marketplace. **Git subdirectory source** Must be `"git-subdir"`. Git repository URL hosting the marketplace. Subdirectory within the repository that contains the marketplace manifest. Optional Git branch or tag to track. Optional full 40-character commit SHA to pin to. Marketplace sources that are strictly enforced. A higher-level list is the ceiling and is not broadened by lower levels. Organization-level `extraKnownMarketplaces` sources are included automatically and do not need to be repeated. Each entry is a source object directly (e.g. `{ "source": "github", "repo": "owner/repo" }`), not wrapped in a `source` key like `extraKnownMarketplaces` values. ### Network, sync, and retention IP restrictions for authenticated Factory web, CLI, and API requests. **Properties** IPv4 addresses or CIDR ranges that may access Factory web, CLI, and API key surfaces. Must contain at least one entry when set. Whether sessions are synced to the cloud. Whether AutoWiki content can be stored in the Factory App and synced to GitHub wikis. Enabled by default; when disabled, all AutoWiki API endpoints are blocked and the CLI skips Cloud Sync and GitHub wiki sync. See [AutoWiki](/software-factory/wiki/overview#autowiki-cloud-sync). Number of days cloud sessions are retained. Between 14 and 365. ### Telemetry Pins your organization's OpenTelemetry sink for every member, so a developer's local `OTEL_*` variables cannot redirect the telemetry stream or switch it off. Read from the org level only: a `telemetry` block in project, folder, or user settings is ignored with a warning. Unlike most org fields, it is distributed to every member (the CLI reads it on each member's machine), so treat every value in it as readable by the whole organization. It does not configure Factory's own collector. See [Telemetry & Analytics](/enterprise/telemetry#org-managed-telemetry-settings). **Properties** Master switch for the customer telemetry pipeline. `true` keeps it on even where `OTEL_CUSTOMER_ENABLED=false`; `false` keeps it off even where `OTEL_CUSTOMER_ENABLED=true`. Omit it to fall back to `OTEL_CUSTOMER_ENABLED`. Airgapped deployments export to your collector only, and run the pipeline only where `endpoint` resolves to a usable collector. Base URL of your OTLP HTTP collector, taking precedence over `OTEL_TELEMETRY_ENDPOINT`. An `http`/`https` URL with no query string or fragment, or a `${VAR_NAME}` reference standing in for the whole URL. Static OTLP headers sent on every export, as a map of header name to value. Values may contain `${VAR_NAME}` references; prefer them over literal collector credentials. When `endpoint` is set, the `OTEL_*` header variables never apply, so omitting `headers` means the sink is called with none. Metric flush interval in milliseconds, taking precedence over `OTEL_METRIC_EXPORT_INTERVAL`. Capped at `2147483647`. Whether message content (user and assistant message text, tool call inputs, and tool results) is exported. `true` exports it even where `OTEL_LOG_MESSAGE_CONTENT` is unset; `false` never exports it, overriding the variable. Omit it to let each machine decide with `OTEL_LOG_MESSAGE_CONTENT`. Content only ever reaches your own sink, never Factory's collector, and is never written when no sink is configured. `granularity: "aggregate"` overrides this field. Whether exported data carries per-individual identity: `user` (default) or `aggregate`. Under `aggregate`, direct and indirect user identifiers are stripped from every datapoint on both sinks and message content is never written. Resolved most-restrictive-wins against `OTEL_TELEMETRY_GRANULARITY`, so either source asking for `aggregate` yields `aggregate`. See [Data granularity](/enterprise/telemetry/privacy#data-granularity). Which convention your collector receives: `legacy` (default) or `genai`. The two replace each other rather than stack. The org value wins over `OTEL_TELEMETRY_FORMAT`, which applies only where this is unset; an unrecognized value resolves to `legacy`. Factory's own collector is unaffected. See [Export formats](/enterprise/telemetry/data-reference#export-formats). ### Feature controls Whether voice dictation is available as an input method to members of the organization. ### Personas For configuration through Enterprise Controls, see [Personas](/enterprise/personas). Controls which work personas members can choose during onboarding and what each persona recommends. Enterprise organizations only. **Properties** Personas offered during onboarding: any of `engineering`, `product`, `design`, `finance`, `marketing`, `sales`, `operations`, `something-else`. Must contain at least one unique entry. Omit it to offer every persona. Connectors recommended to members of a persona during activation. Map of persona to an array of up to 20 unique connector slugs from Factory's connector catalog. Recommendations only surface for connectors currently enabled for the organization; personas without an entry fall back to generic suggestions. ### Droid Computers Whether Factory-managed Droid Computers are enabled for the organization. Emails permitted to use managed Droid Computers. Whether bring-your-own-machine Droid Computers are enabled. Emails permitted to use bring-your-own-machine computers. ### Missions Org-level access control for missions. **Properties** Whether mission access is restricted. Defaults to `false`. User IDs allowed to use missions when access is restricted. ### Org administration Include a `Co-authored-by: Droid` trailer in git commits. Automatically connect to the IDE on session start. Disable automatic CLI updates (useful for managed or pinned deployments). Restrict whether members can see other members of the organization. When `true`, only Owners and Managers can create API keys. Whether advanced analytics are enabled for the organization. Controls the weekly email to organization Managers and Owners that lists users and active service accounts at or above 80% of their monthly credit cap. The summary is sent when this setting is omitted or `false`; set it to `true` to stop the summary. Free-text guidance for the Factory Router. Routing rules for the Factory Router. Each rule has an optional `when` condition and a `guidance` string. Environment variables injected into every Droid session at startup, as a map of name to string value. Only the organization-managed block is honored: a block at any other level, including a `--settings` runtime overlay, is ignored without a warning, so neither an untrusted repository nor a command-line flag can inject variables. Values may reference existing variables with `${VAR_NAME}` syntax, resolved against the environment as it was before the block was applied rather than against other keys in the same block; a reference to a variable the machine does not set resolves to an empty string. A variable already set in the shell is never overwritten, so a developer can always override an injected default. ### Supported model IDs The identifiers used in `allowedModelIds`, `blockedModelIds`, and `requireExplicitOptInModelIds` are the model IDs listed on [Models](/models), which is the single source of truth for available models across providers. Unrecognized IDs are silently dropped rather than failing validation, so policies stay valid across model updates. ## Level responsibilities The same schema applies at every level; the typical division of labor is: {/* sweep-allow: term-bullets */} - **Org** publishes the authoritative policy: allowed models and gateways, BYOK rules, global command allow/deny/block lists, org-standard droids, commands, hooks, and plugin marketplaces, plus retention, cloud sync, network policy, and feature gates. This bundle is distributed to every environment Droid runs in. - **Project and folder** `.factory/` directories (checked into version control) specialize org policy for a codebase: project-specific droids and gateways within the allowed set, hooks that know the repo's tests and deploys, and tighter controls for high-risk repositories. Folder-level directories handle monorepos where subsystems differ. - **User** `~/.factory/` holds personal session defaults where the active policy allows it: a preferred model from the allowed set, preferred reasoning effort, interaction mode, and autonomy level. See [App settings](/factory-app/settings) for the personal preferences surface. Because hard controls remain authoritative, users cannot re-enable disallowed models or tools, loosen command lists, or raise autonomy above an enforced ceiling. --- ## Example: enforcing a model policy Suppose your org wants to allow only approved enterprise models, disallow user-supplied API keys, and force all prompts through a particular LLM gateway. You would: 1. Define the allowed models and gateway endpoints in `modelPolicy` and `customModels`. 2. Set `modelPolicy.allowCustomModels` to `false` to disable user BYOK entirely. 3. Use `modelPolicy.allowedBaseUrls` to require approved gateway endpoints for custom models. Projects and users can still choose **which of the approved models** to use, but cannot break these guarantees. ## Example: environment-specific autonomy Consider an org that wants high autonomy in CI and sandboxed containers but limited autonomy on developer laptops. You could: 1. At org level, set `maxAutonomyLevel` to `high`. 2. In project settings, define environment-aware hooks that inspect environment tags (for example `environment.type=local|ci|sandbox`) and downgrade or block autonomy above `medium` on laptops. 3. Optionally, define stricter folder-level policies for particularly sensitive repos. Users can only choose safer personal defaults within the allowed space. ## Example: disable built-in web search For a connected on-premises org that cannot allow Factory-hosted web search, add this to org-managed settings without enabling full airgap mode: ```json { "webSearchDisabled": true } ``` This disables the CLI's built-in Web Search tool; it does not block all network egress or all provider access. ## Example: full org-managed settings The `mcpPolicy` below allows `docs.example.com` and its subdomains, plus subdomains of `mcp-gateway.example.com`. For example, `https://docs.example.com/mcp` and `https://ticketing.mcp-gateway.example.com/mcp` pass the policy; `https://mcp-gateway.example.com/mcp` does not, because `*.` excludes the parent hostname. These are hostname matchers, not the names used in `mcp.json`. ```json { "sessionDefaultSettings": { "model": "", "reasoningEffort": "high", "interactionMode": "auto", "autonomyLevel": "medium", "specModeModel": "", "specModeReasoningEffort": "medium" }, "maxAutonomyLevel": "high", "builtInToolAutonomyOverrides": { "WebSearch": "high", "FetchUrl": "medium" }, "subagentAutonomyLevel": "medium", "cloudSessionSync": true, "voiceDictationEnabled": true, "includeCoAuthoredByDroid": true, "enableDroidShield": true, "ideAutoConnect": true, "commandAllowlist": ["npm *", "yarn *", "pnpm *", "make *"], "commandDenylist": ["rm -rf /", "sudo *"], "commandBlocklist": ["shutdown", "mkfs", "curl"], "allowManagedHooksOnly": true, "hooks": { "PreToolUse": [ { "matcher": "Execute", "hooks": [ { "type": "command", "command": "\"$FACTORY_PROJECT_DIR\"/.factory/hooks/audit-command.sh", "timeout": 30 } ] } ] }, "customModels": [ { "model": "my-internal-model", "id": "custom:my-internal-model-v1", "index": 0, "baseUrl": "https://llm-gateway.internal.example.com/v1", "apiKey": "${INTERNAL_MODEL_API_KEY}", "provider": "generic-chat-completion-api", "displayName": "Internal Model v1", "maxContextLimit": 128000, "enableThinking": true, "thinkingMaxTokens": 8192, "maxOutputTokens": 16384, "extraHeaders": { "X-Team": "platform" }, "extraArgs": {}, "noImageSupport": false } ], "modelPolicy": { "allowedModelIds": [""], "blockedModelIds": [], "allowCustomModels": false, "allowedBaseUrls": ["https://llm-gateway.internal.example.com/v1"], "allowAllFactoryModels": false }, "mcpPolicy": { "enabled": true, "allowlist": ["https://docs.example.com", "*.mcp-gateway.example.com"] }, "networkPolicy": { "allowedIps": ["10.0.0.0/8", "192.0.2.0/24"] }, "sandbox": { "enabled": true, "mode": "per-command", "filesystem": { "denyRead": ["~/.ssh", "~/.aws"], "denyWrite": ["/etc"] }, "network": { "allowedDomains": ["api.github.com", "registry.npmjs.org"] } }, "restrictApiKeyCreationToManagers": true, "sessionRetentionDays": 90, "advancedAnalyticsEnabled": true, "disableWeeklyUsageSummary": false, "userModelPolicies": { "user-abc-123": { "allowedModelIds": [""], "blockedModelIds": [] } }, "enabledPlugins": { "my-org-plugin@internal-marketplace": true, "experimental-plugin@internal-marketplace": false }, "strictEnabledPlugins": true, "extraKnownMarketplaces": { "internal-marketplace": { "source": { "source": "github", "repo": "my-org/factory-plugins" } } }, "strictKnownMarketplaces": [] } ``` ## Models, gateways, MCP, and tooling The hierarchy governs which models and external systems Droid can reach. The org-level controls are the schema fields above; each area also has a dedicated page: {/* sweep-allow: term-bullets */} - **Models and BYOK**: `modelPolicy` (`allowedModelIds`, `blockedModelIds`, `allowCustomModels`, `allowedBaseUrls`) and `customModels`. See [Models](/models) for IDs and [Custom Models (BYOK)](/model-independence/byok) for gateway and provider setup, including cloud platforms (Bedrock, Vertex, Azure OpenAI) and self-hosted endpoints. - **MCP servers**: `mcpPolicy` (allowlist-only), plus `mcpAutonomyOverrides` and `mcpAutonomyUrlOverrides` for per-server or per-URL autonomy. See [MCP](/harness/mcp). - **Droids, commands, and hooks**: published in the org `.factory` bundle; set `allowManagedHooksOnly` to `true` to ignore project- and user-defined hooks. See [Hooks](/harness/hooks) and [Agent Safety & Controls](/enterprise/llm-safety-and-agent-controls). --- ## Putting it all together Enterprise Controls underpin everything described in the other enterprise pages: - [Identity & Access](/enterprise/identity-and-access): who can change which level of settings. - [Data Flows & Privacy](/enterprise/privacy-and-data-flows): where data and telemetry are allowed to go. - [Deployment Patterns](/enterprise/network-and-deployment): which environments Droid can run in and how it connects. - [Agent Safety & Controls](/enterprise/llm-safety-and-agent-controls): policies for commands, tools, and Droid Shield. - [Compliance & Audit](/enterprise/compliance-audit-and-monitoring): guarantees and telemetry used to prove compliance. By expressing policy once at the right level, you can run Droid across cloud, hybrid, and airgapped environments **without per-machine drift or one-off configuration**. See how command lists, sandbox, and hooks work in practice. Measure adoption and cost with OTEL export and the Analytics API. # Personas Configure which work personas members can choose during onboarding and recommend connectors per persona for a personalized activation experience. A persona describes the kind of work a member does: Engineering, Product, Design, Finance, Marketing, Sales, or Operations. New members pick a persona when they first enter the Factory App, and Factory personalizes their activation around that choice. Every organization gets this experience with Factory's defaults: all personas are offered, and the walkthrough adapts to the chosen role. Enterprise organizations can shape it further. An **Owner** or **Manager** chooses which personas members can select, then assigns recommended [Connectors](/harness/connectors) to each persona so a new member's first steps point at the tools their role actually uses. ## Configure personas In the Factory App, open **Settings → Enterprise Controls → Personas**. The table lists every persona with an availability switch. Turn off personas that do not apply to your organization. At least one persona must stay enabled, so the last enabled switch is locked. Select a persona row to open its defaults. On the **Connectors** tab, check the connectors to recommend for that persona, up to 20 per persona. Only connectors already enabled for the organization appear; enable more under **Enterprise Controls → Connectors** first. Save the persona defaults, then save the Enterprise Controls form. New members see the updated personas and recommendations the next time they onboard. Configuring personas requires an Enterprise organization; personas themselves are available to everyone. Owners and Managers can edit the configuration; other members see it read-only. ## What members experience Persona settings shape a member's first session in the Factory App: | Surface | Personalization | | :------ | :-------------- | | Onboarding role question | The "What best describes your role?" step lists only the personas the organization enabled. | | Walkthrough | The checklist in the new-session view adapts its items to the chosen persona. An engineer is prompted to craft a design doc; a product manager is prompted to map the product surface and draft a spec. | | Connect your tools | When the persona has recommended connectors, the walkthrough's connection prompt names them, and Droid walks the member through connecting each one and how it fits their workflow. | Recommendations do not connect anything on the member's behalf. Each member still authorizes every app with their own identity, and connector calls keep their normal [Autonomy Level](/autonomy-and-safety/auto-run) behavior. ## How persona settings resolve | Situation | Behavior | | :-------- | :------- | | No persona configuration | All personas are offered and the walkthrough uses its generic items. | | Persona without recommended connectors | Members of that persona get generic connector suggestions. | | Recommended connector later disabled for the organization | The recommendation is hidden from members. Persona settings are not pruned, so re-enabling the connector restores the recommendation. | | Member onboarded before the configuration changed | The member's persona choice is already stored; changes apply to members who have not completed onboarding yet. | Persona configuration is stored in the organization's managed settings as `personaSettings`. See the [org-managed settings schema](/enterprise/hierarchical-settings-and-org-control#org-managed-settings-schema) for the raw shape. ## Troubleshooting Persona controls require an Enterprise organization. Confirm that you selected the intended organization and that your account is an Owner or Manager. If your Enterprise organization still does not show the section, contact [support@factory.ai](mailto:support@factory.ai) to confirm it is enabled for your organization. At least one persona must stay enabled. Enable another persona first, then turn off the one you want to remove. Persona recommendations draw from the connectors enabled for the organization. Ask an Owner or Manager to enable the app under **Enterprise Controls → Connectors**, then add it to the persona. Check that the connector is still enabled for the organization. A disabled connector stays in the persona's saved settings but is hidden from members until it is enabled again. Persona selection happens once, during onboarding. Members who already completed onboarding keep their stored persona and do not see the role question again. Enable apps for the organization and manage member connections. Understand org-managed settings, precedence, and the full schema. Manage Factory roles, membership, SSO, and Directory Sync. Personal session defaults members manage in the Factory App. # Internal Plugin Marketplaces Centralize approved Droid plugins across your organization with internal marketplaces and managed settings. Internal Plugin Marketplaces are private plugin catalogs where organizations maintain approved Droid plugins for company-wide distribution. Instead of individual teams discovering and vetting plugins independently, an internal marketplace provides a curated catalog of capabilities that are pre-approved, maintained, and ready to use. For day-to-day plugin usage and authoring, see [Plugins](/harness/plugins). Use this page when you need enterprise distribution, managed marketplace sources, and org-level rollout controls. ## Why use Internal Plugin Marketplaces? | Challenge | Solution | |-----------|----------| | **Inconsistent tooling** across teams | Single source of approved plugins ensures everyone uses the same capabilities | | **Security and compliance review** for each plugin | Vet plugins once at the org level, distribute everywhere | | **Onboarding new developers** | New team members instantly access all approved capabilities | | **Role-specific tooling** | Package plugins by team function (security, frontend, data, etc.) | | **Version control** | Manage plugin versions centrally, roll out updates organization-wide | ## Set up an internal plugin marketplace An internal plugin marketplace is a Git repository containing your organization's approved plugins with a marketplace manifest. ### Repository structure ```text your-org/droid-plugins/ ├── .factory-plugin/ │ └── marketplace.json # Marketplace manifest ├── plugins/ │ ├── security-toolkit/ # Security team plugins │ │ ├── .factory-plugin/ │ │ │ └── plugin.json │ │ └── skills/ │ ├── frontend-standards/ # Frontend team plugins │ │ ├── .factory-plugin/ │ │ │ └── plugin.json │ │ └── skills/ │ ├── data-engineering/ # Data team plugins │ │ └── ... │ └── platform-tools/ # Platform/DevOps plugins │ └── ... └── README.md ``` ### Marketplace manifest Create `.factory-plugin/marketplace.json` to register your plugins: ```json { "name": "acme-corp-plugins", "description": "ACME Corp approved Droid plugins", "owner": { "name": "ACME Platform Team", "email": "platform@acme.com" }, "plugins": [ { "name": "security-toolkit", "description": "Security review, threat modeling, and vulnerability scanning", "source": "./plugins/security-toolkit", "category": "security" }, { "name": "frontend-standards", "description": "React component patterns, accessibility checks, design system integration", "source": "./plugins/frontend-standards", "category": "frontend" }, { "name": "data-engineering", "description": "SQL review, pipeline validation, data quality checks", "source": "./plugins/data-engineering", "category": "data" }, { "name": "platform-tools", "description": "CI/CD helpers, infrastructure review, deployment automation", "source": "./plugins/platform-tools", "category": "platform" } ] } ``` ## Org-level configuration Configure the marketplace at the organization level so it's automatically available to all users. Add to your org-managed settings: ```json { "extraKnownMarketplaces": { "acme-corp-plugins": { "source": { "source": "github", "repo": "your-org/droid-plugins" } } }, "enabledPlugins": { "security-toolkit@acme-corp-plugins": true, "platform-tools@acme-corp-plugins": true } } ``` | Field | Purpose | |-------|---------| | `extraKnownMarketplaces` | Registers the marketplace so users can browse and install plugins | | `enabledPlugins` | Enables plugins and installs them at org scope when missing (optional) | With this configuration: - The marketplace appears automatically when users run `/plugins` - Droid installs the enabled plugins automatically during startup - Users can install additional plugins unless `strictEnabledPlugins` restricts the list ### Restricting marketplaces To prevent users from adding unapproved marketplaces, use `strictKnownMarketplaces`: ```json { "strictKnownMarketplaces": [ { "source": "github", "repo": "Factory-AI/factory-plugins" } ] } ``` When `strictKnownMarketplaces` is set: - Users can only add marketplaces from the approved list - Plugin installations from non-approved marketplaces are blocked Marketplaces the organization declares in `extraKnownMarketplaces` count as approved without being repeated in the list, so `strictKnownMarketplaces` only has to name sources outside your own org settings. A local path in the list has to be absolute or start with `~`, since a relative one points at a different directory for every user, and saving one is rejected. Existing-plugin provenance enforcement is rolling out separately. When it is active for your organization: - Installed plugins whose marketplace is not approved stop loading - Installed plugins whose recorded source differs from the approved source stop loading, even when the marketplace name matches A blocked plugin stays in the installed list and shows the reason in `/plugins`. Approve its marketplace, or have the user reinstall the plugin from a marketplace already on the list. When the name is approved but the source no longer matches, re-add the marketplace to repair the registration, then reinstall the plugin to record the new source. Before existing-plugin enforcement becomes active, audit installed plugins and approve the marketplaces your teams depend on. Once active, Droid checks installed plugins at the next session rather than waiting for an update. If an org list saved by an older CLI contains only relative local paths, Droid denies new installs but keeps existing plugins active. Replace those entries with absolute paths or paths that start with `~`. ### Restricting enabled plugins By default, `enabledPlugins` entries from org, project, and user settings combine. Set `strictEnabledPlugins` to make the entries at the policy level and higher the complete list: ```json { "strictEnabledPlugins": true, "enabledPlugins": { "security-toolkit@acme-corp-plugins": true, "platform-tools@acme-corp-plugins": true } } ``` At org level, project and user settings cannot enable or install other plugins. Their existing entries stay stored but inactive, and can reactivate if the strict policy is removed. Factory's default plugins are also disabled unless they appear in the org list. Older CLI versions do not recognize or enforce `strictEnabledPlugins`. Require a supported version before relying on this policy. See [Hierarchical settings and org control](/enterprise/hierarchical-settings-and-org-control#plugins-and-marketplaces) for precedence and the full settings reference. ## User experience Users manage plugins via the `/plugins` UI: 1. Run `/plugins` to open the plugin manager 2. The Available tab shows plugins from all registered marketplaces including the org marketplace, minus the ones already installed 3. Org-enabled plugins are installed automatically during startup For CLI access: ```bash droid plugin install frontend-standards@acme-corp-plugins droid plugin update frontend-standards@acme-corp-plugins ``` Organization marketplaces are labeled in Marketplaces, and managed installs are labeled in Installed. Available identifies each plugin's source marketplace. ## Organizing plugins Structure your marketplace around the boundaries your platform team already owns. Common patterns include: {/* sweep-allow: term-bullets */} - **Team function**: security, frontend, backend, data, platform, compliance. - **Capability**: code review, testing, documentation, migrations, security. - **Project type**: microservices, monoliths, data pipelines, ML projects. Choose one primary taxonomy so `/plugins` stays predictable for new users. ## Pre-installing plugins For critical capabilities that everyone needs, set `enabledPlugins` in org-managed settings, alongside `extraKnownMarketplaces` as shown in [Org-level configuration](#org-level-configuration). These plugins: - Are installed automatically with `org` scope during startup - Can finish installing after the session becomes usable when a source is slow - Follow the org settings for their version, so `/plugins` offers an org-scoped plugin no update or uninstall action, only its details Move a relative-path org plugin by changing the marketplace pin in org-managed settings. For external Git or npm plugins, update the source entry in the marketplace manifest, then advance the marketplace pin so clients receive that change. ## Version management Relative-path plugins in a Git marketplace use the marketplace checkout commit as their installed version. External Git and npm sources are fetched separately when their plugin is installed or refreshed. Pin a marketplace source with `ref` (branch or tag) or `sha` (full commit SHA) to control its manifest and relative-path plugins: ```json { "extraKnownMarketplaces": { "acme-corp-plugins": { "source": { "source": "github", "repo": "your-org/droid-plugins", "ref": "v2.4.0" } }, "acme-corp-plugins-frozen": { "source": { "source": "github", "repo": "your-org/droid-plugins", "sha": "1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7e8f9a0b" } } } } ``` Common patterns: - Omit `ref` and `sha` - follow the repository's default branch - `ref: "main"` - explicitly follow the `main` branch - `ref: "staging"` - early-adopter channel - `ref: "v1.2.0"` - pin to a release tag - `sha: "<40-char SHA>"` - hard pin to a specific commit; marketplace update validates the checkout but does not fetch or move it See [Plugins · Add and pin marketplaces](/harness/plugins#add-and-pin-marketplaces) for the full source schema. The `version` field in a plugin's own `plugin.json` does not pin anything. Relative-path plugins follow their marketplace pin. External `url` and `git-subdir` entries use `ref` or `sha` on their own source object, while npm entries use `plugins[].source.version`. ## Private repository access For private Git repositories, ensure Droid can authenticate: ### GitHub Enterprise ```bash # Users authenticate via gh CLI gh auth login --hostname github.your-company.com ``` ### GitLab self-hosted ```bash # Configure git credentials git config --global credential.helper store ``` ### SSH-based access Ensure SSH keys are configured for the repository host. ## Local marketplaces For air-gapped environments or when Git access is restricted, you can use local directory marketplaces. This is useful for: - Environments without internet access - Testing plugins before publishing - Internal distribution via shared network drives ### Setting up a local marketplace Create a directory with the standard marketplace structure: ```text /shared/company-plugins/ ├── .factory-plugin/ │ └── marketplace.json └── plugins/ ├── security-toolkit/ │ └── .factory-plugin/ │ └── plugin.json └── code-standards/ └── .factory-plugin/ └── plugin.json ``` ### Adding a local marketplace Via UI: `/plugins` → Marketplaces tab → "Add new marketplace" → enter absolute path Via CLI: ```bash droid plugin marketplace add /shared/company-plugins ``` ### Configuration for auto-registration Use the `local` source type in settings: ```json { "extraKnownMarketplaces": { "company-local-plugins": { "source": { "source": "local", "path": "/shared/company-plugins" } } } } ``` When removing a local marketplace from the marketplace list, Droid does **not** delete the source directory. Only Git-cloned marketplaces have their directories removed on deletion. ## Marketplaces in a subdirectory (`git-subdir`) When your marketplace lives in a subdirectory of a larger repository (rather than at the repo root), use the `git-subdir` source type. It works with any Git host via a repository URL and the path to the marketplace root within it: ```json { "extraKnownMarketplaces": { "acme-corp-plugins": { "source": { "source": "git-subdir", "url": "https://gitlab.com/acme/monorepo.git", "path": "tools/droid-plugins", "ref": "main" } } } } ``` `url` and `path` are required. The selected directory must contain `.factory-plugin/marketplace.json` or `.claude-plugin/marketplace.json`. `ref` (branch or tag) and `sha` (full 40-character commit SHA) are optional and pin the marketplace the same way they do for other Git sources. ## npm-sourced plugins Individual plugins in a marketplace manifest can be sourced from an **npm registry** instead of a Git path. Use the `npm` source type in a plugin's `source` field (it is a per-plugin source only, not a marketplace source): ```json { "name": "acme-corp-plugins", "plugins": [ { "name": "security-toolkit", "description": "Security review and vulnerability scanning", "source": { "source": "npm", "package": "@acme/droid-security-toolkit", "version": "^2.0.0", "registry": "https://npm.internal.acme.com", "authTokenEnvVar": "ACME_NPM_TOKEN" } } ] } ``` | Field | Purpose | |-------|---------| | `package` | npm package name (scoped or unscoped). Required. | | `version` | Exact version, range (`^2.0.0`), or dist-tag (`latest`). Optional. | | `registry` | Custom https registry URL for private packages. Optional. | | `authTokenEnvVar` | Name of the environment variable holding the registry auth token. Optional. | npm-sourced installs record the resolved package version, URL, and integrity when available. A plugin refresh resolves its exact version, range, or dist-tag again. Automatic refresh follows the containing marketplace's update schedule rather than continuously polling the registry. ## Best practices | Practice | Checklist | | --- | --- | | Establish a review process | Require security review for external dependencies, platform-team code review, isolated-environment testing, and documentation before adding plugins to the marketplace. | | Document each plugin | Include what the plugin provides, when to use it, when not to use it, prerequisites, dependencies, and example usage. | | Version semantically | Use **major** versions for breaking command or behavior changes, **minor** versions for backward-compatible capabilities, and **patch** versions for bug fixes. | | Monitor adoption | Track installation counts, active usage metrics, team feedback, issues, and feature requests. | | Plan for deprecation | Announce the timeline, provide a migration path, and keep deprecated plugins available read-only during the transition. | ## Example: financial services org A financial services company sets up its marketplace: **Mandatory plugins** (pre-installed for everyone): - `compliance-checks` - PCI-DSS and SOX compliance validation - `security-scanner` - OWASP vulnerability detection - `audit-logging` - Enhanced audit trail for all Droid actions **Team-specific plugins** (available for install): - `trading-systems` - For quantitative and trading teams - `risk-models` - For risk management teams - `regulatory-reporting` - For compliance teams **Configuration (org-managed-settings.json):** ```json { "extraKnownMarketplaces": { "acme-financial-plugins": { "source": { "source": "github", "repo": "acme-financial/droid-plugins" } } }, "enabledPlugins": { "compliance-checks@acme-financial-plugins": true, "security-scanner@acme-financial-plugins": true, "audit-logging@acme-financial-plugins": true } } ``` This configures the compliance and security plugins for automatic org-scope installation, while specialized teams can add domain-specific capabilities. Plugin basics, installation, and authoring. Scoped subagents that plugins can bundle and distribute. # Organization Management API Administer organization members and usage limits, and review enterprise control history. ## List enterprise control change history `GET /api/v0/organization/enterprise-controls/history` Returns a paginated list of enterprise control setting changes for the organization, ordered by revision descending. Properties for the `settings` object may be added, removed, or renamed between API versions. Clients should parse it as a generic JSON object. ```bash curl 'https://api.factory.ai/api/v0/organization/enterprise-controls/history' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `limit` (`string`) - Query parameter. Maximum number of items to return (1-100) - `cursor` (`string`) - Query parameter. Cursor for pagination **Response:** `200` - Response for status 200 ## Get global user credits limit `GET /api/v0/organization/usage/limits/global` Returns the global per-user credits limit for the organization. This limit applies to all users who do not have an individual override. ```bash curl 'https://api.factory.ai/api/v0/organization/usage/limits/global' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Response:** `200` - Response for status 200 ## Set global user credits limit `PUT /api/v0/organization/usage/limits/global` Set the global per-user credits limit for the organization. This limit applies to all users who do not have an individual override. Pass null to remove the global limit. ```bash curl -X PUT 'https://api.factory.ai/api/v0/organization/usage/limits/global' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Response:** `200` - Response for status 200 ## Get user credits limits `GET /api/v0/organization/usage/limits/users` Returns the global per-user credits limit and any individual user overrides for the organization. ```bash curl 'https://api.factory.ai/api/v0/organization/usage/limits/users' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Response:** `200` - Response for status 200 ## Set user credits limits `PUT /api/v0/organization/usage/limits/users` Set individual per-user credits limits for one or more users. Supports bulk updates (up to 100 users per request). Identify each user by email. Pass null for limit to remove an individual override. ```bash curl -X PUT 'https://api.factory.ai/api/v0/organization/usage/limits/users' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Response:** `200` - Response for status 200 ## List organization users `GET /api/v0/organization/users` Returns a paginated list of users in the Factory organization. ```bash curl 'https://api.factory.ai/api/v0/organization/users' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `limit` (`string`) - Query parameter. Maximum number of items to return (1-100) - `cursor` (`string`) - Query parameter. Cursor for pagination **Response:** `200` - Response for status 200 ## Remove a user `DELETE /api/v0/organization/users` Removes a user from the organization by email. Works for both active members (deletes membership) and pending invitations (revokes the invite). ```bash curl -X DELETE 'https://api.factory.ai/api/v0/organization/users' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `email` (`string `, required) - Query parameter. Email address of the user to remove **Response:** `204` - Response for status 204 ## Invite or update a user `POST /api/v0/organization/users/invite` Sends an invitation to a user to join the organization. If the user is already a member, updates their role to the specified value. If the user has a pending invite, returns the current status without changes. If the organization has reached its seat limit, returns status "error" with a descriptive error message. ```bash curl -X POST 'https://api.factory.ai/api/v0/organization/users/invite' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Response:** `200` - Response for status 200 # Service Accounts API Manage service accounts and their API keys: create, list, update, rotate, and revoke. ## List service accounts `GET /api/v0/service-accounts` Returns all service accounts in the authenticated organization. Requires human authentication with a Manager role or higher; service account credentials are rejected. ```bash curl 'https://api.factory.ai/api/v0/service-accounts' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Response:** `200` - Response for status 200 ## Create a service account `POST /api/v0/service-accounts` Creates a new service account in the authenticated organization. Requires human authentication with a Manager role or higher; service account credentials are rejected. ```bash curl -X POST 'https://api.factory.ai/api/v0/service-accounts' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Response:** `201` - Response for status 201 ## Get a service account `GET /api/v0/service-accounts/{serviceAccountId}` Returns a single service account by ID. Requires human authentication with a Manager role or higher; service account credentials are rejected. ```bash curl 'https://api.factory.ai/api/v0/service-accounts/{serviceAccountId}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `serviceAccountId` (`string`, required) - Path parameter. Service account ID **Response:** `200` - Response for status 200 ## Update a service account `PATCH /api/v0/service-accounts/{serviceAccountId}` Updates a service account description or status. Name is immutable. Requires human authentication with a Manager role or higher; service account credentials are rejected. ```bash curl -X PATCH 'https://api.factory.ai/api/v0/service-accounts/{serviceAccountId}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Parameters** - `serviceAccountId` (`string`, required) - Path parameter. Service account ID **Response:** `200` - Response for status 200 ## Delete a service account `DELETE /api/v0/service-accounts/{serviceAccountId}` Soft-deletes a service account. Tears down all computers owned by the SA and purges stored git credentials (best-effort). Requires human authentication with a Manager role or higher; service account credentials are rejected. ```bash curl -X DELETE 'https://api.factory.ai/api/v0/service-accounts/{serviceAccountId}' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `serviceAccountId` (`string`, required) - Path parameter. Service account ID **Response:** `200` - Response for status 200 ## List service account API keys `GET /api/v0/service-accounts/{serviceAccountId}/api-keys` Returns all API keys for a service account, active and inactive. Each item carries `isRevoked`, `expiresAt`, and `revokedAt` so clients can split them into active/inactive views. ```bash curl 'https://api.factory.ai/api/v0/service-accounts/{serviceAccountId}/api-keys' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `serviceAccountId` (`string`, required) - Path parameter. Service account ID **Response:** `200` - Response for status 200 ## Create a service account API key `POST /api/v0/service-accounts/{serviceAccountId}/api-keys` Mints a new API key for a service account and returns the one-time key value. The value is shown only once and cannot be retrieved later. ```bash curl -X POST 'https://api.factory.ai/api/v0/service-accounts/{serviceAccountId}/api-keys' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Parameters** - `serviceAccountId` (`string`, required) - Path parameter. Service account ID **Response:** `201` - Response for status 201 ## Revoke a service account API key `POST /api/v0/service-accounts/{serviceAccountId}/api-keys/{keyId}/revoke` Revokes an API key for a service account. Once revoked, the key can no longer authenticate. Idempotent. ```bash curl -X POST 'https://api.factory.ai/api/v0/service-accounts/{serviceAccountId}/api-keys/{keyId}/revoke' \ -H 'Authorization: Bearer $FACTORY_API_KEY' ``` **Parameters** - `serviceAccountId` (`string`, required) - Path parameter. Service account ID - `keyId` (`string`, required) - Path parameter. API key ID **Response:** `204` - Response for status 204 ## Rotate a service account API key `POST /api/v0/service-accounts/{serviceAccountId}/api-keys/{keyId}/rotate` DESTRUCTIVE: atomically revokes the existing key and issues a replacement that inherits its name, returning the new one-time key value (shown only once). The previous key stops authenticating the instant this call succeeds, so any client still using it will immediately start failing with 401. Update every consumer with the new value before rotating, or set `gracePeriodMinutes` to keep the old key valid for a short overlap window. ```bash curl -X POST 'https://api.factory.ai/api/v0/service-accounts/{serviceAccountId}/api-keys/{keyId}/rotate' \ -H 'Authorization: Bearer $FACTORY_API_KEY' \ -H 'Content-Type: application/json' \ -d '{ ... }' ``` **Parameters** - `serviceAccountId` (`string`, required) - Path parameter. Service account ID - `keyId` (`string`, required) - Path parameter. API key ID **Response:** `200` - Response for status 200 # Telemetry & Analytics Measure Droid adoption, activity, and cost by exporting OpenTelemetry metrics to your own collector or reading Factory's hosted Analytics API. Enterprise adoption requires more than a good developer experience. You need to understand **who is using Droid, on what, and at what cost**. There are two complementary ways to measure that, and you can use either or both. This page covers setting up the self-hosted export. The rest of the story lives on its own pages: Export Droid's OpenTelemetry **metrics** to your own OTLP-compatible collector. Configuration, examples, and troubleshooting, on this page. Every metric, span, and attribute Droid exports, in both the `droid.*` and GenAI semantic convention formats. Data granularity, message content logging, and what reaches Factory versus your collector. Read aggregated usage, cost, and productivity data from Factory's cloud via the [Analytics API](/api-reference/analytics). Customer-side export is **metrics only by default**; message content can optionally be exported as trace spans to your own endpoint (see [Message content logging](/enterprise/telemetry/privacy#message-content-logging)). Factory's internal tracing is what powers the hosted Analytics API; it is never sent to your collectors. ## Self-hosted OTEL metrics export Droid can export OpenTelemetry (OTEL) metrics to your own OTLP-compatible collector, giving you visibility into Droid activity within your existing observability stack. In connected deployments, metrics are sent to both Factory's collector and yours in the same export cycle. If your collector is unreachable, Factory's own export is not affected. **Spans are yours alone.** Factory's collector never receives a span, in any configuration. Droid builds a trace pipeline only where you have configured a collector of your own that is distinct from Factory's, so with no endpoint set no span is recorded anywhere. ## Configuration There are two places the sink can be configured, and they take precedence in this order: 1. **Org-managed settings.** A `telemetry` block in org-managed settings pins the sink for every member of the organization. See [Org-managed telemetry settings](#org-managed-telemetry-settings). 2. **Environment variables** on the developer's machine, used for any value the organization has not pinned. See [Machine environment variables](#machine-environment-variables). ## Org-managed telemetry settings Environment variables live on the developer's machine, which makes them advisory: anyone can `export OTEL_TELEMETRY_ENDPOINT=...` to redirect the stream or switch it off. Organizations that need the sink enforced can pin it centrally with a `telemetry` block in org-managed settings, edited under **Settings → Enterprise Controls → Raw Configuration**: ```json { "telemetry": { "enabled": true, "endpoint": "https://collector.example.com:4318", "headers": { "Authorization": "Bearer ${OTEL_COLLECTOR_TOKEN}" }, "metricExportIntervalMs": 60000, "logMessageContent": false, "granularity": "user", "format": "legacy" } } ``` Master switch for the customer telemetry pipeline. `true` keeps it on even if the machine sets `OTEL_CUSTOMER_ENABLED=false`; `false` keeps it off even if the machine sets `OTEL_CUSTOMER_ENABLED=true`. Omit it to fall back to `OTEL_CUSTOMER_ENABLED`. In airgapped deployments the pipeline runs only when you have configured a collector of your own; see [Airgapped deployments](#airgapped-deployments). Base URL of your OTLP HTTP collector, taking precedence over `OTEL_TELEMETRY_ENDPOINT`. Must be an `http`/`https` URL with no query string or fragment, or a `${VAR_NAME}` reference standing in for the whole URL. Invalid values are rejected when you save, not on member machines. Because it overrides `OTEL_TELEMETRY_ENDPOINT` on every machine, use a `${VAR_NAME}` reference if machines legitimately need to point at different collectors. Static OTLP headers sent on every export, as a map of header name to value. Values may contain `${VAR_NAME}` references. Headers are transport-only — used to authenticate and route the HTTP export — and are never attached to exported spans or metrics as attributes. Metric flush interval in milliseconds, taking precedence over `OTEL_METRIC_EXPORT_INTERVAL`. Capped at `2147483647`. Whether message content is exported. See [Message content logging](/enterprise/telemetry/privacy#message-content-logging). Whether exported data carries per-individual identity. Defaults to `user`. Set `aggregate` where a works council agreement or a jurisdictional rule forbids per-individual analytics. See [Data granularity](/enterprise/telemetry/privacy#data-granularity). Which convention your collector receives. Defaults to `legacy`. See [Export formats](/enterprise/telemetry/data-reference#export-formats). Factory's own collector is unaffected by this field. Three properties of the block are worth calling out: - **It is read from the organization level only.** A `telemetry` block in project, folder, or user settings is ignored with a warning, so a checked-in repository config or a developer's own `settings.json` cannot redirect the stream. - **It is distributed to every member**, not just Managers and Owners, because the CLI reads it on each member's machine. Treat every value in it as readable by the whole organization. - **It configures your sink only.** Factory's own collector is not affected by it. Prefer a `${VAR_NAME}` reference over a literal collector credential in `headers`, so the secret is distributed to machines out of band (MDM, a secret manager) and never stored in Factory's control plane. If a referenced variable is not set on a machine, Droid disables the customer sink for that session and logs an error naming the missing variable, rather than shipping the literal `${VAR_NAME}` to your collector. When `endpoint` is set in the block, the `OTEL_*` header environment variables never apply to it, because a credential sitting in a developer's shell for unrelated tooling must not be shipped to the organization's collector. Omitting `headers` therefore means the org sink is called with none; put machine-local values in `${VAR_NAME}` references inside `headers` instead. The block is also read once, at startup. A session that begins while logged out has no org settings to read and runs on the `OTEL_*` environment for the life of that process. ## Machine environment variables Set these before launching Droid. Each one configures the sink only where the organization has left that value unpinned; a `telemetry` block in org-managed settings overrides the matching variable. ```bash export OTEL_TELEMETRY_ENDPOINT="https://your-collector.example.com:4318" export OTEL_TELEMETRY_HEADERS="Authorization=Bearer " ``` | Variable | Required | Description | | :------------------------ | :------- | :------------------------------------------------------------------------------------------------------------------- | | `OTEL_TELEMETRY_ENDPOINT` | Yes | Base URL of your OTLP HTTP collector. Metrics are sent to `{endpoint}/v1/metrics`. Falls back to `OTEL_EXPORTER_OTLP_ENDPOINT`. Not required when the organization pins `telemetry.endpoint`. | | `OTEL_TELEMETRY_HEADERS` | No | Comma-separated `key=value` pairs sent as HTTP headers on every export. Values may contain `=` (e.g. base64 tokens). Falls back to `OTEL_EXPORTER_OTLP_HEADERS`. Never applied to an org-pinned endpoint. | | `OTEL_LOG_MESSAGE_CONTENT` | No | Set to `true` or `1` to export message content (user/assistant messages, tool input/results) as trace spans to `OTEL_TELEMETRY_ENDPOINT`. Content is dropped if no customer endpoint is set. | | `OTEL_METRIC_EXPORT_INTERVAL` | No | Metric flush interval in milliseconds. Defaults to `60000`. | | `OTEL_CUSTOMER_ENABLED` | No | Set to `false` to opt this machine out of customer telemetry entirely (useful in tests and local development). This is the same pipeline pinned by `telemetry.enabled`, not a separate stream. | | `OTEL_TELEMETRY_FORMAT` | No | `legacy` or `genai`. Applies only where the organization has left `telemetry.format` unset. Anything unrecognized resolves to `legacy`. See [Export formats](/enterprise/telemetry/data-reference#export-formats). | | `OTEL_TELEMETRY_GRANULARITY` | No | `user` or `aggregate`. Resolved most-restrictive-wins against `telemetry.granularity`, not by precedence. See [Data granularity](/enterprise/telemetry/privacy#data-granularity). | | `OTEL_RESOURCE_ATTRIBUTES` | No | Standard OTEL `key=value,key=value` attributes, applied to your collector only (see [Resource attributes](/enterprise/telemetry/data-reference#resource-attributes)). | | `DROID_PARENT_SESSION_ID` | No | Adds `parent_session.id` to metrics and content spans, for linking headless/CI sub-sessions to a parent. | ## How it works - Once an endpoint is resolved, metrics are sent to it in the same export cycle via a fan-out exporter: no extra timers, no duplication of metric readers. - Failures to your collector do not affect Factory's own export. Each endpoint is isolated. - Metrics use **delta temporality**: each export contains only new values since the last flush (60-second intervals by default). Sum the deltas; do not read a last value. For what arrives on the wire, including the metric catalogs and every attribute, see the [exported data reference](/enterprise/telemetry/data-reference). ## Example configurations ### Generic OTEL collector / Grafana Alloy ```bash export OTEL_TELEMETRY_ENDPOINT="https://collector.example.com:4318" ``` ### Datadog via OTLP ingestion ```bash export OTEL_TELEMETRY_ENDPOINT="http://localhost:4318" ``` ### Datadog direct OTLP intake ```bash export OTEL_TELEMETRY_ENDPOINT="https://otlp.datadoghq.com" export OTEL_TELEMETRY_HEADERS="dd-api-key=" ``` ### New Relic ```bash export OTEL_TELEMETRY_ENDPOINT="https://otlp.nr-data.net:4318" export OTEL_TELEMETRY_HEADERS="api-key=" ``` ### Honeycomb ```bash export OTEL_TELEMETRY_ENDPOINT="https://api.honeycomb.io" export OTEL_TELEMETRY_HEADERS="x-honeycomb-team=,x-honeycomb-dataset=droid-metrics" ``` ## Troubleshooting Verify your collector accepts OTLP HTTP on the `/v1/metrics` path. Confirm `OTEL_TELEMETRY_HEADERS` includes valid auth credentials. Ensure network connectivity from the machine running Droid to your collector. The `telemetry` block is only honored at the organization level. A copy in project, folder, or user settings is ignored, and Droid logs a warning naming the misconfiguration. An `endpoint` or `headers` value that references an environment variable the machine does not set disables the customer sink for that session. Droid logs an error naming the missing variable; set it on the machine or remove the reference. Six resource keys and five identity keys are reserved and cannot be set from the variable. See [Resource attributes](/enterprise/telemetry/data-reference#resource-attributes). There is no org-managed field for resource attributes; set the variable per machine or per runner. Spans require a collector of your own that is distinct from Factory's, so confirm `telemetry.endpoint` or `OTEL_TELEMETRY_ENDPOINT` is set and does not resolve to Factory's collector. If it is set, check whether the organization runs `telemetry.granularity: aggregate` under the `legacy` format, which exports metrics only. Check whether `telemetry.format` was set to `genai`. The two formats replace each other, so your collector now receives `gen_ai.*` instead. See [Export formats](/enterprise/telemetry/data-reference#export-formats). Airgapped deployments export to your collector only, so an unset or unusable `telemetry.endpoint` disables the pipeline entirely rather than falling back to Factory. See [Airgapped deployments](#airgapped-deployments). ## Airgapped deployments Airgapped deployments export to **your collector only**. Droid builds no Factory-bound leg at all, so nothing is sent outward and nothing is queued waiting for a network that is not there. This makes `telemetry.endpoint` (or `OTEL_TELEMETRY_ENDPOINT`) the single telemetry surface in an airgap, and the only one. Without a usable collector of your own, the pipeline does not run: there is no Factory fallback to degrade to. An endpoint that Droid rejects, such as one with an unresolved `${VAR_NAME}` reference or a query string, counts as no collector for this purpose. The `telemetry.enabled` switch still applies on top, and airgapped deployments do not use Factory-hosted analytics. ## Factory-hosted Analytics API In cloud-managed deployments, Factory provides a **hosted analytics view** for platform and leadership teams. It is backed by Factory's internal tracing, which feeds an aggregation pipeline; this internal tracing is separate from the metrics you export above and is never sent to your collectors. The [Analytics API](/api-reference/analytics) exposes aggregated, org-level data, including: - Adoption metrics by org, team, and user. - Model usage and performance trends. - Token consumption and cost estimates for LLM usage. - Tool usage and productivity signals. Use the Analytics API when you want a hosted, aggregated view, particularly for **token and cost data**, which is not part of the customer OTEL metric set. Every metric, span, and attribute Droid exports, in both formats. Data granularity, message content logging, and what reaches Factory. Map telemetry signals to audit trails and regulatory workflows. The org-managed settings surface that pins the telemetry block. # Telemetry Data Reference Reference for every metric, span, and attribute Droid exports over OpenTelemetry, in both the droid.* and GenAI semantic convention formats. This page is the reference for what arrives at your collector once [OTEL export](/enterprise/telemetry) is configured. The [export format](#export-formats) decides which of the two surfaces you receive. ## Export formats Your collector can receive one of two conventions. Factory's own collector receives the `droid.*` metric catalog and is unaffected by this choice. | Format | Your collector receives | | :----- | :---------------------- | | `legacy` (default) | The [`droid.*` metric catalog](#exported-metrics), plus the `droid.message.*` and `droid.tool.*` event spans when [message content logging](/enterprise/telemetry/privacy#message-content-logging) is on. | | `genai` | The [OpenTelemetry GenAI semantic conventions](#genai-semantic-convention-export): `gen_ai.*` spans and `gen_ai.*` metrics, consumable by off-the-shelf GenAI dashboards. | Set it centrally with `telemetry.format` in org-managed settings, or per machine with `OTEL_TELEMETRY_FORMAT`. The organization's value wins: where `telemetry.format` is set to anything at all, the environment variable is not read, and an unrecognized value falls back to `legacy` rather than to the machine. This keeps a developer's shell from switching an organization's collector onto a different convention. **The two formats replace each other; they do not stack.** Under `genai`, your collector receives no `droid.*` metric and no `droid.message.*` / `droid.tool.*` span at all. Rebuild dashboards before you switch an organization over, and expect the switch to be visible as a gap in any `droid.*` query. Durations also change unit. Every `gen_ai.*` duration is in **seconds**, per the semantic conventions. Every `droid.*` duration is in **milliseconds**. ## Exported metrics Under the default `legacy` format, all metrics use the `droid.*` namespace. Every instrument declares a unit. | Metric | Instrument | Unit | Meaning | | :----- | :--------- | :--- | :------ | | `droid.code.files_modified` | Counter | `{file}` | Files created or updated by an edit tool. | | `droid.code.files_read` | Counter | `{file}` | Files read. | | `droid.code.lines_modified` | Counter | `{line}` | Lines added and removed, split by an `operation` attribute. | | `droid.git.commits` | Counter | `{commit}` | Commits created. | | `droid.git.pull_requests` | Counter | `{pull_request}` | Pull requests created. | | `droid.tool.invocations` | Counter | `{invocation}` | Tool executions. | | `droid.tool.execution_time` | Histogram | `ms` | Tool execution duration. | | `droid.command.blocked` | Counter | `{command}` | Commands blocked by policy. Carries the tool name only, never the command text. | | `droid.mcp.tool_invocations` | Counter | `{invocation}` | MCP tool calls. | | `droid.skill.invocations` | Counter | `{invocation}` | Skill runs. | | `droid.skill.installed` | Counter | `{skill}` | Inventory of installed non-builtin skills, emitted once per session. | | `droid.hook.invocations` | Counter | `{invocation}` | Hook commands run. | | `droid.slash_command.invocations` | Counter | `{invocation}` | Slash commands run. | | `droid.auth.login_success` | Counter | `{login}` | Successful logins. | | `droid.session.count` | Counter | `{session}` | Sessions started. Compaction successors and forks are not counted. | | `droid.repo.metadata` | Counter | `{repository}` | One per workspace, carrying repository metadata as attributes. | `droid.skill.installed` was previously emitted once per model call, which multiplied the count by the number of calls in a session. It is now emitted once per session, so values recorded before Droid `0.199.0` are inflated and are not comparable with later ones. ## Common attributes Every data point includes these attributes automatically: | Attribute | Description | | :------------------ | :---------------------------------------------------- | | `user.id` | Authenticated user ID. Removed under [`aggregate` granularity](/enterprise/telemetry/privacy#data-granularity). | | `user.email` | Authenticated user email. Removed under [`aggregate` granularity](/enterprise/telemetry/privacy#data-granularity). | | `organization.id` | Organization ID | | `session.id` | Current Droid session ID | | `model.id` | Active model ID (when available) | | `parent_session.id` | Parent session, when `DROID_PARENT_SESSION_ID` is set | Tool-specific attributes (`tool.name`, `mcp.server`, `skill.name`, etc.) are included where applicable. Resource attributes include `service.name` (`cli`), `service.version`, `os.type`, `os.version`, `host.arch`, and `terminal.type`, plus `telemetry.granularity` when the organization runs in aggregate mode. ## GenAI semantic convention export Set `telemetry.format` to `genai` and your collector receives the OpenTelemetry GenAI semantic conventions instead of the `droid.*` surface. This is what off-the-shelf GenAI dashboards expect, so you can point one at Droid without writing custom queries. ### Spans Each user turn produces one flat trace: a root span with the model calls and tool runs as direct children. Tool spans are siblings of model calls, never nested inside them. | Span | Kind | Parent | | :--- | :--- | :----- | | `invoke_agent droid` | Internal | None. This is the trace root. | | `chat ` | Client | `invoke_agent droid` | | `execute_tool ` | Internal | `invoke_agent droid` | Attributes, by span: | Attribute | Span | Notes | | :-------- | :--- | :---- | | `gen_ai.operation.name` | all | `invoke_agent`, `chat`, or `execute_tool`. | | `gen_ai.agent.name` | root | Always `droid`. | | `gen_ai.conversation.id` | all | The Droid session ID. | | `session.id` | all | The same value, kept under its general-convention name so you can join against metrics. | | `error.type` | all | Present on failure. On a tool span, the literal `rejected` means the call was denied by policy rather than failed. | | `gen_ai.request.model` | `chat` | Omitted when no model ID is known. | | `gen_ai.provider.name` | `chat` | See [Provider names](#provider-names). | | `gen_ai.response.finish_reasons` | `chat` | An array of one string. | | `server.address`, `server.port` | `chat` | The model API host. The port is omitted when the scheme default applies. | | `gen_ai.tool.name`, `gen_ai.tool.call.id` | `execute_tool` | A model-authored name or ID that fails a shape check is replaced with `invalid_tool_name` or `invalid_tool_use_id`. | | `gen_ai.tool.type` | `execute_tool` | Always `function`. | | `gen_ai.input.messages` | root | Message content only. See [Message content logging](/enterprise/telemetry/privacy#message-content-logging). | | `gen_ai.output.messages` | `chat` | Message content only. | | `gen_ai.tool.call.arguments`, `gen_ai.tool.call.result` | `execute_tool` | Message content only. | `gen_ai.input.messages` and `gen_ai.output.messages` are JSON strings holding the conventions' message array: ```json [{ "role": "user", "parts": [{ "type": "text", "content": "..." }] }] ``` ### Metrics Every GenAI instrument is a **histogram**, including the two call counts. To total tool calls, sum the histogram rather than reading a counter. | Metric | Unit | Meaning | | :----- | :--- | :------ | | `gen_ai.client.operation.duration` | `s` | Model call duration. | | `gen_ai.client.operation.time_to_first_chunk` | `s` | Model call time to first streamed chunk. | | `gen_ai.execute_tool.duration` | `s` | Tool execution duration. | | `gen_ai.invoke_agent.duration` | `s` | Turn duration. | | `gen_ai.invoke_agent.inference_calls` | `{inference_call}` | Model calls per turn. | | `gen_ai.invoke_agent.tool_calls` | `{tool_call}` | Tool calls per turn. | Datapoints carry `gen_ai.operation.name`, plus `gen_ai.request.model` and `gen_ai.provider.name` on the model instruments, `gen_ai.tool.name` on the tool instrument, `gen_ai.agent.name` on the turn instruments, and `error.type` on failures. The [common attributes](#common-attributes) are stamped on top of those. ### Provider names `gen_ai.provider.name` is the model vendor, derived from the model ID prefix, not Factory's routing provider. It resolves to one of `anthropic`, `openai`, `gcp.gemini`, `x_ai`, `deepseek`, or `mistral_ai`. A model whose vendor has no value in the conventions gets no provider attribute rather than a guessed one, so do not build a dashboard that requires the attribute to be present. Two behaviors will otherwise look like gaps in your data: - **A denied tool call produces a span but no duration datapoint.** The call never ran, and a zero-second measurement would distort the distribution of the calls that did. Denied calls still count toward `gen_ai.invoke_agent.tool_calls`. - **Subagent turns are separate trace roots.** The `parent_session.id` linkage that the `droid.*` surface carries has no equivalent in the GenAI conventions, so a subagent's trace cannot be joined to its parent's under this format. ## Resource attributes To attach your own identity or lineage attributes (for example `team.id`, `repo`, `environment`), set the standard OTEL variable per machine: ```bash export OTEL_RESOURCE_ATTRIBUTES="team.id=platform,environment=ci" ``` Your pairs reach **your collector only**. They are never sent to Factory's, including a pair that spoofs an identity key. They are applied in three places: - On the **resource** of exported metrics. - On the **resource** of exported spans. - Spread onto individual **metric datapoint attributes**, which is legacy behavior kept because existing dashboards query it. There is no span-attribute equivalent. Two collisions resolve against you. Droid's own resource values win, so a pair setting `service.name`, `service.version`, `os.type`, `os.version`, `host.arch`, or `terminal.type` is ignored. On datapoints, the identity keys `user.id`, `user.email`, `organization.id`, `session.id`, and `parent_session.id` are dropped outright rather than merged, so an operator-set pair cannot impersonate a user. There is **no** org-managed (Raw Configuration) field for resource attributes, and `headers` cannot carry them. To attribute headless or CI runs, set `OTEL_RESOURCE_ATTRIBUTES`, and `DROID_PARENT_SESSION_ID` for lineage, in each runner's environment. Configure the OTEL export that delivers this data. Data granularity and message content logging. # Telemetry Privacy Controls Control whether Droid telemetry carries per-user identity or message content, and see exactly what reaches Factory versus your own collector. Two org-managed controls decide what Droid telemetry contains: `telemetry.granularity` for per-individual identity, and `telemetry.logMessageContent` for message content. This page covers both, and where each kind of data can go. ## What reaches Factory versus your collector | Data | Factory's collector | Your collector | | :--- | :------------------ | :------------- | | Metrics | Yes (never in an [airgapped deployment](/enterprise/telemetry#airgapped-deployments)) | Yes, once an endpoint is configured. | | Spans, including all message content | Never, in any configuration. | Only with a collector of your own, and content only with the opt-in below. | | [`OTEL_RESOURCE_ATTRIBUTES`](/enterprise/telemetry/data-reference#resource-attributes) pairs | Never. | Yes. | `telemetry.granularity` applies at the source, before export, so it binds Factory's collector and yours equally. `telemetry.logMessageContent` only controls whether message content is written to your own span export (Factory never receives spans). ## Data granularity Some organizations operate under a works council agreement or a jurisdictional rule that forbids per-individual analytics. `telemetry.granularity` is the control for that, applied at the source so identifying data is never written rather than filtered downstream. | Mode | Behavior | | :--- | :------- | | `user` (default) | Datapoints carry per-user identity. | | `aggregate` | `user.id`, `user.email`, `commit.hash`, `pr.number`, and `pr.url` are removed from every datapoint on both collectors, and message content is never written regardless of any other setting. | `organization.id` survives aggregate mode, because it is the dimension you aggregate over. So do `session.id`, `parent_session.id`, `model.id`, and every workload attribute such as `tool.name`, `repo.*`, and `skill.name`. Data in aggregate mode carries the resource attribute `telemetry.granularity: aggregate`, so a downstream consumer can tell which mode produced it. Resolution is **most-restrictive-wins**, not a precedence order: either `telemetry.granularity` or `OTEL_TELEMETRY_GRANULARITY` asking for `aggregate` yields `aggregate`, and an organization set to `user` cannot override a machine set to `aggregate`. A value that cannot be read resolves to `aggregate`. Aggregate mode overrides every content opt-in, including `telemetry.logMessageContent: true`. It also changes what spans you get: - **`legacy` plus `aggregate` produces no spans at all.** The `droid.message.*` and `droid.tool.*` event spans are per-user by construction, so there is nothing left to strip them down to. Metrics still flow. - **`genai` plus `aggregate` keeps the full trace**, with the four content attributes stripped. Use this combination if you need trace structure under a compliance mandate. ## Message content logging Message content export is off by default. Metrics fan out to both Factory and your collector; message content does not. The organization's controls are authoritative: | Control | Effect | | :------ | :----- | | `telemetry.granularity: aggregate` in org-managed settings | Content is never exported, overriding every setting below. See [Data granularity](#data-granularity). | | `telemetry.logMessageContent: true` in org-managed settings | Content is exported on every member's machine, whether or not they set `OTEL_LOG_MESSAGE_CONTENT`. | | `telemetry.logMessageContent: false` in org-managed settings | Content is never exported, even on a machine that sets `OTEL_LOG_MESSAGE_CONTENT`. | | Property omitted from org-managed settings | Each machine decides for itself with `OTEL_LOG_MESSAGE_CONTENT`. | To enable message content export for the current process (for example, for session auditing) where the organization has left the decision open, set: ```bash export OTEL_LOG_MESSAGE_CONTENT=true # or 1 ``` When enabled, Droid emits message content as OTEL **trace spans**, on the `/v1/traces` path of your endpoint. Where it lands depends on your [export format](/enterprise/telemetry/data-reference#export-formats). Under `legacy`, content arrives as four span types: - User messages (`droid.message.user`) - Assistant responses (`droid.message.assistant`) - Tool calls and their inputs (`droid.tool.call`) - Tool results (`droid.tool.result`) Under `genai`, the same content arrives as four attributes on the [GenAI spans](/enterprise/telemetry/data-reference#spans): `gen_ai.input.messages`, `gen_ai.output.messages`, `gen_ai.tool.call.arguments`, and `gen_ai.tool.call.result`. **Message content is exported raw, with no redaction.** There is no PII scrubbing and no secret detection; the only transformation is truncation of any single attribute to 32,000 characters. These spans contain verbatim prompts, responses, tool inputs (including file contents and command arguments), and tool results. Make your data-classification decision accordingly before enabling. **Spans are sent only to the collector you configure.** Droid does not fan out any span to Factory's collector. - A customer endpoint (`telemetry.endpoint` or `OTEL_TELEMETRY_ENDPOINT`) is **required**. With none set, Droid builds no trace pipeline, so content is never written rather than being written and dropped. - If your endpoint resolves to the same URL as Factory's collector, it does not count as a collector of your own: no span is recorded and content logging is disabled. Adding credentials to the URL does not make it distinct. - Metrics export is unaffected and continues to fan out to both Factory and your collector. Configure the OTEL export these controls govern. The full picture of what leaves a developer's machine and why. # Factory Analytics API REST API for organization and personal usage, Factory Standard Credits consumption, tool usage, and productivity metrics. The Factory Analytics API returns organization-level and self-scoped usage data for Factory. Query Factory Standard Credits consumption, tool invocations, user activity, and productivity metrics across your organization or for the authenticated user. --- ## Authentication All requests require a Factory API key in the `Authorization` header. ```bash Authorization: Bearer fk-your-api-key ``` [Factory API keys settings](https://app.factory.ai/settings/api-keys) ### Permissions The organization-level endpoints require the **Manager** or **Owner** role. The personal cost endpoint is available to members with the **User** role and returns data only for the authenticated user. --- ## Base URL `https://api.factory.ai/api/v1/analytics` --- ## Response format All responses follow a consistent envelope structure: ```json { "data": [ ... ], "meta": { ... } } ``` | Field | Type | Description | | :----- | :----- | :------------------------------------------------------------- | | `data` | array | Array of result objects (one per day, or per group when using `group_by`) | | `meta` | object | Request metadata: `org_id`, `start_date`, `end_date`, and pagination info for `/users` | --- ## Endpoints The Analytics API provides six endpoints, each focused on a specific category of metrics: | Endpoint | Description | | :--------------- | :------------------------------------------------------ | | `/tokens` | Organization-wide Factory Standard Credits by model and user | | `/cost/me/query` | Self-scoped Factory Standard Credits and attribution metrics | | `/tools` | Tool invocations and autonomy metrics | | `/activity` | Daily, weekly, and monthly active users | | `/productivity` | File operations and git activity | | `/users` | Per-user metrics with pagination | --- ## Understanding `group_by` Several endpoints support a `group_by` parameter. Here's how it works: - **Without `group_by`**: Returns one row per day with nested breakdowns (e.g., `by_tool`, `daily_active_users_by_client`). Use this when you want all dimensions in a single response. - **With `group_by`**: Flattens one of those nested arrays into separate rows. Each row has a `group_key` field identifying the dimension value. Use this when piping data into tools that expect flat rows (spreadsheets, BI tools, time-series databases). For example, `/activity` without `group_by` returns `daily_active_users_by_client` as an object. With `group_by=client`, you get separate rows for `terminal-ui`, `web`, and `non-interactive-cli` - useful for plotting each client type as its own line on a chart. ## Factory Standard Credits usage Returns daily Factory Standard Credits consumption across your organization. Start date in `YYYY-MM-DD` format End date in `YYYY-MM-DD` format Set to `model` to group results by model Date in `YYYY-MM-DD` format Factory Standard Credits consumed, computed from raw input and raw output tokens with cache discounts Raw input tokens sent to model Tokens generated by model Tokens read from prompt cache Tokens written to prompt cache Breakdown per model Breakdown per user ```bash # Factory Standard Credits usage for a date range curl -H "Authorization: Bearer $FACTORY_API_KEY" \ "https://api.factory.ai/api/v1/analytics/tokens?startDate=2026-01-14&endDate=2026-01-28" # Grouped by model curl -H "Authorization: Bearer $FACTORY_API_KEY" \ "https://api.factory.ai/api/v1/analytics/tokens?startDate=2026-01-15&endDate=2026-01-15&group_by=model" ``` --- ## Personal Factory Standard Credits usage Returns Factory Standard Credits and attribution metrics for the authenticated user. The API derives the user ID from the bearer credential and does not accept a user selector, so callers cannot query another user. This endpoint is available to Enterprise organizations with Analytics enabled. Selects the metric returned in `data`. Start date in `YYYY-MM-DD` format. End date in `YYYY-MM-DD` format. Use yesterday or earlier. | Value | Description | | :--------------------------- | :----------------------------------------------- | | `headline_daily` | Daily Factory Standard Credits consumed | | `cost_summary` | Total sessions, messages, credits, and averages | | `top_sessions` | Sessions with the highest credit consumption | | `by_model` | Credit consumption by model | | `fsc_by_model_daily` | Daily credit consumption grouped by model | | `fsc_by_multiplier_daily` | Daily credit consumption grouped by multiplier | | `by_ticket` | Credit consumption attributed to tickets | | `by_pr` | Credit consumption attributed to pull requests | | `my_activity_summary` | Activity totals for the authenticated user | | `my_time_spent` | Estimated time spent by the authenticated user | | `my_intent_breakdown` | Credit consumption grouped by intent | | `my_completed_tickets_daily` | Daily completed-ticket counts | | `my_completed_tickets` | Completed-ticket details | | `my_merged_prs_daily` | Daily merged-pull-request counts | | `my_merged_pr_detail` | Merged-pull-request details | The fields inside `data` depend on `queryKey`. For `headline_daily`, each row contains: ```json { "data": [ { "date": "2026-08-01", "fsc": 123456 } ], "meta": { "org_id": "org_01HPMQ8ABCDE7Y7PR3TTZY4KLM", "start_date": "2026-08-01", "end_date": "2026-08-31" } } ``` Rows for the requested `queryKey`, scoped to the authenticated user. Organization associated with the bearer credential. Start of the requested UTC date range. End of the requested UTC date range. ```bash # Daily personal usage curl -H "Authorization: Bearer $FACTORY_API_KEY" \ "https://api.factory.ai/api/v1/analytics/cost/me/query?queryKey=headline_daily&startDate=2026-08-01&endDate=2026-08-31" ``` --- ## Tool usage Returns daily tool invocations, MCP usage, skills, slash commands, and autonomy metrics. Start date in `YYYY-MM-DD` format End date in `YYYY-MM-DD` format Set to `tool_name` to group results by tool ```json { "data": [ { "date": "2026-01-15", "tool_calls": 45000, "by_tool": [ { "tool": "Read", "invocations": 12500 }, { "tool": "Edit", "invocations": 8200 }, { "tool": "Execute", "invocations": 6100 } ], "mcp_users_with_mcp": 42, "mcp_by_server": [ { "server": "github", "invocations": 1200 }, { "server": "notion", "invocations": 850 } ], "skills_invocations": 320, "skills_by_name": [ { "name": "browser", "count": 180 }, { "name": "frontend-ui", "count": 95 } ], "slash_commands_invocations": 1500, "slash_commands_by_name": [ { "name": "review", "count": 420 }, { "name": "test", "count": 380 } ], "hooks_invocations": 2800, "hooks_by_event": [ { "event": "PostToolUse", "matcher": "*.ts", "command": "eslint --fix", "count": 1200 } ], "web_users": 42, "autonomy_ratio_avg": 8.5, "autonomy_ratio_p50": 6.2, "autonomy_ratio_p90": 18.4, "tool_calls_per_session_avg": 45.2, "user_turns_per_session_avg": 5.3, "tool_autonomy_level_ratio": { "auto_high": 0.35, "auto_medium": 0.42, "auto_low": 0.18, "manual": 0.05 } } ], "meta": { "org_id": "org_01HPMQ6ABCDE...", "start_date": "2026-01-15", "end_date": "2026-01-15" } } ``` Date in `YYYY-MM-DD` format Total tool invocations Breakdown by tool name Users who used MCP servers Invocations per MCP server Total skill activations Breakdown by skill Total slash command uses Breakdown by command Total hook executions Breakdown by event type Users who used web/workspace interface Average tool calls per user turn Median autonomy ratio 90th percentile autonomy ratio Average tool calls per session Average user messages per session Distribution of autonomy levels When `group_by=tool_name`, returns one row per tool per day inside `data`: ```json { "data": [ { "date": "2026-01-15", "group_key": "Read", "tool_calls": 12500 }, { "date": "2026-01-15", "group_key": "Edit", "tool_calls": 8200 } ], "meta": { "org_id": "org_01HPMQ6ABCDE...", "start_date": "2026-01-15", "end_date": "2026-01-15" } } ``` --- ## User activity Returns daily, weekly, and monthly active users along with session counts. Start date in `YYYY-MM-DD` format End date in `YYYY-MM-DD` format Set to `client` to group by client type ```json { "data": [ { "date": "2026-01-15", "daily_active_users": 128, "weekly_active_users": 312, "monthly_active_users": 485, "daily_active_users_by_client": { "terminal-ui": 95, "web": 42, "non-interactive-cli": 18 }, "sessions": 890, "messages": 12500, "user_messages": 4200 } ], "meta": { "org_id": "org_01HPMQ6ABCDE...", "start_date": "2026-01-15", "end_date": "2026-01-15" } } ``` Date in `YYYY-MM-DD` format Unique users on this day Unique users in trailing 7 days Unique users in trailing 30 days DAU breakdown by client type Total sessions started Total messages (user + assistant) Messages from users only | Client | Description | | :------------------- | :--------------------------------------- | | `terminal-ui` | Interactive CLI sessions | | `web` | Factory App | | `non-interactive-cli`| Headless/automated CLI (`droid exec`) | When `group_by=client`, returns one row per client type per day inside `data`: ```json { "data": [ { "date": "2026-01-15", "group_key": "terminal-ui", "daily_active_users": 95 }, { "date": "2026-01-15", "group_key": "web", "daily_active_users": 42 } ], "meta": { "org_id": "org_01HPMQ6ABCDE...", "start_date": "2026-01-15", "end_date": "2026-01-15" } } ``` --- ## Productivity Returns daily file operations and git activity. Start date in `YYYY-MM-DD` format End date in `YYYY-MM-DD` format ```json { "data": [ { "date": "2026-01-15", "files_created": 245, "files_edited": 1820, "by_extension": [ { "extension": ".ts", "count": 890 }, { "extension": ".tsx", "count": 420 }, { "extension": ".py", "count": 310 } ], "by_language": [ { "language": "TypeScript", "count": 1310 }, { "language": "Python", "count": 310 } ], "git_commits": 156, "git_prs_created": 42 } ], "meta": { "org_id": "org_01HPMQ6ABCDE...", "start_date": "2026-01-15", "end_date": "2026-01-15" } } ``` Date in `YYYY-MM-DD` format New files created by agent Existing files modified by agent Operations per file extension Operations per programming language Commits made via agent Pull requests created via agent --- ## Per-user metrics Returns detailed metrics per user with cursor-based pagination. Start date in `YYYY-MM-DD` format End date in `YYYY-MM-DD` format Users per page, 1-100 (default: 20) User ID for pagination (from `next_cursor`) ```json { "data": [ { "user_id": "user_01HPMQ7NXKHM7Y7PR3TTZY3JZS", "user_email": "developer@company.com", "date": "2026-01-15", "tool_calls": 1250, "billable_tokens": 450000, "primary_model": "claude-sonnet-4-5-20250929", "primary_model_tier": "standard", "files_created": 12, "files_edited": 85, "git_commits": 8, "git_prs_created": 2, "mcp_calls": 45, "skill_calls": 8, "slash_commands": 22, "hooks": 120, "sessions": 15, "messages": 180, "user_messages": 62, "assistant_messages": 118, "autonomy_ratio": 9.2, "delegation_level": "auto-high", "languages": ["TypeScript", "Python", "Go"] } ], "meta": { "org_id": "org_01HPMQ6ABCDE...", "start_date": "2026-01-15", "end_date": "2026-01-15", "has_more": true, "next_cursor": "user_01HPMQ8ABCDE7Y7PR3TTZY4KLM" } } ``` Unique user identifier User email Date in `YYYY-MM-DD` format Tool invocations by this user Factory Standard Credits consumed by this user Most-used model Model tier (`standard` or `thinking`) Files created Files edited Commits made Pull requests created MCP tool invocations Skill activations Slash command uses Hook executions Sessions started Total messages User messages only Assistant messages Tool calls per user turn Primary autonomy mode Programming languages worked in | Level | Description | | :------------ | :------------------------------------------------- | | `auto-high` | Maximum autonomy, minimal confirmations | | `auto-medium` | Balanced autonomy with some confirmations | | `auto-low` | Limited autonomy, frequent confirmations | | `spec` | Specification mode, planning before execution | | `manual` | Full manual control, confirm each action | Use cursor-based pagination to iterate through users: ```bash # First page curl -H "Authorization: Bearer $FACTORY_API_KEY" \ "https://api.factory.ai/api/v1/analytics/users?startDate=2026-01-15&endDate=2026-01-15&limit=50" # Next page curl -H "Authorization: Bearer $FACTORY_API_KEY" \ "https://api.factory.ai/api/v1/analytics/users?startDate=2026-01-15&endDate=2026-01-15&limit=50&cursor=user_01HPMQ8ABCDE7Y7PR3TTZY4KLM" ``` ## Important constraints ### Date requirements All dates must be `YYYY-MM-DD`. UTC only (no timezone parameter). Data is available through yesterday (UTC). Requesting today's date returns a `400` error. Available from January 14, 2026. The `/cost/me/query` endpoint accepts a maximum range of 90 days. ### Rate limits Rate limits vary by plan. [Contact us](mailto:support@factory.ai) for specifics or if you need higher limits for dashboard or automation use cases. --- ## Errors The API returns standard HTTP status codes: | Status | Description | | :----- | :--------------------------------------------------- | | `400` | Invalid date format, today's date requested, or limit out of range | | `401` | Missing or invalid API key | | `403` | Insufficient role, organization tier, or Analytics access | | `500` | Internal error | ### Error response format ```json { "title": "Bad Request", "detail": "Cannot query today's date - analytics data has a 24-hour lag", "status": 400, "requestId": "req_01HPMQ9WXYZ..." } ``` --- ## Data pipeline Analytics data flows through the following pipeline: ```text CLI/Daemon → OTEL Events → BigQuery (raw) → dbt models → API ``` {/* sweep-allow: term-bullets */} - **Source**: OpenTelemetry spans from the CLI and daemon - **Processing**: Daily batch aggregation via dbt - **Availability**: Data is available the day after it's generated --- ## Data quality notes A few known data quality considerations: {/* sweep-allow: term-bullets */} - **MCP server names**: Some duplicates exist due to case sensitivity (e.g., `axiom` vs `Axiom`) - **Tool names**: Approximately 0.006% of entries contain parsing artifacts - **User counts**: A user active on multiple clients counts once in DAU but appears in each client breakdown --- ## Use cases ### Cost monitoring dashboard Track usage trends and identify cost drivers: ```bash # Daily usage for the month curl -H "Authorization: Bearer $FACTORY_API_KEY" \ "https://api.factory.ai/api/v1/analytics/tokens?startDate=2026-01-14&endDate=2026-01-28" ``` ### Adoption tracking Monitor DAU/WAU/MAU and identify adoption patterns: ```bash # Activity metrics with client breakdown curl -H "Authorization: Bearer $FACTORY_API_KEY" \ "https://api.factory.ai/api/v1/analytics/activity?startDate=2026-01-14&endDate=2026-01-28&group_by=client" ``` ### Team productivity reports Measure output and efficiency: ```bash # Productivity metrics curl -H "Authorization: Bearer $FACTORY_API_KEY" \ "https://api.factory.ai/api/v1/analytics/productivity?startDate=2026-01-14&endDate=2026-01-28" ``` ### Individual performance Export per-user metrics for team leads: ```bash # Paginate through all users curl -H "Authorization: Bearer $FACTORY_API_KEY" \ "https://api.factory.ai/api/v1/analytics/users?startDate=2026-01-15&endDate=2026-01-15&limit=100" ``` # Compliance & Audit Map Droid activity to audit, compliance, and monitoring workflows with Trust Center resources, audit events, and telemetry. Security and compliance teams need clear answers to **who did what, when, where, and with which data**. Droid fits into that posture through Factory's Trust Center, Factory-side audit events for cloud-managed features, customer-owned telemetry exports, and workflow logs in systems such as GitHub Actions. Use this page for audit and compliance posture. For telemetry setup and the hosted Analytics API, see [Telemetry & Analytics](/enterprise/telemetry); for exported metric names, see the [Telemetry Data Reference](/enterprise/telemetry/data-reference). --- ## Certifications and Trust Center Factory maintains an enterprise-grade security and compliance program, including: - **SOC 2 Type II** - **ISO 27001** - **ISO 42001** Our **Trust Center** provides up-to-date reports, security architecture documentation, and sub-processor lists. Use it as the primary reference for security and compliance reviews. --- ## Compliance restrictions Factory can apply organization-specific compliance restrictions for Enterprise customers across: - Managed inference. - Managed compute. - Cloud sync. These restrictions cannot be changed through Enterprise Controls. Adding, removing, or changing a restriction requires a written request to Factory. Factory enforces all restrictions at the **backend request layer**. For every authenticated request, the backend checks the requesting organization's compliance policy and rejects access to restricted capabilities before the capability handles the request. --- ## Audit trails and events There are two complementary sources of audit information: 1. **Factory-side audit logs** (for cloud-managed features). 2. **Customer-side OTEL telemetry** emitted by Droid. ### Factory-side audit logs (cloud-managed) When you use Factory's hosted services, the control plane records key events such as: - Authentication events and SSO/SCIM changes. - Org and project configuration updates. - Policy changes (model allow/deny lists, autonomy limits, Droid Shield settings, hooks configuration). - Administrative actions in the web UI. Use the Trust Center for current compliance materials. For audit-log export or security-system integrations, coordinate the supported path during your enterprise engagement. ### Customer-side OTEL telemetry Droid exports OTEL **metrics by default** as fine-grained activity data inside your own systems. You can also [opt in to customer-only message content trace spans](/enterprise/telemetry/privacy#message-content-logging) for session auditing. Exported metrics include: - Tool and MCP tool invocations, including execution duration. - Commands blocked by policy. - Code modification activity (files read and modified, lines modified). - Git activity (commits and pull requests created) and successful logins. Every data point carries `organization.id`, `session.id`, and `user.id` (under `telemetry.granularity: user`) attributes, so you control how long these signals are retained and how they are correlated with other systems such as CI/CD pipelines, SIEMs, and case management tools. Because the metrics land in collectors you operate, you can enrich or redact attributes against your own policies before forwarding them, keeping **ownership of telemetry in your hands** even when you use Factory's cloud-managed features. Where a works council agreement or a jurisdictional rule forbids per-individual analytics, pin [`telemetry.granularity: aggregate`](/enterprise/telemetry/privacy#data-granularity) in org-managed settings. User identifiers are then stripped at the source on both collectors, and message content is never written, whatever a member sets on their own machine. `organization.id` survives as the aggregation dimension, so activity and policy signals stay usable for compliance reporting without identifying individuals. For aggregated usage, token, and cost data, use Factory's hosted [Analytics API](/api-reference/analytics), which is backed by Factory's internal tracing. The full exported-metric catalog is owned by the [Telemetry Data Reference](/enterprise/telemetry/data-reference); OTLP collector configuration and provider examples by [Telemetry & Analytics](/enterprise/telemetry). This page covers only how that data supports audit and compliance workflows. ## Securing the GitHub Action and CI The Factory Droid GitHub Action ([`Factory-AI/droid-action`](https://github.com/Factory-AI/droid-action)) runs **entirely inside GitHub Actions** on runners you control. It does not provision external compute, code is checked out transiently for the workflow run and discarded afterward, and authentication uses your Factory API key subject to your org's model allowlists, rate limits, and policies. ### Authentication and authorization {/* sweep-allow: term-bullets */} - **Factory API key.** Store `FACTORY_API_KEY` as a GitHub Actions secret; never commit it. The key authenticates Droid Exec sessions and is subject to your org policies. - **GitHub App tokens.** When using the Factory Droid GitHub App, the app requests a short-lived installation token scoped to the repository, automatically revoked after the workflow completes. - **User permission verification.** Before executing any `@droid` command, the action verifies the triggering user has **write access** to the repository. Bots are rejected unless explicitly allowed via the `allowed_bots` input. ### Permission scoping The generated workflow requests only the GitHub permissions the action needs: ```yaml permissions: contents: write # Read code, write for fixes pull-requests: write # Comment on and update PRs issues: write # Comment on issues id-token: write # OIDC token for secure auth actions: read # Read workflow run metadata ``` You can further restrict permissions in your workflow file based on your security requirements. ### Wired security inputs | Input | Purpose | | --------------------------- | -------------------------------------------------------------------------------------------------- | | `allowed_bots` | Comma-separated list of bot usernames allowed to trigger, or `*` for all. Default: none. | | `automatic_security_review` | Run security review automatically on non-draft PRs without requiring a manual `@droid` mention. | | `security_model` | Override the model used for security review candidate generation and full-repository scans. | For the full security review workflow, modes, and methodology, see [Security Review](/software-factory/security-review). ### Audit and monitoring for CI - **Workflow logs.** All Droid activity is recorded in GitHub Actions workflow runs, with timestamps, step status, and links to comments or changes made. - **Factory telemetry.** If your org exports OTEL metrics, Droid Exec sessions from GitHub Actions are included in your metrics, tagged with `session.id` and tool/activity attributes. ### Deployment recommendations #### For security-conscious organizations 1. **Use repository or org secrets** for `FACTORY_API_KEY`. 2. **Review workflow permissions** so the workflow requests only what it needs. 3. **Restrict bot access** - keep `allowed_bots` empty unless you have a specific need. 4. **Enable branch protection** to require PR reviews before merging Droid-assisted changes. 5. **Monitor workflow runs** in your GitHub Actions logs regularly. #### For regulated environments {/* sweep-allow: term-bullets */} - **Self-hosted runners** in your controlled environment. - **Model allowlists** via Factory org policies to restrict which models Droid can use. - **Integrate with SIEM** by exporting GitHub Actions logs and Factory telemetry to your security monitoring tools. ## Regulatory and industry use cases Factory is designed to support organizations operating under strict regulatory regimes. While implementation details differ, common patterns include: | Use case | Common pattern | | --- | --- | | Financial services | Use hybrid or airgapped deployments for systems subject to strict data residency and record-keeping requirements. Route LLM traffic through gateways that implement your bank's data policies, and use OTEL telemetry plus hooks so Droid activity is visible in your SIEM. | | Healthcare and PHI | Deploy Droid in environments that never expose protected health information to external LLMs. Use model allowlists that include only providers and gateways that meet your PHI handling requirements, and use approved gateways, DLP hooks, and Droid Shield to reduce exposure. | | National security and defense | Rely on fully airgapped deployments with on-prem models and collectors. Treat Droid as an internal tool whose artifacts and logs never leave your network, and integrate OTEL plus hooks with mission-specific monitoring and incident response tooling. | --- ## Deployment and configuration for compliance teams To integrate Droid into your compliance and monitoring stack: 1. **Decide on deployment pattern**: cloud-managed, hybrid, or fully airgapped. 2. **Define model and gateway policies**: which providers and gateways are allowed, and where. 3. **Configure OTEL collectors and destinations**: ensure all Droid telemetry flows into your SIEM and observability tools. 4. **Set up hooks and Droid Shield**: enforce DLP, approval workflows, and environment-specific controls. 5. **Document policies and mappings**: connect Droid controls to your internal control framework and regulatory obligations. Most of this configuration is expressed through [Enterprise Controls & Managed Settings](/enterprise/hierarchical-settings-and-org-control). Set up OTEL export and read the hosted Analytics API for audit data. Choose hybrid or airgapped patterns for regulated environments. # Audit Log Organization-level audit events that record who did what in your Factory organization. The Factory audit log records organization-level events so you can track who did what in your organization. Each event captures **who** acted, **where** the action originated, **what** was affected, and structured details about the change. **Enterprise Feature** -- The audit log is available to **Enterprise organization owners only**. Owners can view the audit log in the Factory web app under Team Settings, or via the public API. --- ## Overview Every audit event includes: - **Actor** -- the user or service principal who initiated the action (or `null` for system-initiated events). - **Source** -- the surface that originated the event (e.g. web settings pages, public API, identity provider webhooks). - **Target** -- identifiers for the affected entity (e.g. user, service account, integration). - **Payload** -- structured, event-specific details. Contains only IDs and enum values, never secrets. - **Timestamp** -- ISO 8601 timestamp of the event. --- ## Event categories Audit events cover the following categories of organization activity: - **Usage limits** -- per-user token usage limit changes. - **Membership** -- role changes, invitations, and member removals. - **API keys** -- creation and deletion of user and service-account API keys. - **Service accounts** -- creation, modification, deletion, and credential/grant changes. - **Integrations** -- connection, disconnection, configuration, and availability toggles. - **Organization lifecycle** -- creation and deactivation of organizations. - **Managed settings** -- updates to org-managed settings (with revision tracking for before/after diffing). - **Analytics settings** -- updates to org-wide analytics preferences. See how audit events fit into your broader audit, compliance, and monitoring posture. Export OTEL telemetry and read the hosted Analytics API for fine-grained activity data. # Factory API Public API for Factory platform. Requires authentication via the `Authorization: Bearer` header. ## Authentication The Factory Public API authenticates every request with a bearer token. Pass your Factory API key in the `Authorization` header as a `Bearer` token. Requests without a valid API key receive a `401` response. Factory API key or JWT token for authentication ```bash Authorization: Bearer $FACTORY_API_KEY ``` ## Create an API key Create API keys in the Factory App before you call the API. Keys are shown once, so copy the value immediately and store it in your secret manager or shell environment. 1. Open [Factory API keys settings](https://app.factory.ai/settings/api-keys). 2. Create a new API key and give it a name that matches the integration. 3. Copy the key, store it securely, and load it as an environment variable. ```bash export FACTORY_API_KEY="fk-your-api-key" ``` ## Base URL All endpoints are relative to the production base URL. Prepend it to each operation path shown on the group pages. ```text https://api.factory.ai ``` API version 0.1.0. ## Resource groups The API is organized into 6 resource groups. Each group page lists its operations with parameters, request and response schemas, and a curl example. - [CI Automations API](https://docs.factory.ai/api-reference/automations-ci.md): Manage CI automation workflows: scan repositories, list jobs, and open workflow PRs. (6 endpoints) - [Droid Computers API](https://docs.factory.ai/api-reference/computers.md): Provision and manage Factory computers: create, inspect, refresh, restart, and delete. (15 endpoints) - [Organization Management API](https://docs.factory.ai/api-reference/organization.md): Administer organization members and usage limits, and review enterprise control history. (8 endpoints) - [Service Accounts API](https://docs.factory.ai/api-reference/service-accounts.md): Manage service accounts and their API keys: create, list, update, rotate, and revoke. (9 endpoints) - [Droid Sessions API](https://docs.factory.ai/api-reference/sessions.md): Create and drive Droid sessions: manage their lifecycle, settings, and messages. (9 endpoints) - [AutoWiki API](https://docs.factory.ai/api-reference/wiki.md): Generate and manage AutoWiki runs: create runs, browse pages, search, and export. (10 endpoints) # Full Changelog Every notable change across the Factory App, Droid CLI, and other product surfaces, newest first. ## CLI v0.209.0 · Desktop v0.166.0 (September 1, 2026) Session archiving from chat, a new default model, faster startup, and steadier app views ### New features - **Archive sessions from chat** - Archive the current session with /archive or Ctrl+X, and restore sessions from the Archived tab ### Improvements - **New default model** - Sessions with no model selected now default to GPT-5.6 Sol - **Faster startup with organization settings** - The CLI now starts on time when organization settings load slowly, applying them once they arrive ### Bug fixes - **Diff text selection** - Text selected in a diff no longer clears when the session sidebar updates (app) - **Environment panel visibility** - The environment panel now stays visible while repository details load or refresh (app) - **File link line targets** - Opening a file link now centers the target line in the file view (app) - **Droid Core model suggestion** - Running out of credits now suggests an available Droid Core model instead of a retired one ## CLI v0.208.1 · Desktop v0.165.1 (August 29, 2026) 60% off GPT and Grok models ### Improvements - **GPT and Grok discount** - GPT and Grok models now bill at 60% off for a limited time, with the discount shown in the model selector ## CLI v0.208.0 · Desktop v0.165.0 (August 29, 2026) Sharper image reading, resilient MCP servers, faster slash commands, Windows hook fixes, and verified domain invitations ### New features - **Verified domain invitations** - Organization owners can now require people on verified email domains to join by invitation instead of creating a new organization ### Improvements - **Sharper image reading** - Screenshots, diagrams, and tables now keep fine detail when Droid resizes images - **New session progress** - The /new and /clear commands now show progress while your next session is created - **You.com MCP server** - You.com is now available in the MCP server registry for web search and research - **Faster slash commands** - Built-in slash commands now run immediately instead of waiting for command discovery to finish ### Bug fixes - **Mission milestone validation** - Mission milestones now run their validation steps even after a session is interrupted, paused, or resumed - **Windows plugin hooks** - Plugin hooks now run on Windows when their command path contains spaces - **Unresponsive MCP servers** - Unresponsive MCP servers no longer stall the CLI, and a notice names the server that failed - **Airgapped session warnings** - Airgapped sessions no longer log blocked connection warnings, and /bug explains it cannot send reports - **Exit command** - Typing /exit now quits the CLI instead of running a different suggested command - **Local session loading** - Opening a local session from the sidebar now shows a loading state while its details load (app) - **Desktop integration sign-in** - Connecting an integration from Desktop now uses the organization you selected (app) ## CLI v0.206.0 · Desktop v0.163.0 (August 28, 2026) Faster CLI startup, reliable session controls, project guidance across repositories, and resilient commands ### Improvements - **Cross-repository guidance** - Droid now applies project guidance when searching or listing files outside the current repository - **Faster CLI startup** - The CLI now accepts input sooner by postponing work until it is needed ### Bug fixes - **Session reconnections** - Long-running sessions now continue cleanly when Droid needs to reconnect - **Responsive session controls** - Session archiving, renaming, and selection no longer hang when favorites fail to refresh - **First message cancellation** - You can now cancel your first message immediately after sending it - **Custom model context limits** - Droid now recognizes more custom model context limit errors and recovers normally - **Sleeping laptop commands** - Running commands no longer time out solely because your laptop went to sleep - **Connected tool cancellation** - Interrupting a turn now cancels active connected tool calls instead of leaving the session stuck ## CLI v0.205.0 · Desktop v0.162.0 (August 26, 2026) Account switching, unread sessions, smoother updates, and clearer file handling ### New features - **Mark sessions unread** - You can now mark a session as unread from its sidebar menu for later review (app) ### Improvements - **Smoother background updates** - Background updates now keep the CLI responsive while you type ### Bug fixes - **VS Code windows** - Opening a project in VS Code now creates a new window and handles folder paths correctly (app) - **Smoother account switching** - Switching organizations now starts a fresh session, while logout returns directly to the login screen - **Local session creation** - Local sessions can now start without waiting for organization defaults to load (app) - **Pending session requests** - Pending questions and approvals now remain available when a session reconnects (app) - **Clearer file failures** - Failed file changes now display a clear error state without misleading diff details (app) ## CLI v0.204.0 · Desktop v0.161.0 (August 25, 2026) Droid diagnostics, quicker one-shot commands, reliable MCP connections, steadier session resumes, and clean response completion ### New features - **Droid diagnostics** - Use Droid doctor to diagnose connectivity and local setup issues with focused or JSON output ### Improvements - **Quicker one-shot commands** - One-shot commands now start without waiting for unrelated connections unless the command needs them ### Bug fixes - **MCP connections** - MCP connections now work reliably across the CLI and app without repeatedly asking you to reconnect - **Resumed session state** - Resuming a session now restores pending questions and approvals without stale working states or duplicate replay - **Completed responses** - Responses now finish cleanly instead of hanging after Droid has completed its answer ## CLI v0.203.0 · Desktop v0.160.0 (August 25, 2026) Queued messages are now editable, skill commands work after other text, and session switching, voice dictation, and large attachments are steadier ### New features - **Edit queued messages** - You can now move a queued or pending steering message back into the composer to change its text or attachments before it is sent (app) ### Improvements - **Retention change warnings** - Turning on or shortening session retention now warns you that existing sessions past the limit are permanently deleted before you save ### Bug fixes - **Inline skill commands** - You can now start a skill command after other text in your request, and text that begins with an unmatched slash submits normally instead of being blocked - **Session switching** - Switching between sessions no longer flashes, drops focus in the composer, or scrolls the page unexpectedly (app) - **Session reload on focus** - A session now reloads when you focus the window again, so it recovers after the app has been minimized (app) - **Voice dictation at usage limits** - Voice dictation now stops cleanly when you reach a usage limit instead of continuing to record audio it cannot transcribe (app) - **Large document attachments** - Attaching a very large spreadsheet or document no longer stalls the session on every turn (app) ## CLI v0.202.0 · Desktop v0.159.0 (August 21, 2026) The CLI starts faster, admins can turn image generation off, and a batch of app fixes land for sessions, billing, and markdown tables ### New features - **Image generation control** - Org admins can now turn off image generation for everyone in the organization - **Voucher redemption** - You can now redeem a voucher from your billing settings, including on an annual plan (app) ### Improvements - **Faster CLI startup** - The CLI now starts up faster - **Markdown table rendering** - Markdown tables in a session now render more clearly (app) - **Message actions on hover** - The actions for a message now appear when you hover over it (app) - **Generated document downloads** - Documents Droid generates now finish as files you can download (app) - **Apply changes tooltip** - A tooltip now explains why Apply changes is unavailable (app) ### Bug fixes - **Lifecycle hooks** - Lifecycle hooks now run again instead of being skipped - **Local plugin caches** - A damaged local plugin cache now repairs itself instead of leaving your plugins unavailable - **Chinese translations** - Simplified Chinese text in the interface is now translated correctly - **Unarchived sessions** - Your sessions now load right after you unarchive one (app) - **Session state across pages** - Your open session is kept when you move to another page in the app and come back (app) - **Overage preference updates** - Your overage preference now shows the new setting right after you change it (app) - **Self-hosted Git connections** - You can again start a self-hosted Git provider connection from Source Control settings (app) - **Stray startup notification** - The app no longer shows a startup notification when it connects to your local machine (app) - **Complete org member list** - The org members list now shows every member of your organization ## CLI v0.200.0 · Desktop v0.157.0 (August 20, 2026) Background processes report real process IDs, MCP connections stay authenticated across reconnects, and the model list is searchable in the app ### Improvements - **Failed pull confirmation** - When a Git pull fails, Droid now asks before changing your working tree to fix it instead of doing so on its own - **Faster session startup** - File search tooling now loads only when it is needed, so a session is ready to use sooner - **Model list search** - The model list now has a search box, so you can find a model without scrolling through the whole list (app) ### Bug fixes - **Background process IDs** - Starting a command in the background now reports its real process ID, so you can check on it or stop it - **MCP server reauthentication** - Connections to MCP servers now stay authenticated instead of asking you to sign in again on the next reconnect - **Inherited subagent reasoning** - A subagent can now inherit the session's reasoning setting instead of being forced onto a fixed one (app) ## CLI v0.199.0 · Desktop v0.156.0 (August 18, 2026) The explorer droid can use every read-only tool, agent questions are easier to answer from the keyboard, and diagrams and long sessions render better in the app ### New features - **SDK custom system prompts** - The TypeScript SDK can now set a custom system prompt for a session, and it is kept when the session resumes - **Managed BYOK model defaults** - Org admins can now pick BYOK custom models as managed defaults for sessions and subagents in Enterprise Controls - **Service account model policies** - Service accounts now appear in Enterprise Controls model policies, so admins can set which models each one may use ### Improvements - **Explorer tool access** - The built-in explorer droid can now use every read-only tool, including command execution, web search, and MCP tools, while file edits stay blocked - **Question prompt keyboard controls** - You can now pick answers with Space or the number keys, move between questions with Tab, and go back to edit an answer you already gave - **Mermaid diagram rendering** - Diagrams in a session now render with the official Mermaid engine, so more diagram types display correctly (app) - **Faster session rendering** - Long sessions and streaming command output now render with less lag, so scrolling and typing stay responsive while the agent runs (app) - **Mission model settings** - You can now choose the models and reasoning defaults a mission uses from its settings (app) ### Bug fixes - **Agent question context** - The context the agent includes with a question now stays visible while you answer it ## CLI v0.198.0 · Desktop v0.155.0 (August 17, 2026) Diff lines wrap by default, plus fixes for diff search, the model selector, empty sessions, and mission models ### Improvements - **Diff line wrapping** - Long lines in a diff now wrap by default so you can read them without scrolling sideways (app) - **Clearer empty response error** - When a model finishes without producing an answer, the message now explains why and suggests retrying or switching models ### Bug fixes - **Diff search navigation** - Stepping to the next or previous diff search match now scrolls to it instead of stopping when the match is still loading (app) - **Model selector search** - The model selector keeps focus in its search box so your typing is no longer lost while filtering models (app) - **Chat activity while thinking** - The chat keeps showing the running tool while the agent is thinking instead of clearing it (app) - **Empty draft sessions** - Sessions that were opened but never used no longer pile up in your session list and resume picker - **Missions after sleep** - A mission now moves a stuck worker onto a fresh one after your machine sleeps instead of hanging - **Mission model defaults** - A mission now starts with the mission models from your defaults, and every connected client shows the models the run is actually using ## CLI v0.197.0 · Desktop v0.154.0 (August 15, 2026) More accurate voice dictation and steadier option navigation in prompts ### Improvements - **Voice dictation accuracy** - Voice dictation now transcribes what you say more accurately (app) ### Bug fixes - **Prompt option navigation** - Moving through the options in a question or permission prompt no longer jumps to whichever option the mouse is resting on (app) ## CLI v0.196.0 · Desktop v0.153.0 (August 14, 2026) Forking now keeps you in your current session and credit usage moves to its own page, plus fixes for the pull request indicator, SessionStart hook context, and credit rates for two models ### Improvements - **Fork stays in place** - Forking a session now keeps you in the original session and prints a resume command you can use to open the copy - **Credit usage page** - Credit usage now has its own page instead of living inside Analytics (app) ### Bug fixes - **Pull request indicator** - The linked pull request indicator stays visible when a status lookup briefly fails instead of disappearing - **SessionStart hook context** - Context added by a SessionStart hook is kept for the session instead of being dropped - **Model credit rates** - Credit usage for Grok 4.5 and GPT-5.3-Codex fast mode is now reported at the correct rate ## CLI v0.195.0 · Desktop v0.152.0 (August 13, 2026) Escape now confirms before interrupting a run, plus fixes for sidebar width, text selection after resizing, recent sessions in the sidebar, repeated polling tool calls, and plugin file watchers ### Improvements - **Escape interrupt confirmation** - Pressing Escape while the agent is running now asks whether to stop it, and remembers your choice (app) ### Bug fixes - **Stable sidebar width** - The sidebar keeps the width you set when you resize the window or collapse and reopen it (app) - **Text selection after resizing** - Text stays selectable and clicks keep working after a panel resize is interrupted (app) - **Recent sessions in the sidebar** - Your most recent sessions stay visible in the sidebar instead of being crowded out by sessions with many child sessions (app) - **Polling tool calls** - A turn no longer stops early when the same tool call is repeated to watch output that keeps changing - **Plugin file watching** - Droid no longer uses up your system's file watchers when you have plugin marketplaces installed, and cleans up files left behind by an interrupted install ## CLI v0.194.0 · Desktop v0.151.0 (August 12, 2026) Session message access and a specification mode helper in the TypeScript SDK, plus fixes for pinned mission proposals, custom model credential errors, copied line breaks, and the idle todo spinner ### New features - **Session messages in the SDK** - Sessions returned by the TypeScript SDK session list can now fetch their own messages directly - **Specification mode SDK helper** - The TypeScript SDK can now take a session out of specification mode with a single call ### Bug fixes - **Pinned mission proposals** - A pinned mission proposal now keeps its heading and full text instead of rendering garbled - **Custom model credential errors** - A custom model whose endpoint has no working credentials now reports the problem right away instead of retrying the same request several times - **Copied text line breaks** - Copying text out of a session now keeps its line breaks instead of collapsing everything onto one line (app) - **Idle todo spinner** - The todo list no longer shows a spinner when nothing is running (app) ## CLI v0.193.0 · Desktop v0.150.0 (August 11, 2026) Faster session search, fixes for repeated tool calls, lowercase proxy variables, the credits usage chart, and diff loading, plus the MiniMax M2.5 retirement ### Improvements - **Faster session search** - Searching your session history now returns results faster ### Bug fixes - **Repeated tool call recovery** - A turn now recovers when the same tool call repeats instead of getting stuck - **Lowercase proxy variables** - Droid now honors lowercase proxy environment variables such as https_proxy in addition to the uppercase ones - **Credit usage chart dates** - Daily totals in the credits usage chart now land on the correct day (app) - **Diff loading retries** - A diff that fails to load no longer retries in a loop (app) ### Deprecations - **MiniMax M2.5 retirement** - MiniMax M2.5 is no longer available for selection ## CLI v0.190.0 · Desktop v0.147.0 (August 7, 2026) Claude Opus 5 in Auto model, HTTPS proxy support, and GitLab CI setup for automated QA, plus fixes for GitLab authentication, empty responses, and disconnected sessions ### New features - **QA setup for GitLab CI** - The install-qa skill can now set up automated QA on GitLab CI and include video evidence in its reports ### Improvements - **Opus 5 in Auto model** - Auto model now upgrades to Claude Opus 5 when a task needs a stronger model - **HTTPS proxy support** - Droid now uses your HTTPS_PROXY setting when it connects, so sessions work behind a corporate proxy - **Clearer connection errors** - A session that cannot connect at startup now explains what went wrong instead of showing a generic failure ### Bug fixes - **Empty response recovery** - A turn now recovers when the model returns an empty response instead of ending with an error - **GitLab authentication per host** - GitLab commands now authenticate against the host you are working with instead of a single default - **Usage dashboard filters** - The usage dashboard filter bar now wraps so every filter stays visible (app) - **Disconnected session state** - A session that loses its connection now keeps its state instead of resetting (app) - **Archiving a running session** - Archiving a session that is still running now offers to stop it first instead of leaving it running (app) ## CLI v0.189.0 · Desktop v0.146.0 (August 6, 2026) Session filtering in the command palette, faster session default changes, and fixes for status line failures, custom model key helpers, and empty responses ### New features - **Command palette session filter** - The command palette now filters your sessions as you type instead of showing a fixed list (app) ### Improvements - **Faster session default changes** - Changing a session default in settings now takes effect immediately instead of pausing while it saves ### Bug fixes - **Status line failures** - A custom status line that fails now names the problem and points you to `/statusline` instead of leaving a blank row - **Custom model key helpers** - Custom models that get their key from a helper command now start up instead of failing without a static API key - **Empty final responses** - A turn that ends with an empty response now finishes successfully instead of failing with a no-output error - **Openable file path links** - File paths in a response are only shown as links when they can actually be opened, and the tooltip says where they will open (app) - **Recent session ordering** - Sending a message moves that session back to the top of your recent sessions right away (app) ## CLI v0.188.0 · Desktop v0.145.0 (August 4, 2026) Fixes for plugin marketplace status, authentication for manually added MCP servers, diff colors, the thinking timer, member pickers, and opening markdown files ### Bug fixes - **Marketplace status reporting** - A plugin marketplace that is working is no longer reported as missing - **Manually added MCP servers** - MCP servers you add by hand can now finish authenticating instead of failing to connect (app) - **Thinking timer stability** - The thinking timer keeps counting correctly when you move between views instead of restarting (app) - **Diff colors by theme** - Added and removed lines in a diff now use the colors that match your light or dark theme (app) - **Paginated member pickers** - Member pickers now page through everyone in your organization instead of showing only the first results (app) - **Markdown in default app** - Markdown files now open in your default app instead of being blocked (app) ## CLI v0.187.0 · Desktop v0.144.0 (August 4, 2026) Usage limits filtered by limit type, plus fixes for PDF attachments, plugin marketplace updates, and file diff counts ### Improvements - **Usage limits by type** - Usage limits can now be filtered by limit type so you can narrow the list to the one you are looking for (app) - **Voucher redemption confirmation** - Redeeming a voucher that changes your plan now shows what will change and asks you to confirm first (app) ### Bug fixes - **PDF attachment fallback** - Attaching a PDF now falls back to its extracted text when the selected model cannot read PDFs, instead of failing - **Skills with broken symlinks** - A skill with an unresolved symlink is now skipped instead of stopping your other skills from loading - **Marketplace updates with local changes** - Updating a plugin marketplace now succeeds when its checkout has local changes instead of failing - **Inline file diff counts** - File diffs show their added and removed line counts again (app) - **Updates after canceled quit** - The app keeps checking for updates after you cancel quitting instead of waiting until the next launch (app) ## CLI v0.186.0 · Desktop v0.143.0 (July 31, 2026) Support for GitHub Enterprise and self-hosted GitLab remotes, clickable links in tmux, and fixes for transcript performance, command denylists, and table scrollbars ### New features - **Self-hosted Git remotes** - Repositories hosted on GitHub Enterprise or a self-hosted GitLab are now recognized as Git remotes ### Improvements - **Streaming transcript performance** - Long streaming responses no longer slow the transcript down while they render (app) - **Cmd-click hint on file paths** - File path links now show a tooltip explaining that Cmd-click opens the file (app) ### Bug fixes - **Clickable links in tmux** - Links in the terminal are now clickable when you run the CLI inside tmux - **Command denylist matching** - A denylisted command no longer blocks unrelated commands that happen to share a subcommand name - **Table scrollbar overlap** - Scrollbars in tables no longer sit on top of row content (app) ## CLI v0.185.0 · Desktop v0.142.0 (July 30, 2026) A per-server MCP connect timeout, plus fixes for stuck approval prompts, spec mode colors, and sign-in across multiple windows ### New features - **Per-server MCP connect timeout** - MCP server configs accept a `connectTimeout` so a slow-starting server gets more time to connect instead of failing at startup ### Bug fixes - **Stuck approval prompts** - An approval prompt no longer stays wedged when the request behind it is already gone, so Enter and Esc respond again - **Spec mode message color** - Messages you send in spec mode keep the spec color when the view repaints or you resume the session, instead of looking like a normal turn - **Sign-in with multiple windows** - Sign-in and MCP authorization now finish in the window that started them instead of hanging when more than one window is open (app) - **File diff animations** - Opening and closing a file diff animates again instead of snapping into place (app) ## CLI v0.183.0 · Desktop v0.140.0 (July 29, 2026) Multi-select questions from the agent, a fullscreen file preview, and a round of app improvements and fixes ### New features - **Multi-select questions** - Questions the agent asks you can now accept more than one answer, so you can pick several options at once - **Fullscreen file preview** - The file preview panel can now be expanded to fullscreen and back with a toggle (app) ### Improvements - **MCP tool details** - Tool calls from MCP servers now show their details in the transcript instead of hiding them (app) - **Queued attachment previews** - Attachments on a queued message now show a preview before the message is sent (app) ### Bug fixes - **Git connection errors** - Errors while connecting a Git provider are now shown in the app instead of failing without explanation (app) - **Pinned session dragging** - Dragging a pinned session no longer unpins it when another target handles the drop (app) - **Show in folder on WSL** - Revealing a file in its folder now works for paths inside WSL (app) ## CLI v0.182.0 · Desktop v0.139.0 (July 28, 2026) Audit log scope filter, user removal in the Admin Console, and a round of CLI session and authentication fixes ### New features - **Audit log scope filter** - The audit log can now be filtered by scope, so you can narrow it to the events you are looking for (app) - **Admin Console user removal** - Admins can now remove a user from the organization directly in the Admin Console (app) - **Image copy context menu** - Images in a session can be copied to the clipboard from a right-click menu (app) ### Improvements - **Recurring prompts in Loop** - Creating, listing, and cancelling recurring prompts is now one Loop tool instead of separate cron tools - **Image thumbnails in previews** - Artifact preview cards in chat now show a thumbnail of the image instead of a generic card (app) ### Bug fixes - **Forked session sync** - Sessions you fork now sync to the cloud instead of staying only on the machine that forked them - **Expired session errors** - The CLI now reports an authentication error when the stored session token has gone missing - **Re-authentication default** - Re-authenticating is now the preselected option when your session has expired - **BYOK model names** - Bring-your-own-key models no longer show an extra suffix after the model name - **Session and mission pickers** - Typing `q` in the session or mission picker now filters the list instead of closing it - **Duplicate AGENTS.md** - AGENTS.md files are no longer added to context twice when discovery runs more than once - **Header state** - The header now follows tool execution as it happens instead of lagging behind the agent - **Transcript order** - Session transcripts stay in the right order when the system clock jumps - **Reconnect after slow startup** - A local session that takes a long time to start now reconnects instead of staying disconnected (app) - **Side Chat focus** - Running `/btw` now moves focus into Side Chat (app) - **Repository access during setup** - Connecting a Git provider no longer fails when a single repository in the account cannot be read (app) ## CLI v0.181.0 · Desktop v0.138.0 (July 27, 2026) Slash commands while the agent runs, MCP OAuth resource override, and subagent and Git fixes ### New features - **Slash commands while the agent runs** - Slash commands now work during an active agent turn, so you no longer have to wait for it to finish - **MCP OAuth resource override** - MCP server configs accept a `resource` field to send a different OAuth resource indicator, or omit it entirely ### Improvements - **Subagent delegation** - The agent now delegates to subagents more selectively when using OpenAI models - **Subagent results** - The agent reuses a subagent's report instead of re-reading the files it already covered, making follow-up work faster ### Bug fixes - **Git index lock** - Checkouts now retry when another process is holding the Git index lock instead of failing - **BYOK models** - Bring-your-own-key models are now scoped to the selected computer (app) ## CLI v0.180.0 · Desktop v0.137.0 (July 24, 2026) Gemini 3.6 Flash support, a revamped keyboard shortcuts dialog, and faster session search ### New features - **Gemini 3.6 Flash support** - Gemini 3.6 Flash is now available in the model selector - **Disable built-in skills in the app** - You can now turn off Factory-provided built-in skills from the app (app) ### Improvements - **Keyboard shortcuts help** - Redesigned the keyboard shortcuts help dialog so shortcuts are easier to scan (app) - **MCP readiness** - Agent turns now wait until your MCP servers finish loading, so their tools are available from the first message - **Session search** - Session search, sorting, and paging now run server-side for faster results in long histories, and you can look up a session by its ID (app) - **MCP sign-in** - MCP authentication in the app now matches the CLI flow (app) ### Bug fixes - **Spec proposal order** - Spec proposals now stay in chronological order in the chat transcript - **Image loading** - Fixed images that failed to load in the desktop app, and cards now hide broken images instead of showing a placeholder (app) ## CLI v0.179.0 · Desktop v0.136.0 (July 23, 2026) Sandbox and macOS hardening, plus mission, dictation, and app UI fixes ### Improvements - **Sandbox hardening** - Strengthened isolation of processes running inside the command sandbox - **macOS hardening** - Hardened macOS releases against code injection ### Bug fixes - **Mission approvals** - The mission proposal body now renders once during approval instead of repeating - **Capacity errors** - The CLI now backs off gracefully when a model is temporarily at capacity and no alternative is available - **Voice dictation** - Long voice dictations are no longer cut off (app) - **Visualization flicker** - Streamed visualizations no longer flicker while rendering (app) - **Autonomy chip** - The autonomy chip in the composer now keeps a fixed width (app) - **GitHub Enterprise Server repositories** - Repository lists are now correctly scoped to your organization (app) ## CLI v0.178.0 · Desktop v0.135.0 (July 22, 2026) Keep-awake setting, system certificate trust, plus MCP, model, and invite fixes ### New features - **Keep-awake setting** - A new setting keeps your machine awake during agent runs so long tasks aren't interrupted by sleep (app) ### Improvements - **System certificate trust** - The CLI and desktop now trust your system's certificates at startup, improving connectivity in environments with custom certificate authorities ### Bug fixes - **MCP confirmations** - Autonomy upgrades now apply correctly to MCP tool confirmations - **MCP OAuth** - Signing in to MCP servers over OAuth is now more reliable - **Hook display** - Hook metadata now stays on a single row - **Verified domains on invite** - Invitations now respect your organization's verified domains (app) ## CLI v0.177.0 · Desktop v0.134.0 (July 21, 2026) Open artifacts in your default app, plus mission, transcript, and app UI fixes ### New features - **Open artifacts externally** - Artifact preview cards can now open in your default app (app) ### Improvements - **Code block font** - Code blocks now render in Geist Mono for clearer monospaced text (app) - **Attachment-only messages** - You can now send chat messages that contain only attachments (app) ### Bug fixes - **Approval details toggle** - Option+E dead-key accents no longer interfere with toggling approval details - **Transcript repaint** - The transcript now repaints correctly after viewing plan details - **Computer removal** - Removing a computer no longer shows a spurious session-archiving warning - **Mission interrupts** - Missions now pause correctly when the orchestrator is interrupted - **Agent browser cleanup** - Orphaned browser automation processes are now cleaned up reliably - **Unavailable sessions** - Sessions that can't be loaded now show a clear unavailable state (app) - **Skill colors** - Loaded skills now display with the correct colors (app) ## CLI v0.176.0 · Desktop v0.133.0 (July 20, 2026) Windows ARM64 desktop, disable builtin skills, plus session recovery and Auto model fixes ### New features - **Disable builtin skills** - New `--disable-builtin-skills` flag turns off Factory-provided skills while keeping your own skill sources - **Windows ARM64 desktop** - The desktop app now ships a native Windows ARM64 build (app) - **Usage cap alert emails** - Admins now receive email alerts as usage approaches configured caps (app) ### Bug fixes - **Session recovery** - Cancelling a tool no longer corrupts session history, and existing corruption is repaired automatically - **Auto model retries** - Auto model turns now retry automatically when a model is unavailable - **Disconnected prompts** - Permission and input prompts are now suppressed while disconnected - **macOS sleep** - Notification sounds no longer prevent macOS from going to sleep (app) - **PDF generation** - PDF generation errors now surface immediately instead of hanging (app) ## CLI v0.175.0 · Desktop v0.132.0 (July 17, 2026) Auto-opening file previews, task ID prefixes, and /btw, model settings, wiki, and plugin fixes ### New features - **File previews** - Generated files now open in a preview automatically (app) ### Improvements - **Task ID prefixes** - `TaskOutput` and `TaskStop` now accept task ID prefixes so you can reference background tasks more easily ### Bug fixes - **/btw mid-stream** - `/btw` now handles mid-stream forks correctly and submits follow-ups - **Model settings** - Spec and normal mode model settings now persist correctly - **Wiki shareable links** - Wiki runs can now generate shareable links - **Plugins overlay** - The plugins overlay now closes on Esc from any tab ## CLI v0.173.0 · Desktop v0.130.0 (July 15, 2026) Skill disable controls plus loop, plugin, hook, and settings fixes ### New features - **Skill controls** - You can now disable individual skills and see how skill settings resolve across scopes ### Bug fixes - **Repeated tool calls** - The agent now stops looping when it repeats the same tool call - **Plugin loading** - Enabled plugins now load correctly across all install scopes - **Hook execution** - Hooks now run exactly once per event - **Settings writes** - Concurrent settings updates are no longer overwritten by stale cached writes - **Mission Control** - The chat now repaints correctly when you close Mission Control (app) - **Sidebar update button** - Fixed alignment of the sidebar update button (app) ## CLI v0.172.0 · Desktop v0.129.0 (July 14, 2026) Mission side panel, add project folder, and PDF, MCP auth, and markdown fixes ### New features - **Mission side panel** - Added a unified Mission side panel for managing missions (app) - **Add project folder** - New button to add a local project folder, streamlining local project setup (app) ### Improvements - **Consistent text editing** - Text-editing keys now behave consistently across all input fields ### Bug fixes - **`/clear` preserves model** - `/clear` now keeps your current model and autonomy settings - **Password-protected PDFs** - Password-protected PDFs no longer lock up the session - **MCP authentication** - MCP sign-ins now refresh tokens silently and avoid duplicate authentication prompts - **Markdown line breaks** - Soft line breaks in markdown now render correctly ## CLI v0.171.0 · Desktop v0.128.0 (July 13, 2026) MCP connect timeouts, theme detection over SSH, and mission, model, and notification fixes ### Improvements - **Context usage on resume** - Context usage now displays immediately when you resume a session - **Windows installer architecture** - The CLI installer now detects the native Windows architecture ### Bug fixes - **MCP connection timeout** - stdio MCP servers now time out on connect instead of hanging - **Grep and Glob timeouts** - Grep and Glob timeouts now surface as actionable errors - **Theme detection over SSH** - Fixed theme auto-detection over remote terminal sessions - **Todo status colors** - Todo status colors now render correctly - **Completion notification** - The turn-completion notification now fires before the CLI exits - **Subagent reports** - The end-of-turn todo reminder no longer swallows subagent reports - **Mission settings** - Fixed updating mission settings in the desktop app (app) - **Mission notifications** - Mission mode now fires audio and badge notifications (app) - **Model defaults** - Model defaults now persist correctly (app) - **Reasoning tooltip** - The reasoning tooltip now dismisses when you start typing (app) - **Local storage failures** - The app no longer breaks when local storage is unavailable (app) ## CLI v0.170.0 · Desktop v0.127.0 (July 10, 2026) Legacy file encoding support, plus browser and composer refinements ### Improvements - **Browser panel suggestions** - Local suggestions in the browser panel are now deduplicated and better ranked (app) - **Prompt history recall** - Refined recall of previous prompts in the composer (app) ### Bug fixes - **Legacy file encodings** - File tools now preserve legacy text encodings such as ISO-8859-2 and CP1250 instead of converting them ## CLI v0.169.0 · Desktop v0.126.0 (July 9, 2026) macOS process sandbox, spec rendering fixes, and MCP session improvements ### New features - **macOS command sandbox** - Added whole-process sandbox support on macOS for safer command execution ### Improvements - **Clearer send button** - The composer now explains why the send button is disabled (app) - **Safer web tools** - WebSearch and FetchUrl now require at least low autonomy before they run ### Bug fixes - **Image-only messages** - Fixed sending messages that contain only an image - **Plan todos** - Plan todo lists are now bounded and the full todo displays correctly - **Wide tables** - Wide markdown tables now scroll horizontally - **MCP footer** - The MCP status footer now appears immediately - **MCP session notices** - MCP status and auth notices now stay scoped to their session instead of leaking into untouched sessions - **Spec rendering** - Spec contents now render in the approval details screen, and the spec body shows when the plan heading matches the title - **Update rollback** - Hardened the update rollback flow (app) ## CLI v0.168.0 · Desktop v0.125.0 (July 8, 2026) Pin local sessions, working directory in the desktop app, and CLI and app fixes ### New features - **Pin local sessions** - You can now pin local sessions that aren't synced to the cloud (app) - **Working directory in desktop** - The desktop app now brings your current working directory into the session (app) ### Improvements - **Slash command search** - Improved slash command search results - **Skills count** - The header now shows an accurate skills count - **Wiki permalinks** - Added repository permalink support in wikis ### Bug fixes - **Session token usage** - Session token usage is now counted only on successful turns - **Unread sessions** - Unrelated unread sessions are preserved when a window regains focus (app) - **Notification hook** - The `idle_prompt` Notification hook now fires on cancelled turns - **Reasoning effort** - `set-default` now preserves your session reasoning effort - **Bash mode** - Pressing backspace on empty input now exits bash mode - **Spec proposals** - Empty spec proposal bodies are now rejected - **Ghostty Open In** - Fixed launching Ghostty from the Open In deeplink (app) - **Window close warning** - The app now only warns on window close when it will actually quit the app (app) - **MCP OAuth screen** - Aligned the MCP OAuth callback screen ## CLI v0.167.0 · Desktop v0.124.0 (July 7, 2026) Port-forwarding for droid computers, faster startup and search, and Windows and Linux fixes ### New features - **Port-forward for droid computers** - New `droid computer port-forward` command to forward ports from a droid computer to your local machine ### Improvements - **Faster app startup** - Startup boot phases now run in parallel, reducing time to a usable window (app) - **Faster session search** - Session search is quicker and returns better-ranked results - **Full-screen help** - Help now opens as a full-screen overlay you can dismiss with `?` or `Esc` - **Slash menu filtering** - Filtered slash menu results are flattened with a clearer category block - **Cleaner model names** - Model names display more clearly across the app (app) - **Cleaner chat transcript** - Internal tool lookups no longer clutter the chat transcript (app) - **Update rollback** - The app can now roll back to a previous version when needed (app) ### Bug fixes - **Image paste on Windows** - Fixed pasting images into the composer on Windows - **Safe exit during updates** - You can now exit safely while a background update is in progress - **SessionStart hooks** - A hung `SessionStart` hook no longer blocks session initialization - **Completion sound** - Restored the completion sound - **Windows home directory** - Fixed resolving the home directory for non-ASCII Windows usernames - **Draft attachments** - Draft attachments are now preserved when the screen repaints - **Sandboxed hooks** - Fixed hook command execution under the sandbox - **Detailed transcript freeze** - Capped detailed transcript tool results to stop `Ctrl+O` from freezing - **New session shortcut** - The new-session shortcut now ignores Shift (app) - **Truncated sessions** - Fixed context handling in truncated sessions - **Linux clipboard** - Text clipboard is now Wayland-first on Linux - **Special characters** - Typing `å`, `Å`, or `´` is no longer treated as an Option shortcut - **Unavailable tools** - Calls to unavailable tools are now clearly marked as errors - **Highlighted code** - Highlighted code text is now preserved (app) - **Airgap mode** - Improved the reliability of airgap mode (app) ## CLI v0.164.0 · Desktop v0.121.0 (July 2, 2026) Grouped slash commands, cloud sync control, and MCP and login fixes ### New features - **Grouped slash commands** - Slash commands are now organized by category for easier browsing - **Cloud sync control** - New setting to control whether your data syncs to the cloud (app) ### Improvements - **Linear model picker** - Model names now display cleanly in the Linear delegation picker (app) - **Computer provisioning wizard** - Clarified copy in the droid computer provisioning wizard (app) ### Bug fixes - **MCP tool names** - Tool names containing dots are now accepted per the MCP specification - **MCP authentication** - Restored sign-in to MCP servers that reject certain OAuth client registration - **Background processes in droid exec** - Agent-spawned background processes now keep running after `droid exec` exits - **Editor for allow/deny lists** - The CLI now honors `$EDITOR` when editing allow and deny lists on headless hosts - **Gemini reasoning** - Fixed handling of invalid Gemini reasoning blocks - **Droid computer SSH** - Fixed connecting to a droid computer by its exact name - **Cross-org login** - Fixed login issues when working across multiple organizations (app) ## CLI v0.162.0 · Desktop v0.119.0 (June 30, 2026) Cloned repos in projects, lower memory use, and rendering and auth fixes ### New features - **Cloned repos in projects** - Cloned repositories now appear in the new-session projects dropdown (app) ### Improvements - **Lower memory usage** - Reduced memory use during long and idle sessions - **Word-wise navigation** - Option+Left/Right now jumps between words in secondary text inputs ### Bug fixes - **Resumed hooks** - Hook rows now render correctly after you resume a session - **File encoding** - Edited files now keep their original encoding instead of always saving as UTF-8 - **MCP authentication** - Restored basic auth fallback when signing in to MCP servers over OAuth - **Org switching** - Fixed selecting the active organization when switching organizations, including in the EU (app) ## CLI v0.161.0 · Desktop v0.118.0 (June 29, 2026) Hooks manager overhaul, chat transcript search, and faster searches ### New features - **Hooks manager overhaul** - Redesigned the hooks manager for viewing and configuring your hooks - **WezTerm setup** - Added terminal setup support for WezTerm - **Chat transcript search** - Search within a chat transcript (app) ### Improvements - **Faster searches** - Avoided unnecessary search cache rebuilds ## CLI v0.159.0 · Desktop v0.116.0 (June 25, 2026) Archive projects with clearer MCP authentication guidance ### New features - **Archive projects** - Archive projects you're no longer working on to keep your project list tidy (app) ### Improvements - **Clearer MCP auth guidance** - When an MCP server fails to authenticate, the CLI now explains how to resolve it ### Bug fixes - **MCP config errors** - MCP configuration parse errors are now surfaced clearly instead of failing silently ## CLI v0.158.0 · Desktop v0.115.0 (June 24, 2026) Plugin toggles, tab slash-command completion, and login and rendering fixes ### New features - **Enable and disable plugins** - Turn individual plugins on or off with a new toggle ### Improvements - **Tab to complete slash commands** - Press Tab to complete slash commands - **MCP authentication** - Improved handling of MCP authentication responses and desktop redirects - **Account credit checkout** - Subscribe to or upgrade your plan using account credit ### Bug fixes - **Invalid API key login** - The CLI now exits with an error on an invalid `FACTORY_API_KEY` instead of stalling at login - **Cancel AskUser prompts** - Cancelling an AskUser prompt now interrupts the agent instead of letting it continue - **Duplicate tool approvals** - Approval-gated tool executions now render only once - **Windows background processes** - Fixed cleanup of background processes on exit on Windows - **Document skills location** - Document skills are now written to your current working directory - **Dark mode toggles** - Dark mode toggles are now visible and loading spinners animate correctly (app) - **Editable composer** - The composer stays editable during session blockers (app) ## CLI v0.157.2 · Desktop v0.114.2 (June 24, 2026) Slack follow-up prompt fix ### Bug fixes - **Slack follow-up prompt** - Fixed a duplicate "tag to follow up" prompt that could appear in Slack threads ## CLI v0.157.0 · Desktop v0.114.0 (June 23, 2026) External prompt editing and input and onboarding fixes ### New features - **Edit prompts externally** - Open and edit your prompt in an external editor for composing longer messages ### Bug fixes - **Caps-lock cursor keys** - Fixed cursor key handling when caps lock is on - **Onboarding crash** - Fixed a crash that could occur when connecting before onboarding finished (app) ## CLI v0.156.2 · Desktop v0.113.2 (June 22, 2026) Conversation rewind and reliability fixes ### New features - **Rewind conversation** - The new `/rewind-conversation` command rewinds the chat and restores your files to an earlier point (app) - **Org hooks governance** - Organizations can now centrally control hooks settings across their members ### Improvements - **Clickable message links** - URLs in your messages are now clickable - **Light terminal colors** - Improved white rendering in the light terminal theme ### Bug fixes - **Reasoning indicator** - Non-thinking models no longer show a reasoning indicator - **Execute step shimmer** - Fixed a stale loading shimmer left on completed execute steps - **Windows hooks** - Preserved command quoting for hooks on Windows - **MCP server logos** - MCP server logos now keep their original colors ## CLI v0.152.0 · Desktop v0.109.0 (June 18, 2026) Command-line autonomy flag, org MCP and marketplace controls, and reliability fixes ### New features - **Autonomy flag** - Set Droid's autonomy level straight from the command line with `droid --auto high` - **Org MCP and marketplace controls** - Organizations can centrally govern MCP server and plugin marketplace access - **In-app membership controls** - Manage organization members directly in the app (app) ### Improvements - **Private npm registry plugins** - Plugin sources can now authenticate to private npm registries using an auth token environment variable - **Onboarding copy** - Clearer guidance during Droid setup onboarding (app) ### Bug fixes - **Status after interrupt** - Droid's status now clears correctly after you interrupt it - **Event hooks** - Multiple hooks configured for the same event are now combined correctly - **Summaries across directories** - Fixed conversation summaries when your working directory changes - **Streaming placeholder** - Fixed a stale placeholder that could linger in the chat (app) - **Connection stability** - Reduced session UI flicker during brief connection drops (app) ## CLI v0.151.0 · Desktop v0.108.0 (June 17, 2026) Branded macOS installer with MCP and session reliability fixes ### New features - **Skill slash commands in editors** - Skill slash commands are now available when connecting from external editors ### Improvements - **macOS installer branding** - The macOS DMG installer is now branded (app) - **MCP results as markdown** - MCP tool results returned as JSON now render as formatted markdown ### Bug fixes - **Session reconnection** - Sessions now stay rendered and reconnect instead of redirecting to a Session Unavailable page (app) - **Org switcher** - The desktop organization switcher now shows the currently selected organization (app) - **Desktop MCP sign-in** - Improved desktop MCP authentication callback handling (app) - **Composer actions layout** - Composer actions now stay on a single line (app) ## CLI v0.150.0 · Desktop v0.107.0 (June 16, 2026) npm plugin sources and long-session performance ### New features - **npm plugin sources** - Plugin marketplace definitions can now reference plugins published as npm packages - **Quit confirmation** - Desktop now asks for confirmation before quitting or closing would interrupt an active Droid session (app) ### Improvements - **Long session performance** - Improved chat performance in long sessions (app) - **MCP authentication page** - Polished the MCP authentication success page - **Overdue invoice banner** - Billing settings now show a banner when an invoice is overdue (app) ### Bug fixes - **Rate limit reset display** - Fixed the rate limit reset time shown for sub-hour resets ## CLI v0.149.0 · Desktop v0.106.0 (June 15, 2026) Canva MCP authentication fix ### Bug fixes - **Canva MCP authentication** - Added a fallback so connecting the Canva MCP server works more reliably ## CLI v0.148.1 · Desktop v0.105.0 (June 15, 2026) Air-gapped CLI binary distribution fix ### Bug fixes - **Air-gapped binaries** - Air-gapped CLI binaries are now properly signed and published so air-gapped installs receive the latest builds ## CLI v0.148.0 · Desktop v0.105.0 (June 14, 2026) MCP connection status, org-wide auto-update control, and reliability fixes ### New features - **MCP connection status** - `droid mcp list` and your session now show each MCP server's connection and authentication status - **Org-wide auto-update control** - Organizations can turn off CLI auto-updates with a new `disableAutoUpdate` setting ### Improvements - **Log rotation** - The CLI now rotates its log files to avoid filling up disk space - **Faster MCP loading in editors** - MCP servers load without blocking when connecting from external editors - **Streamlined desktop sign-up** - Signing up on desktop now takes you straight to login (app) - **Clearer unsaved changes banner** - The unsaved changes banner is now easier to spot (app) ### Bug fixes - **Duplicate BYOK models** - Removed duplicate models from the bring-your-own-key model selector - **BYOK custom endpoints** - Fixed compatibility with custom bring-your-own-key model endpoints - **Duplicate transcript** - Fixed a duplicate transcript appearing in the session view - **Plugins slash command** - The `/plugins` command now appears before a session finishes initializing (app) - **Wiki reader dragging** - Restored window dragging in the wiki reader (app) - **MCP authentication notices** - Improved how MCP authentication prompts are shown ## CLI v0.147.0 · Desktop v0.104.0 (June 12, 2026) Command blocklist, MCP OAuth improvements, and welcome screen sign-in ### New features - **Command blocklist** - The new `commandBlocklist` setting hard-stops specified commands so they can never run, even under full autonomy or skipped approvals - **Welcome screen sign-in** - Added log-in and sign-up CTAs to the welcome screen (app) ### Improvements - **MCP OAuth** - MCP servers can now authenticate using Client ID Metadata Documents (CIMD), with improved compatibility for servers that require a resource indicator - **Updated tooltips** - Refreshed tooltips for clearer guidance ### Bug fixes - **Ctrl+Enter tooltip** - Fixed the Ctrl+Enter tooltip - **Desktop MCP OAuth banner** - The MCP OAuth banner now points at the MCP button (app) ## CLI v0.146.0 · Desktop v0.103.0 (June 12, 2026) Browser-based desktop onboarding, air-gapped mode improvements, and chat, MCP, and onboarding fixes ### New features - **Browser-based desktop onboarding** - Desktop onboarding now opens in your browser for a smoother first-run setup (app) ### Improvements - **Queue message shortcut** - Queue a message with Ctrl+Enter - **Earlier session titles** - Droid now names your session earlier in the conversation - **Air-gapped BYOK defaults** - Air-gapped mode auto-picks a sensible default BYOK model, and configuration errors point you at `settings.json` ### Bug fixes - **Chat cursor on resize** - The chat input now keeps your cursor position when the terminal is resized - **Context summarization hangs** - Fixed a rare hang that could occur while summarizing long conversations - **Content filtering** - Responses blocked by a content filter are now surfaced as a clear moderation error - **Desktop MCP OAuth** - Desktop MCP OAuth now flows through the hosted callback (app) - **Desktop onboarding top bar** - The top bar can now be dragged on desktop onboarding screens (app) - **Data retention dialog** - The data retention dialog no longer interferes with desktop navigation (app) ## CLI v0.144.2 · Desktop v0.101.2 (June 10, 2026) MCP authentication fixes ### Bug fixes - **MCP authentication** - Improved MCP sign-in reliability for servers that advertise no scopes - **Desktop MCP OAuth banner** - The MCP OAuth banner no longer overlaps the top bar (app) ## CLI v0.144.1 · Desktop v0.101.1 (June 9, 2026) More reliable context summarization when a model declines a request ### Bug fixes - **Context summarization fallback** - When a model declines to summarize your conversation, Droid now falls back gracefully instead of retrying repeatedly ## CLI v0.144.0 · Desktop v0.101.0 (June 9, 2026) Session file suggestions, safer command risk handling, and MCP reliability fixes ### New features - **Session file suggestions** - The new session page now suggests relevant files as you get started (app) ### Improvements - **Safer command risk handling** - Commands that could destroy your working tree are now treated as high risk - **tmux key support** - Improved keyboard handling when running inside tmux ### Bug fixes - **MCP reliability** - Improved reliability of MCP authentication and connectors ## CLI v0.143.1 · Desktop v0.100.1 (June 9, 2026) Maintenance and stability release ### Improvements - **Maintenance and stability** - Routine maintenance and reliability updates under the hood ## CLI v0.143.0 · Desktop v0.100.0 (June 8, 2026) Richer Jira details ### New features - **Self-serve organization deactivation** - Deactivate your organization yourself, with subscription cancellation now available in billing settings (app) ### Improvements - **Richer Jira details** - Fetching a Jira issue now includes more fields, comments, and linked-issue URLs - **Quieter hooks** - Hooks can hide their output block in the TUI by setting `suppressOutput` ### Bug fixes - **Windows home directory** - Fixed home directory resolution on Windows to avoid username corruption (app) - **Session integrity** - Fixed an issue that could overwrite an existing session ## CLI v0.142.0 · Desktop v0.99.0 (June 5, 2026) GitLab code review setup and chat stability fixes ### Improvements - **GitLab code review setup** - The install-code-review skill now configures review pipelines using a GitLab CI Component ### Bug fixes - **Duplicate chat messages** - Fixed duplicate messages appearing in chat - **Diff viewer counts** - The diff viewer now refreshes its change counts correctly (app) ## CLI v0.141.0 · Desktop v0.98.0 (June 4, 2026) More reliable command output and CLI fixes ### Improvements - **Large command output** - Command execution now handles very large outputs more reliably ### Bug fixes - **Input hint placement** - The running hint now stays below your typed input - **Guidelines in non-git projects** - Droid no longer looks for guideline files in parent directories when you're not inside a git repository ## CLI v0.140.1 · Desktop v0.97.1 (June 4, 2026) Maintenance and stability release ### Improvements - **Maintenance and stability** - Routine maintenance and reliability updates under the hood ## CLI v0.140.0 · Desktop v0.97.0 (June 3, 2026) Claude Opus 4.8 and Gemini 3.5 Flash support, MCP tool search, incident response, and stability fixes ### New features - **Claude Opus 4.8** - Added support for Claude Opus 4.8, including a faster Fast Mode variant - **Gemini 3.5 Flash** - Added support for Gemini 3.5 Flash - **MCP tool search** - Droid can now search and load MCP tools on demand, keeping large MCP setups fast and context-efficient - **Incident response** - Droid can automatically investigate production alerts posted to a configured Slack channel and work toward a root-cause analysis ### Improvements - **Accurate status banners** - Spinner status banners now reflect the true state of the current turn - **Context menu text actions** - Added text actions to the right-click context menu (app) ### Bug fixes - **Terminal write errors** - The CLI no longer crashes when the terminal rejects output writes - **Paste handling** - Fixed a crash on certain pasted input and corrected handling of split bracketed pastes - **Agent loop stalls** - The agent loop no longer runs unbounded when turns produce no output - **Empty chat placeholder** - Restored the placeholder shown in empty chats (app) - **Machine connection crash** - Fixed a crash when loading sessions with older machine connection types (app) ## CLI v0.139.0 · Desktop v0.96.0 (June 2, 2026) Custom commands and skills in the slash menu and MCP fixes ### New features - **Custom commands and skills in the slash menu** - Your custom commands and skills now appear directly in the slash menu ### Bug fixes - **MCP schema compatibility** - Improved handling of MCP servers with incompatible tool schemas - **Clearing MCP server fields** - You can now clear fields when editing an MCP server - **Local session labels** - Local sessions are now labeled "Local" instead of "Ephemeral" (app) ## CLI v0.138.0 · Desktop v0.95.0 (June 1, 2026) MCP management upgrades, Mermaid export, and stability fixes ### New features - **Interactive MCP management** - Add, remove, and list MCP servers interactively - **MCP credential variables** - MCP server configs can now expand credential variables at runtime - **Configurable MCP call timeout** - Set a custom timeout for MCP tool calls - **`mcp off` shortcut** - Quickly disable MCP servers with a new shortcut - **MCP servers for custom droids** - Choose which MCP servers are available to custom droids - **Per-server MCP risk** - Configure command risk levels for each MCP server - **Mermaid export** - Export Mermaid diagrams from chat (app) ### Improvements - **Highlighted new models** - New models now appear highlighted at the top of `/models` - **MCP embedded resources** - MCP servers' embedded resources are now surfaced in chat - **Composer loading state** - The message composer now shows a loading state (app) - **Session archiving** - Unified the session archive flow across the app and desktop (app) - **Pinned messages** - Clicking a pinned message now scrolls to its source (app) ### Bug fixes - **Turn cancellation** - Canceling a turn now takes effect immediately - **Session rename** - Renamed session titles are now trimmed of extra whitespace (app) ## CLI v0.137.0 · Desktop v0.94.0 (May 29, 2026) Usage stats, airgapped BYOK, and stability fixes ### Bug fixes - **Duplicate thinking output** - Thinking output no longer renders twice - **MCP permission persistence** - MCP permissions now persist across all confirmation paths - **Context usage meter** - Fixed alignment of the context usage meter ## CLI v0.136.0 · Desktop v0.93.0 (May 28, 2026) Deep review skill, side questions in the app, session compression controls, and stability fixes ### New features - **Deep review skill** - New `/deep-review` skill for a more thorough, multi-pass code review - **Side questions in the app** - The `/btw` side questions panel is now available in the app (app) - **Auto-compression toggle** - Turn automatic session compression on or off to manage long sessions - **Configurable compression model** - Choose which model handles automatic session compression ### Improvements - **Org details in `/status`** - The `/status` command now shows your organization details - **Standardized hooks config path** - Hooks configuration now uses a standardized file path with automatic migration - **Faster large pastes** - Sped up pasting large amounts of text - **Tool summary hover affordance** - Tool summaries now show a hover affordance for available actions (app) ### Bug fixes - **Autoscroll fixes** - Fixed several autoscroll issues in the session view (app) - **Reasoning-only turns** - Fixed sessions stalling when the assistant returned only reasoning - **DeepSeek tool calls** - Improved reliability of multi-step tool calls when using DeepSeek models ## CLI v0.135.0 · Desktop v0.92.0 (May 27, 2026) Adaptive terminal theme, agent-browser update, and stability fixes ### New features - **Adaptive terminal theme** - The CLI now auto-detects your terminal background and picks a matching theme ### Improvements - **Agent-browser update** - Updated the bundled `agent-browser` skill to a newer release with better automation reliability - **Subagent type badges** - Subagent type badges in `TaskOutput` now render in PascalCase for easier reading ### Bug fixes - **Message draft attachments** - Message draft attachments are now preserved when switching contexts (app) - **Session activity indicator** - Improved the activity indicator in the session panel (app) - **Enterprise Controls forms** - Empty arrays and records are no longer written when saving Enterprise Controls (app) - **Enterprise Controls layout** - Aligned control heights across Enterprise Controls forms (app) ## CLI v0.134.0 · Desktop v0.91.0 (May 26, 2026) Stability fixes and UI polish ### Improvements - **Figma helper skill** - The built-in Figma skill has been renamed from `figma-mcp-promotion` to `figma-mcp-helper` - **System-managed settings** - System-managed settings are now loaded from standard platform paths ### Bug fixes - **File tool rendering** - File tool inputs now wait for streaming to complete before rendering to avoid showing partial data - **Slash command suggestions** - Slash command suggestion rows now keep a consistent height - **Pending command approvals** - Capped the number of pending command approvals shown at once - **New session sidebar animation** - Smoothed the animation when starting a new session from the sidebar (app) - **Session titles in sidebar** - Session titles are now preserved correctly in the sidebar (app) - **AskUser question labels** - Question labels now pluralize correctly when multiple questions are asked ## CLI v0.133.1 · Desktop v0.90.1 (May 25, 2026) Context window modal, audio cues, sidebar session menu, and stability fixes ### New features - **Context window** - New `/context` slash command and modal for inspecting the current session's context usage - **Audio cues** - Optional sounds when a session finishes or is awaiting your input (app) - **Sidebar session menu** - Right-click sessions in the sidebar for quick actions (app) ### Bug fixes - **Final output visibility** - The final output now stays visible above the input box - **New session draft cleared** - New session drafts now clear after submitting (app) - **Linear Computer setup link** - The "set up Computer" call to action in Linear now points to the correct settings page - **Desktop zoom stability** - Hardened desktop zoom handling against host zoom errors (app) ## CLI v0.132.0 · Desktop v0.89.0 (May 22, 2026) Render deployments and stability fixes ### New features - **Deploy to Render** - New `/deploy-to-render` skill helps you deploy projects to Render ### Improvements - **Security review scope confirmation** - The `/security-review` skill now confirms the scope and audit branch with you before running - **Faster EU connections** - Reduced connection latency for users in the EU region ### Bug fixes - **Disabled model selector rows** - Restored disabled rows in the model selector so all options are visible - **Missions across reconnects** - Active missions now keep running through network reconnects - **Mission computer in mission control** - The active mission computer now appears on the mission control page (app) - **Windows desktop setup** - Fixed setup and scripts for the Windows desktop app (app) - **Rich text autofocus** - Improved autofocus reliability for rich text inputs (app) ## CLI v0.131.0 · Desktop v0.88.0 (May 21, 2026) Maintenance release ### Improvements - **Maintenance release** - with under-the-hood improvements. ## CLI v0.130.0 · Desktop v0.87.0 (May 20, 2026) Model favorites, spec save directory, paste expansion, and stability fixes ### New features - **Model favorites** - Mark and quickly access your preferred models from the model selector - **Spec save directory** - Configure where specs are saved with a new setting - **Expand pasted placeholders** - Press Alt+Shift+V (Opt+Shift+V on macOS) to expand truncated paste placeholders inline ### Improvements - **Refined model selector** - Streamlined the model selector submenu - **BYOK model parity** - Brought the BYOK model experience in line with the rest of the model picker - **Clearer help output** - Improved formatting of the `help` command output ### Bug fixes - **Image paste on Linux** - Pasting images now works on Linux, with a graceful fallback when clipboard data is unavailable - **MCP OAuth clients** - Static MCP OAuth clients are now supported during authentication - **OAuth provider errors** - 4xx responses from connected OAuth providers are now surfaced as user-facing errors - **ApplyPatch new files** - New files in ApplyPatch results now render correctly as creates - **CLI update reliability** - Prevented restart races during CLI daemon updates - **PowerShell under `exec`** - PowerShell commands now run non-interactively in `exec` mode - **Slack first-mention indicator** - Slack now shows the typing indicator from the very first mention - **Daemon auth on reconnect** - Daemon auth tokens now refresh when reconnecting (app) - **Desktop paid onboarding** - The onboarding finish screen is now skipped on the desktop paid flow (app) ## CLI v0.129.0 · Desktop v0.86.0 (May 18, 2026) Streaming file hook events, faster large sessions, and stability fixes ### New features - **Streaming file hook events** - File hook execution events now stream so you can follow progress as hooks run ### Improvements - **Faster large session rendering** - Large session transcripts now render faster by trimming at the cutoff before derivation ### Bug fixes - **Esc cancels tool activity** - Pressing Esc now reliably cancels in-progress tool activity - **MCP tools with complex schemas** - Listing MCP tools no longer fails when schemas contain unresolvable references - **Claude Code subagent import** - Subagent import is now skipped when there is no project root to import into - **Session load file snapshots** - File snapshots are now initialized when reloading a saved session - **ACP todo updates** - Hardened parsing of todo updates over ACP against missing fields - **Read-only badge layout** - Read-only badges no longer get compressed in narrow layouts (app) - **Pinned group new session action** - Removed an unintended new-session action on pinned groups (app) - **Desktop session recovery** - Desktop sessions now recover after a period of daemon inactivity (app) ## CLI v0.128.0 · Desktop v0.85.0 (May 15, 2026) Stability fixes ### Bug fixes - **Stale Computer registrations** - Sessions now reconcile stale Droid Computer registrations automatically ## CLI v0.126.0 · Desktop v0.83.0 (May 14, 2026) Bulk session archiving, faster Computer restarts, EU data residency, and stability fixes ### New features - **Bulk session archiving** - New settings page for managing and bulk-archiving sessions (app) - **EU data residency** - Sessions and LLM inference for EU customers now stay within an EU-only data plane - **Built-in security review** - The `/security-review` skill is now bundled with the CLI by default - **Copy bug report ID** - Added a copy button next to bug report IDs (app) ### Improvements - **Faster Computer restarts** - Sped up restarting Droid Computers (app) - **Lighter CLI startup** - Reduced redundant requests during CLI startup - **Persistent desktop preferences** - Desktop preferences are now saved alongside other settings (app) - **Windows pending update cleanup** - Cleaned up stale pending updates on Windows ### Bug fixes - **Structured output validation** - Strengthened validation of structured tool output - **Task token count** - Removed an incorrect zero token count from Task tool results - **Windows installer packaging** - Repaired Windows installer packaging - **Limits reset on plan change** - Plan changes now correctly reset usage limits (app) - **MCP servers for running sessions** - MCP servers now load when reopening already-running sessions (app) - **Local daemon warning during setup** - Suppressed a spurious local daemon warning during initial setup (app) - **Session not found errors** - Fixed spurious session not found errors (app) - **Code block styling** - Aligned code block styling in chat (app) - **Cross-surface integrations** - Prevented integrations from sending duplicate postbacks across surfaces - **QA droid CLI auth** - QA droid CLI now pairs with the matching Factory API host ## CLI v0.125.1 · Desktop v0.82.1 (May 14, 2026) Maintenance release ### Improvements - **Maintenance release** - with under-the-hood improvements. ## CLI v0.125.0 · Desktop v0.82.0 (May 13, 2026) ACP config options, larger Task results, and stability fixes ### New features - **ACP config options** - The CLI now advertises configuration options to ACP clients and supports `session/set_config_option` for per-session overrides ### Improvements - **Larger Task tool results** - Raised the truncation limit on Task tool results so more output is preserved ### Bug fixes - **Stale tool shimmer** - Tool calls no longer leave a stale shimmer behind in the transcript - **Stale Slack shimmer heartbeat** - Stopped a stale shimmer heartbeat from lingering in Slack threads - **Return to session on interaction prompt** - The CLI now returns to the session view when a user interaction box surfaces - **Gemini thinking signatures on replay** - Gemini thinking signatures are now preserved when sessions replay - **Title bar path on hover** - The title bar now shows the full working directory path on hover (app) ## CLI v0.124.0 · Desktop v0.81.0 (May 12, 2026) Deep and shallow security review modes, login redirect URIs, and stability fixes ### New features - **Deep and shallow security review modes** - `/security-review` now supports dedicated deep and shallow review modes - **Login redirect URIs** - The login flow now supports a redirect URI so you land where you started after signing in (app) ### Improvements - **Sharper fuzzy search** - Refined fuzzy search behavior in the CLI - **Minimum context savings shown** - Context savings now display a clear minimum percent for easier interpretation - **Aligned MCP status indicators** - MCP server status indicators now line up consistently - **Faster setting updates** - Settings changes are applied in place, avoiding unnecessary MCP and skill refetches (app) - **Light mode polish** - Small component styling refinements in light mode ### Bug fixes - **Bounded resume history** - Resume now renders a bounded history so large sessions stay responsive - **Message send flash** - Fixed a visual flash when submitting a message (app) - **Desktop tutorial in desktop app** - The "install desktop" onboarding step is now skipped when you're already in the desktop app (app) ## CLI v0.123.1 (May 12, 2026) Maintenance release ### Improvements - **Maintenance release** - with under-the-hood improvements. ## CLI v0.123.0 · Desktop v0.80.0 (May 11, 2026) Factory token usage in mission control, model deprecation notices, and stability fixes ### New features - **Factory token usage in Mission Control** - Factory token usage is now persisted to session settings, with a toggle to show it in Mission Control - **Model deprecation notices** - Factory now surfaces notices when a model is scheduled for deprecation - **Optimistic message submit** - Messages can now be submitted before the session connection is ready (app) ### Improvements - **Clearer parallel execute approvals** - Parallel execute approval prompts now make it clearer what's being approved - **Chat notice display** - Refreshed how chat notices render (app) ### Bug fixes - **Spec reasoning tab cycle** - Restored cycling through tabs in the spec reasoning view - **Duplicate spec rendering** - Specs no longer render twice - **Padded multiline status lines** - Multiline status lines now render with consistent padding - **Subagent tokens in Mission Control** - Mission Control token counts now include subagent tokens - **Queued messages during /compress** - Messages typed while `/compress` is running are now queued instead of dropped - **Stale /missions prompt** - Fixed a stale `/missions` prompt lingering after dismissal - **Provider tool name handling** - Improved handling of tool names returned by model providers - **Stuck session loading** - Fixed sessions that could get stuck loading (app) - **Ghost update badge** - Fixed a spurious update-available badge caused by an update manifest mismatch (app) ## CLI v0.122.0 · Desktop v0.79.0 (May 8, 2026) Manager-only API keys, tighter Computer name validation, and stability fixes ### Improvements - **Manager-only API key creation** - Organization API keys can now only be created by managers (app) - **Stricter Computer name validation** - Computer names are limited to lowercase letters, digits, and hyphens, with clearer error states (app) ### Bug fixes - **Working directory from computer wizard** - Working directory is now seeded correctly when navigating from the computer wizard to a new session (app) - **Duplicate collapsed groups** - Fixed duplicate collapsed groups in the CLI - **Droid header spacing** - Restored a missing trailing space in the Droid header - **Extra usage on subscription end** - Remaining extra usage is now transferred to your account balance when a subscription ends (app) ## CLI v0.121.0 · Desktop v0.78.0 (May 7, 2026) Thinking duration, smoother scrollback, better session titles, and stability fixes ### New features - **Thinking duration** - CLI and app now show how long a droid spent thinking - **Scroll to most recent prompt** - You can now scroll down directly to the most recent prompt in the CLI ### Improvements - **Smoother transcript scrollback** - Cached static transcript markdown rendering for faster scrolling - **Better session titles** - Session titles are now well-formed, with redundant syncs deduplicated - **Tighter /sessions layout** - Column widths in `/sessions` are now distributed more evenly - **Faster setup failure feedback** - Setup now fails fast with a clearer error when the daemon can't connect ### Bug fixes - **Pasting long blocks** - Fixed issues when pasting large blocks of text into the CLI - **ACP tool call lifecycle** - Repaired the tool call lifecycle for ACP sessions - **Droid binary download** - Fixed an installation error when downloading the droid binary on some systems - **Droid Computer reconnection** - Droid Computers now reconnect automatically on activity (Factory App) - **Clearer daemon error messages** - Improved daemon error messaging in the Factory App ## CLI v0.120.0 · Desktop v0.77.0 (May 6, 2026) Per-user Droid Computer controls, Slack auto-run buttons, Windows performance, and stability fixes ### New features - **Per-user Droid Computer controls** - Admins can now toggle managed and bring-your-own-machine (BYOM) Droid Computers on a per-user basis (Factory App) ### Improvements - **Windows performance** - Sped up CLI responsiveness on Windows - **More reliable OpenAI file edits** - File edits now use the native patch tool for better reliability on OpenAI models ### Bug fixes - **Wrapped hyperlink targets** - Hyperlinks that wrap across lines now preserve the correct target - **Working directory after Droid Computer change** - The working directory is now restored when the selected Droid Computer changes (Factory App) - **Non-ASCII paths on Windows** - Native Windows tools now handle non-ASCII paths correctly - **Repeated context management notices** - Stopped showing the same context management notice multiple times in a row - **Windows daemon port retry** - Daemon now retries binding correctly when a port is already in use on Windows - **Plan cache after upgrade** - Organization cache now refreshes after plan changes so the new plan appears immediately (Factory App) - **Droid Shield reliability** - Additional fixes to Droid Shield - **Applying patch header** - Removed a redundant "applying patch" header - **Droid message crash** - Fixed a rare crash when rendering droid messages (Factory App) - **Payment methods for new organizations** - Organizations without a billing profile now see an empty payment method list instead of an error (Factory App) - **Outdated thinking settings hint** - Removed a stale hint related to thinking settings ## CLI v0.119.0 · Desktop v0.76.0 (May 6, 2026) Windows setup script, MCP menu revamp, sticky user messages, Wiki Run deletion, and stability fixes ### New features - **Windows setup script** - New setup script simplifies installing the Droid CLI on Windows - **MCP menu revamp** - Redesigned MCP menu for clearer server and tool management - **Sticky user messages** - User messages stay anchored as you scroll, with a full-session view (Factory App) - **Delete Wiki Runs** - You can now delete Wiki Runs directly from the Factory App ### Improvements - **Factory App auto-updater UX** - Overhauled the Factory App auto-updater experience for a clearer update flow - **Droid Computer wizard polish** - Streamlined the new Droid Computer wizard (Factory App) - **Cleaner ask-user prompts** - Polished the ask-user prompt UI in the CLI - **Refined Thinking/Exploring previews** - Tightened the look of Thinking and Exploring summary previews - **Light mode chat polish** - Improved chat styling in light mode - **Missions overage error clarity** - Distinguishes Missions overage network errors from generic failures ### Bug fixes - **Status line working directory** - Status line now uses the last known working directory instead of a stale path - **AutoWiki history dropdown** - AutoWiki history dropdown is no longer cut off by the surrounding layout (Factory App) - **BYOK reasoning effort range** - Full reasoning effort range now applies to custom BYOK models - **Stale custom provider routing** - Fixed stale routing persisting across custom BYOK provider switches - **Download Factory App banner on web** - "Download Factory App" banner no longer appears on the web - **Session switching flicker** - Fixed flickering when switching between sessions (Factory App) - **Hibernated Droid Computers for delegations** - Linear and Slack delegations now wake hibernated Droid Computers reliably - **Credits dashboard pricing** - Credits dashboard now works for accounts with multiple price intervals (Factory App) ## CLI v0.118.1 · Desktop v0.75.1 (May 5, 2026) Custom working directory for sessions created via the public API ### New features - **Custom working directory for sessions** - The public Sessions API now accepts an optional `cwd` field to start Droid Computer-backed sessions in a specific directory ### Bug fixes - **Clearer working directory errors** - Invalid working directories now fail fast with a structured error instead of silently falling back to the wrong directory ## CLI v0.118.0 · Desktop v0.75.0 (May 4, 2026) Smarter readiness hints, bulk user limit updates, live streaming in the app, and CLI stability fixes ### New features - **Live streaming in the Factory App** - Factory App now streams partial text and thinking updates as a droid works - **Bulk Usage Limit updates** - Admins can update Usage Limits for multiple users at once (Factory App) ### Improvements - **Smarter readiness hints** - Readiness hints now detect more cases where guidance helps before continuing - **Localized review findings** - Review findings now match your language instead of the source file's language - **Plus plan rename** - Pro Plus has been renamed to Plus throughout the product - **Directory-managed roles** - Role dropdown is now disabled for users managed by directory sync (Factory App) ### Bug fixes - **Paused response recovery** - Improved recovery from paused model responses - **System notification color** - Fixed color of system notification body copy - **Slack session link unfurls** - Session links no longer auto-unfurl in Slack threads - **Slack persona on replies** - Droid persona is now preserved on Slack replies ## CLI v0.117.0 · Desktop v0.74.0 (May 3, 2026) Members settings page, Individual plans, billing address at checkout, and CLI stability fixes ### Improvements - **Individual plans** - Self-serve plans are now grouped under Individual, with smoother loading states (Factory App) - **Sharper LLM error classification** - Overloaded and context-length errors now surface as their own error types ### Bug fixes - **Execute summary while streaming** - Execute tool summary is now hidden while a response is still streaming - **Chat log polish** - Minor visual fixes to the chat log - **/missions reminder in Mission Mode** - The /missions system reminder no longer shows when you're already in Mission Mode - **Empty template cloud session** - Submitting a cloud session with an empty template now works (Factory App) - **Thinking signature recovery** - Improved recovery from interrupted thinking signatures ## CLI v0.116.0 · Desktop v0.73.0 (May 2, 2026) Runtime-aware Execute tool, no-report readiness hint, payment method in onboarding, and CLI stability fixes ### New features - **Runtime-aware Execute tool** - Execute tool descriptions now adapt to your runtime, with WSL-specific guidance and an expanded bashism detector - **No-report readiness hint** - CLI surfaces a dynamic hint when a report isn't needed - **Payment method in onboarding** - You can now add a payment method during the final onboarding step (Factory App) ### Bug fixes - **Lingering prompts after resume** - User prompts are now cleared correctly after a droid resumes - **Gemini turn handling** - Scoped Gemini reasoning signals to the active turn for more reliable continuations - **Billing error classification** - Payment-required responses now surface as billing errors instead of generic invalid-request errors - **Working directory persistence** - Validated working directory is now persisted and locked until the Droid Computer connects (Factory App) - **Droid Computer details on transient errors** - Droid Computer details now remain visible during transient query errors (Factory App) - **Factory App updater download** - Factory App updater now downloads the new version only after you click to install ## CLI v0.115.0 · Desktop v0.72.0 (May 1, 2026) PDF support for Gemini 3.1 Pro, GLM-5.1 reasoning, bulk image drag-and-drop, PDF uploads in app, and CLI stability fixes ### New features - **PDF support for Gemini 3.1 Pro** - Gemini 3.1 Pro now accepts PDF inputs - **GLM-5.1 reasoning** - Added reasoning support for GLM-5.1 - **Bulk image drag-and-drop** - Drop multiple images into the CLI at once - **PDF uploads** - Upload PDF files directly in chat (Factory App) - **No Droid Computer state on new session** - The new session page now handles cases when no Droid Computers are available (Factory App) ### Improvements - **Resilient session resume** - Resuming a session that references a retired model now triggers a clean model migration - **Cloud sync stability** - Hardened cloud sync against session race conditions ### Bug fixes - **Terminal tab title** - Terminal tab title now updates when the active session changes - **Phantom lines in narrow terminals** - Fixed phantom lines that could appear in narrow terminals - **Mermaid terminal rendering** - Guarded Mermaid diagram rendering in the terminal against crashes - **Clearer auth errors** - Improved authentication error messages - **ACP session cancellation** - You can now cancel ACP sessions, and the correct auth code is shown - **Spinner rendering** - Fixed a spinner rendering issue - **Linear delegation matching** - Relaxed matching in the Linear delegation picker (Factory App) - **Session timeout messages** - Cleaner timeout messages and fixed a working directory error (Factory App) - **Connection failure toasts** - Removed noisy connection failure toasts (Factory App) - **Free trial onboarding** - Fixed a free trial onboarding issue (Factory App) - **Subagent collapse** - Subagents now collapse correctly even when the parent is selected (Factory App) - **AutoWiki export validation** - Hardened filename validation when exporting Wikis (Factory App) ## CLI v0.114.0 · Desktop v0.71.0 (April 30, 2026) BYOM admin toggle, Droid Shield push-scan refinement, Windows Python execution fix, and CLI stability fixes ### New features - **BYOM admin toggle** - Organization admins can now enable or disable bring-your-own-machine (BYOM) Droid Computers from enterprise settings (Factory App) ### Improvements - **Droid Shield push scan** - Push scanning now targets only commits not yet on any remote, reducing duplicate scans - **AutoWiki UI polish** - Visual refinements across the AutoWiki interface (Factory App) ### Bug fixes - **Agent loop continuation** - Agent loop now continues when a model response has no visible output - **Bash result spacing** - Removed an unwanted gap between bash commands and their results - **AskUser answer alignment** - Your own answer row in AskUser is now aligned with other rows - **Windows Python execution** - Python commands run on Windows now use a UTF-8 environment by default ## CLI v0.113.0 · Desktop v0.70.0 (April 29, 2026) Mission Control shortcuts, deny-list explanations, force-kill occupied ports, draft messages, and CLI stability fixes ### New features - **Mission Control shortcuts** - New `g`/`G` keyboard shortcuts to jump between workers and the features list in Mission Control - **Deny-list command explanations** - CLI now explains why a command was deny-listed when prompting for approval - **Force-kill occupied ports** - Reclaim a port that's already in use directly from the CLI - **Readiness hints** - CLI now surfaces occasional readiness hints to help guide your next step - **Draft messages** - Save and resume in-progress chat messages between sessions (Factory App) ### Improvements - **Randomized Academy quiz options** - Quiz answer choices are now randomized per question (Factory App) - **Cloud Machines status label** - Renamed Cloud Machines status from "deprecated" to "superseded" (Factory App) - **Draggable window during startup** - Factory App window stays draggable while the app finishes loading ### Bug fixes - **Git AI attribution** - Git AI attribution now records commit authors correctly - **Mission skill validation** - Mission skills are now validated when starting a mission run - **Archived sessions** - CLI no longer pulls archived sessions - **Slack delegation formatting** - Slack-delegated sessions now respond using Slack-native formatting - **Fewer Git diff failures** - Reduced intermittent Git diff failures during sessions - **Invite seat limits** - Invite flow now consistently enforces organization seat limits (Factory App) - **Default model selection** - New sessions now validate the persisted model against available models (Factory App) - **Long repository URLs** - Avoided request errors when repository names made GET URLs too long (Factory App) - **Academy callouts** - Academy markdown callouts now render reliably (Factory App) - **Chat scrolling** - Fixed several scrolling issues in the chat area (Factory App) - **macOS Factory App updater** - Fixed the Factory App auto-updater on macOS ## CLI v0.112.0 · Desktop v0.69.0 (April 28, 2026) DeepSeek V4 Pro support, MCP elicitation resume, security review on by default, and CLI stability fixes ### New features - **DeepSeek V4 Pro** - Added support for the DeepSeek V4 Pro model - **Resume MCP elicitations** - Elicitation prompts from MCP tools can now be resumed after interruption - **Extra-usage payment status** - Failed and succeeded extra-usage payment statuses are now surfaced in the payment modal (Factory App) ### Improvements - **Security review on by default** - The security-review skill is now enabled by default - **Non-retryable 4xx responses** - 4xx responses (other than 429) are no longer retried, so failures surface faster - **Persistent Slack loading status** - Slack loading indicators now stay visible past status updates for clearer progress feedback (Factory App) ### Bug fixes - **ApplyPatch absolute paths** - Absolute paths are now honored verbatim in the ApplyPatch tool - **Working directory redraw** - Fixed working directory redraw and fork `cwd` overrides - **Missions model settings** - Missions model settings are now snapshotted at launch so they don't drift mid-run - **Custom statusline** - Custom statusline now uses session-scoped data instead of leaking across sessions - **Model-switch tool history** - Tool history is now sanitized when switching models to prevent rendering issues - **Context window error detection** - Recognize additional context-window error phrasings so recovery is more reliable - **Spec approval comments** - Spec approval comments now align with conversation messages - **Factory App local default** - Factory App now consistently defaults to local connections - **Onboarding loading state** - Fixed a stuck loading state during onboarding (Factory App) ## CLI v0.111.0 · Desktop v0.68.0 (April 28, 2026) Italian translations, interval loop command, refreshed hooks UI, faster Grep, and CLI stability fixes ### New features - **Italian translations** - The Droid CLI is now available in Italian - **`interval` loop command** - New CLI command for running tasks on an interval - **Refreshed hooks UI** - Updated UI for managing hooks - **Custom commands in Task prompts** - Task tool prompts now expand resolved custom slash commands - **Chat transcript scroll keybinds** - New keyboard shortcuts to jump between turns in the chat view - **Missions Rate Limits** - Missions now handle Rate Limit responses gracefully - **Missions entry overage gating** - Missions entry now respects your organization's overage preference - **Per-run automation viewer** - New visual viewer to inspect each automation run (Factory App) - **Model access table** - New UI for managing organization-level model access (Factory App) - **Factory Academy certification tiers** - Detect certification tiers and expanded quizzes to 30 questions (Factory App) ### Improvements - **Faster Grep tool** - Grep tool now uses a hybrid ripgrep engine for better performance - **Stronger secret scrubbing** - Expanded secret scrubber patterns and applied scrubbing to Read tool output - **Faster `/new` and `/clear`** - Daemon session is now eagerly initialized on `/new` and `/clear` - **Centered settings on wide screens** - Settings content is now capped at 1024px and centered on wide displays (Factory App) - **Clearer schedule timezones** - Automation schedule timezone and run history timestamps are clearer (Factory App) - **Devtools shortcut in Factory App** - Re-enabled the devtools keyboard shortcut in production builds - **Clearer seat-limit errors** - Clearer error message when inviting members beyond your seat limit (Factory App) - **Native Slack loading status** - Slack now uses native loading status indicators (Factory App) ### Bug fixes - **Subagents during silent tools** - Task subagents are now kept alive during silent execution - **Execute preview rendering** - Bounded Execute tool preview rendering to prevent layout issues - **Execute approval rendering** - Stabilized rendering during Execute tool approvals - **Image paste on Linux/SSH** - Fixed image paste in Ghostty on Linux and over SSH - **Skill and droid preview scroll** - Skill and droid previews now scroll with arrow keys - **Chat Esc sequence** - Codified Esc key behavior in chat - **Spec mode editor flow** - Unblocked the spec edit flow when the editor opens via the system default - **Terminal indicator on Windows** - Hidden the terminal indicator on Windows in the footer - **Spec confirmation status** - Spec confirmation status is now gated on tool confirmation state - **Missions wake-lock leaks on Linux** - Prevented wake-lock leaks during Missions runs on Linux - **Skills rendering** - Fixed a parsing issue when rendering skills - **Subagent nesting** - Fixed nesting of subagents in the UI (Factory App) - **Cloud-synced Droid Computer sessions** - Cloud-synced Droid Computer sessions now display correctly (Factory App) - **Login screen error toast** - Suppressed an `/auth/me` error toast on the login screen (Factory App) - **Slack thread replies** - Slack now only forwards thread replies via `app_mention` (Factory App) - **Empty GitHub repos** - Added support for empty GitHub repositories (Factory App) - **Daemon directory paths** - Directory paths are now normalized in the daemon and render properly in the frontend (Factory App) ## CLI v0.109.3 · Desktop v0.66.3 (April 26, 2026) Auto-fallback to Droid Core, per-tool MCP autonomy overrides, Factory App system theming, and CLI stability fixes ### New features - **Auto-fallback to Droid Core** - Sessions that hit Rate Limits now automatically fall back to Droid Core when your overage preference is set to Droid Core - **Per-tool MCP autonomy overrides** - New `mcpAutonomyOverrides` setting lets you configure autonomy levels for individual MCP tools - **System theming** - Factory App now follows your system theme - **New Factory App updater** - Improved auto-update flow for the Factory App - **Per-channel Slack auto-run model** - Choose which model is used for each Slack auto-run channel (Factory App) ### Improvements - **Larger Droid Shield buffer** - Increased Droid Shield max buffer from 20MB to 64MB to handle larger outputs - **Faster hibernation detection** - Droid Computer hibernation is now detected more quickly - **Sorted model list** - CLI model list is now sorted by release date, newest first - **Cleaner MCP tool parameters** - MCP tool parameters show truncated JSON instead of `[object Object]` ### Bug fixes - **Lost subagent output** - Fixed an issue where subagent output could be lost - **MCP tool selection refresh** - Tool selection now refreshes after MCP updates - **Slash command initialization** - Slash commands now wait for initialization to finish before running - **Spec mode approval comments** - Spec approval comments now work in both live and resumed sessions - **`--auto` level in exec runs** - `--auto` level is now applied in exec runs even when the session inherits spec mode - **Tool confirmation state** - Restored correct tool execution state after a tool-confirmation approval - **Settings theme control alignment** - Fixed alignment of the settings theme control - **Stray hover glyphs** - Removed unintended active hover glyphs - **Windows worker reliability** - Added retries for Windows Job Object spawn failures - **Settings folder watcher** - Switched to native recursive watching for the settings folder - **Cron schedules in UTC** - Scheduled Missions runs now evaluate cron expressions in UTC - **Clearer paywall errors on Droid Computer routes** - Droid Computer routes now return a clear error when accessed from a non-paying organization - **Slack multi-turn delegation** - Fixed multi-turn delegation when running from Slack - **Factory App auth timeout fallback** - Fixed an auth timeout fallback in the Factory App - **App UI polish** - Fixed various spacing, color, and chevron issues (Factory App) ## CLI v0.109.1 · Desktop v0.66.1 (April 23, 2026) Fast-mode toggle, Droid Computer improvements, Missions artifact validation, and CLI stability fixes ### New features - **Fast-mode toggle** - Single opt-in toggle to enable fast-mode model variants - **Droid Computer registration improvements** - Smoother register/remove flow with better handling of duplicate Droid Computer names - **Missions artifact validation** - Missions artifacts are now schema-validated on create, edit, and apply-patch ### Improvements - **Faster Windows worker startup** - Reduced worker startup latency on Windows - **Cleaner Ctrl+O transcripts** - Verbose transcript bodies are now hidden for easier reading - **Quieter CLI logs** - Reduced noisy log volume during normal use - **Mermaid rendering** - Various improvements to Mermaid diagram rendering ### Bug fixes - **Spec mode approval flow** - Stabilized the approval flow in spec mode - **IDE diagnostics on disconnect** - Guarded against errors when the IDE client is disconnected - **Background process cleanup** - Background processes are now cleaned up when the CLI exits - **Sessions during silent tools** - Sessions are kept alive during silent tool execution - **Session tool IDs** - Session `enabledToolIds` are now treated as additive - **Secret scanner on large diffs** - Improved secret scanner reliability on large diffs - **`/rename` persistence** - `/rename` now persists across daemon sync and refreshes the cloud title (Factory App) - **Auto-allow-new toggle** - Toggling auto-allow-new no longer clears every model (Factory App) - **Integration modal** - Stabilized the Integration Management modal's height and scrolling (Factory App) - **On-demand pagination** - Replaced "Load All" with proper on-demand pagination (Factory App) - **GitLab integration reliability** - Fixed an out-of-memory issue in the GitLab integration (Factory App) ## CLI v0.108.0 · Desktop v0.65.0 (April 22, 2026) Consolidated Missions files, compute usage on billing pages, and CLI stability fixes ### New features - **Consolidated Missions files** - Files related to Missions now live together under `~/.factory` for easier discoverability - **Compute usage on billing pages** - Billing and usage pages now surface compute usage alongside other usage data (Factory App) ### Bug fixes - **Missions wake-lock on Linux** - Missions wake-lock now also blocks idle and lid-switch suspend so long-running Missions runs aren't interrupted - **Terminal focus feedback loop** - Fixed a terminal focus-in feedback loop in the TUI - **JSON-RPC error handling** - Improved error handling for JSON-RPC communication - **Subagents in sidebar** - Restored subagents in the sidebar (Factory App) ## CLI v0.106.0 · Desktop v0.63.0 (April 21, 2026) New /btw command, upgraded /copy selector, full-project security audit, and CLI stability fixes ### New features - **`/btw` command** - New `/btw` slash command for sending side messages to the current session - **Upgraded `/copy` selector** - Improved selector UI for the `/copy` command - **Full-project security audit** - New full-project audit mode in the security-review skill - **`/copy` slash command** - `/copy` slash command and user message hover actions are now available in the app (Factory App) - **Search by session ID** - Search sessions directly by session ID (Factory App) - **Member role filter** - Filter the organization members list by role (Factory App) ### Improvements - **Timeout error handling** - Improved handling of timeout errors with longer retry delays - **Corporate network error detection** - Corporate firewall, DNS, and TLS errors are now treated as non-retryable - **Context limit classification** - Requests exceeding a model's context window are now classified as context-limit errors with clearer messaging ### Bug fixes - **Subagent permission requests** - Subagent permission requests are now auto-rejected instead of crashing the session - **Auto suggestions after command** - Auto suggestions no longer appear after a command is accepted - **Auth error retries** - 401 and 403 authentication errors are no longer retried - **Default model preserved** - The user's configured default model is no longer rewritten on reads - **Dismissed handoff items** - Dismissed handoff items are now preserved during validation - **Bug report fallback** - Bug reports now fall back to local submission when the background daemon is unavailable - **Factory API key scoping** - The Factory API key list is now scoped to your current organization (Factory App) - **Slack session owner across pages** - The saved Slack Session Owner is now correctly shown across paginated org member lists (Factory App) - **Single pinned session** - Fixed an issue that allowed multiple pinned sessions in the sidebar (Factory App) ## CLI v0.105.1 · Desktop v0.62.0 (April 21, 2026) Spec mode option picker refinements and CLI stability fixes ### Improvements - **Spec mode option selection** - Simplified the flow for choosing between multiple proposal options in spec mode - **`droid computer list` output** - Hides the default "active" status column for cleaner output ### Bug fixes - **Kimi tool call handling** - Fixed an issue with Kimi tool calls that could cause stream errors - **Auto-update on exit** - CLI now waits for in-flight auto-updates to finish before exiting to avoid interrupted installs - **Droid Computer connection reliability** - Better detection and recovery when a Droid Computer connection stalls - **Droid Computer connection warning** - Suppressed a spurious "Failed to get computer connection state" warning (Factory App) - **Onboarding wizard display** - Fixed display issues in the onboarding wizard (Factory App) - **Send button when disconnected** - Send button is now properly disabled in disconnected sessions (Factory App) - **Banner import error** - Fixed a banner import error (Factory App) - **Sidebar grouping collapse** - Fixed unrelated groups collapsing together in the sidebar (Factory App) ## CLI v0.105.0 · Desktop v0.61.0 (April 20, 2026) Kimi K2.6 support, simplify built-in skill, AutoWiki / GitHub Wiki sync, and CLI stability fixes ### New features - **Kimi K2.6 support** - Added support for the Kimi K2.6 model - **Simplify built-in skill** - Simplify is now a built-in skill, no installation required - **AutoWiki / GitHub Wiki sync** - Wikis generated by AutoWiki can now sync to GitHub Wiki, with additional Wiki reader fixes - **Managed Droid Computer enterprise control** - New managed Droid Computer toggle in enterprise controls (Factory App) - **Missions access policy** - Enterprise admins can now control access to Missions - **Session title in browser tab** - Factory App sessions now use the session title as the page title - **Revamped Droid Computer creation** - Redesigned flow for creating Droid Computers (Factory App) - **Nested sidebar** - Sidebar now supports nesting along with other polish improvements (Factory App) - **Droid Computer connection status** - Machine selector now shows the connection status for each Droid Computer (Factory App) ### Improvements - **Compaction model for BYOK** - Compaction now defaults to your current model and is used for BYOK - **PowerShell syntax warning** - Execute tool now detects PowerShell and warns when Bash syntax is used - **Windows path normalization** - Glob tool now returns normalized backslash paths on Windows ### Bug fixes - **Spec option selection** - Fixed the user-selected spec option not being passed to the model - **Execute tool header** - Fixed the last character being clipped in the Execute tool header - **Invalid JSON rendering** - JSON renderer now hides output when the JSON is invalid - **Review guidelines skill** - CLI now handles a missing review-guidelines skill gracefully - **Tool cancel interrupts** - Follow-up prompts now wait for tool-cancel interrupts to settle - **TUI daemon session switching** - Previous TUI daemon session now closes only after the next one loads - **Session title updates** - Daemon session title updates now reflect in the terminal tab - **Missions worker permissions** - Missions worker permission requests are now auto-denied - **Thinking traces in web chat** - Thinking traces now appear in session chat on the web - **Ripgrep startup** - Fixed intermittent startup errors when initializing the bundled ripgrep binary - **Truncated stream handling** - Improved error handling when model response streams are truncated - **Stuck "Show more"** - Fixed a stuck "Show more" button after creating a session on a new Droid Computer (Factory App) - **Slack session-owner picker** - Picker now uses server-side search with infinite scroll (Factory App) - **Sidebar hover and selection** - Fixed hover states and highlighting in the sidebar (Factory App) - **401 errors when logged out** - Suppressed 401 errors after logging out (Factory App) - **Stuck loading state** - Fixed a stuck loading state in the app (Factory App) - **Add Project modal directory** - Modal now picks up the current directory from the native browse dialog (Factory App) - **Local machine option** - Hidden outside of the Factory App ## CLI v0.104.0 · Desktop v0.60.0 (April 17, 2026) Custom ripgrep path, BYOK config in bug reports, and CLI stability fixes ### New features - **Custom ripgrep path** - New environment variable to point Factory at a custom ripgrep binary - **BYOK config in bug reports** - Bug reports now include resolved BYOK configuration to make troubleshooting easier ### Improvements - **Full Execute command in approval** - The approval dialog now shows the full command that will run - **Selected confirmation option visibility** - The selected option in tool confirmation prompts is now bolded for better visibility - **Hidden daemon log output** - Background daemon output no longer appears in the session ### Bug fixes - **Spec edits in confirmation flow** - Fixed spec edits being dropped during the confirmation flow - **Manual spec edits respected** - Fixed manual edits to specs not being preserved during handoff - **Image upload compression** - Fixed image upload compression issues - **Assistant active timer** - Fixed the assistant active timer not persisting across unmounts - **Daemon connection and bug reports** - Fixed daemon connection issues and bug report errors (Factory App) - **Local reconnect on machine change** - Local machine reconnect now works reliably when switching connections, with a clearer warning when not connected (Factory App) - **Tool rendering parity** - Improved DefaultTool and TaskTool rendering to match the CLI (Factory App) - **ApplyPatch diff rendering** - Fixed ApplyPatch tool diff rendering for the CLI input format (Factory App) ## CLI v0.103.0 · Desktop v0.59.0 (April 16, 2026) Sessions page redesign, dead code detection in reviews, and CLI stability fixes ### New features - **Sessions page redesign** - The sessions page has been refreshed with a new layout and design (Factory App) ### Improvements - **Denied command confirmation** - Redesigned the confirmation UI when running blocked or risky commands - **Dead code detection in reviews** - Code review now detects dead and unused code - **QA report formatting** - QA skill now includes structured report format sections ### Bug fixes - **File overwriting** - File creation and patch tools can now properly overwrite existing files - **Model selection after spec approval** - Fixed the session model not being restored after approving a spec - **Inactive session reload** - Fixed inactive TUI sessions not reloading on demand - **Parallel tool calls in spec mode** - Fixed tool calls running in parallel incorrectly during spec approval - **Skill details refresh** - Fixed selected skill details not updating when the skills list changes (Factory App) ## CLI v0.102.0 · Desktop v0.58.0 (April 15, 2026) Security review checks, AutoWiki images, and CLI stability fixes ### New features - **Security review checks** - Code review now includes security checks for OWASP Top 10 vulnerabilities, LLM-specific risks, and supply chain issues - **AutoWiki images** - AutoWiki pages now support and display embedded images ### Improvements - **Max autonomy enforcement** - App now properly enforces max autonomy settings (Factory App) - **Large PR review support** - CLI now gracefully handles large PR diffs by using local git diff as a fallback ### Bug fixes - **Spec approval comments** - Fixed spec approval comments not being preserved during handoff - **Tool result spacing** - Fixed excessive margin between tool header and result in the CLI - **Mermaid diagram rendering** - Fixed rendering issues with mermaid diagrams on the web - **Personal integrations visibility** - Fixed personal integrations not showing for regular users (Factory App) ## CLI v0.101.0 · Desktop v0.57.0 (April 14, 2026) New /install-slack-app command, spinner redesign, and CLI stability fixes ### New features - **`/install-slack-app` command** - New `/install-slack-app` command to set up the Slack app integration for your workspace ### Improvements - **CLI spinner redesign** - Updated CLI spinner animation for a refreshed look - **Usage Limit warning** - Warning banner now shown when setting a global Usage Limit (Factory App) - **Droid Computer source control warning** - Warning displayed when creating a Droid Computer without source control configured (Factory App) - **Usage Limits clarity** - Usage Limits now clearly display values in millions of Factory Standard Credits (Factory App) ### Bug fixes - **File tool accuracy** - File tools now use the actual filesystem state for more reliable results - **File edit previews** - Fixed structured previews not displaying correctly for file edit results - **Duplicate spec sessions** - Fixed duplicate sessions being created during spec handoff ## CLI v0.100.0 · Desktop v0.55.0 (April 13, 2026) API key creation now available for all organization members ### Improvements - **API key creation for all members** - All organization members can now create API keys, previously limited to managers and owners (Factory App) ## CLI v0.99.0 · Desktop v0.54.0 (April 10, 2026) Incremental AutoWiki refreshes, themed mermaid diagrams, AutoWiki search, and CLI stability fixes ### New features - **Incremental AutoWiki refreshes** - AutoWiki refreshes now only update changed pages for faster updates - **Themed mermaid diagram colors** - Mermaid diagrams now render with per-element colors that match your current theme - **AutoWiki search and exports** - AutoWiki now supports keyword search, skill browsing, and page exports (Factory App) ### Bug fixes - **Theme background override** - Fixed terminal theme background being overridden unexpectedly - **BYOK base URL switching** - Fixed client connections not being refreshed after switching custom API base URLs - **Context model inheritance** - Fixed context management not using the current session model - **Spec editor terminal handoff** - Fixed spec editor not properly completing terminal handoff - **MCP auth cancellation** - Fixed pending MCP authentication not being cancelled when closing the MCP panel - **Tool result display** - Fixed a crash when tool results contain non-string values - **Custom model stability** - Fixed custom model IDs changing unexpectedly between sessions - **Directory not found error** - Directory not found errors now appear immediately when confirming a directory (Factory App) - **BYOK Rate Limits** - BYOK users can now continue chatting when platform Rate Limits are reached (Factory App) - **Ghost tool calls** - Fixed phantom tool calls appearing during message processing - **Chat input** - Fixed various chat input issues (Factory App) - **Droid Computer terminal** - Fixed terminal display issues when using Droid Computers (Factory App) - **Machine config modal** - Fixed machine configuration modal not displaying correctly (Factory App) - **Reasoning and model selection** - Fixed reasoning level and model selection issues (Factory App) ## CLI v0.98.0 · Desktop v0.53.0 (April 9, 2026) Install code review skill, session renaming, Spotlight search, and CLI stability fixes ### New features - **Install code review skill** - New built-in skill for setting up automated code review in your project - **Session renaming** - Rename sessions directly from the app (Factory App) - **macOS Spotlight discoverability** - Factory App is now easier to find via macOS Spotlight search ### Improvements - **Install QA skill upgrade** - The install-qa skill now learns from test failures to improve over time - **Faster AutoWiki loading** - Optimized AutoWiki page loading for reduced latency (Factory App) ### Bug fixes - **Terminal palette override** - Fixed terminal color palette being overridden unexpectedly - **Droid Shield SSH key detection** - Fixed Droid Shield incorrectly flagging SSH public keys as secrets - **Droid Computer sessions display** - Fixed sessions not appearing when run on Droid Computers without a local directory (Factory App) - **Droid Computer SSH** - Fixed SSH connectivity issues with Droid Computers - **Tool status text** - Restored the "invoking tools" status indicator during tool execution - **Escape key after spec handoff** - Fixed Escape key cancellation not working after spec approval handoff - **Session sidebar** - Fixed various sidebar display issues (Factory App) - **Certificate loading** - Fixed certificate loading issues that could affect connectivity - **Chat and tool results** - Fixed display issues with chat messages and tool results (Factory App) - **Local machine connection toggle** - Fixed the local machine connection toggle not being respected when turned off (Factory App) ## CLI v0.97.0 · Desktop v0.52.0 (April 8, 2026) New /cwd command, GLM-5.1 model support, diff viewer for Droid Computers, and CLI stability fixes ### New features - **`/cwd` command** - New `/cwd` command and `--cwd` startup flag to set or change the working directory for your session - **GLM-5.1 model support** - Added GLM-5.1 as an available model option - **Diff viewer for Droid Computers** - Added diff viewer support for viewing file changes in Droid Computers (Factory App) ### Improvements - **Execute approval display** - Pending Execute approvals are now truncated for readability while preserving the full view via Ctrl+O ### Bug fixes - **AskUser layout** - Fixed alignment and spacing issues in the AskUser prompt (Factory App) - **Theme consistency on startup** - Fixed terminal theme overrides not being applied consistently on startup - **Kitty terminal over SSH** - Fixed Ctrl+D handling in Kitty terminal when connected over SSH - **Session settings** - Fixed settings not being applied correctly when creating a new session ## CLI v0.96.0 · Desktop v0.51.0 (April 7, 2026) New /install-qa command, design.md support, Cmd+N for new sessions, and CLI stability fixes ### New features - **`/install-qa` command** - New `/install-qa` command to generate automated QA skills for your project - **`design.md` support** - CLI now supports `design.md` files alongside `agents.md` for providing design guidelines to your droid - **Cmd+N for new session** - Use Cmd+N keyboard shortcut to quickly start a new session (Factory App) - **Session disambiguation** - Sessions with the same name now display the full path in a tooltip for easier identification (Factory App) - **Pre-session bug reports** - Submit bug reports from the Factory App even when no session is active ### Improvements - **Bash mode output redesign** - Bash mode output now matches the execute tool display for a more consistent look - **MCP tool confirmation display** - Restyled MCP tool display in the confirmation popup for better readability - **Light mode colors** - Improved color consistency in light mode by migrating hardcoded colors to the theme system - **Archived session access** - Archived sessions can now be viewed in read-only mode and restored when accessed directly (Factory App) - **Consistent hover states** - All option items now have consistent hover state styling (Factory App) - **Factory App UI polish** - Various visual refinements across the Factory App ### Bug fixes - **Duplicate tool call messages** - Fixed duplicate tool call messages and missing updates during tool execution - **Missions feature counts** - Fixed incorrect feature counts in `/missions` - **Escape key handling** - Fixed prompt handling issues after pressing Escape - **Terminal input security** - Hardened terminal input capture for improved security - **Session creation** - Fixed an issue where sessions could be created with an empty working directory (Factory App) - **Profile dropdown styling** - Refined profile dropdown and menu item styling (Factory App) - **Font size consistency** - Fixed font size mismatch in the Droid Computers section (Factory App) - **Sidebar layout** - Fixed session group header to span the full sidebar width (Factory App) ## CLI v0.95.0 · Desktop v0.50.0 (April 6, 2026) Redesigned help menu, green progress bars, faster session reload, and CLI stability fixes ### New features - **Redesigned help menu** - The help menu has been completely revamped with improved layout and organization - **Green progress bars** - Progress bars now display in green while commands are running ### Improvements - **Faster session reload** - Sessions reload significantly faster for a smoother experience - **CLI performance** - General performance improvements for faster CLI responsiveness - **Light mode color refinements** - Improved color palette and theme refinements for light mode (Factory App) - **Scrollable session view** - The entire session view is now scrollable for easier navigation (Factory App) - **Theme-aware menu colors** - Selection colors in slash command menus now adapt to the current theme - **Keyboard shortcut footer** - Improved keyboard shortcut display in slash command menus - **Back navigation** - Secondary pages now show a cleaner "Back to App" navigation button (Factory App) ### Bug fixes - **Execute tool name display** - Fixed the Execute tool name being truncated in pending state - **Session freeze prevention** - Fixed session freezes caused by interactive commands in Bash Mode and Execute - **Tool execution after interrupts** - Tools no longer continue executing after being interrupted - **Subagent working directory** - Fixed subagents not inheriting the correct working directory - **ANSI escape handling** - Fixed ANSI escape sequences leaking into the input prompt - **Todo list display** - Fixed visual issues in the todo list display - **Custom model validation** - Fixed validation errors when using custom models in Missions workers - **Mission Control model selector** - Fixed the model selector in Mission Control - **Exec mode reasoning output** - Reasoning events now properly appear in stream-json exec output - **Sidebar collapse on Windows** - Fixed sidebar collapse button location on Windows (Factory App) - **Loading spinner** - Fixed loading spinner display issues (Factory App) - **Chat layout spacing** - Fixed left margin alignment in chat composer and content area (Factory App) ## CLI v0.94.0 · Desktop v0.49.0 (April 3, 2026) /limits usage panel, credit usage countdown, model selector expansion, and CLI stability fixes ### New features - **/limits usage panel** - New `/limits` command shows an interactive credit usage breakdown with bucket views - **Credits Balance countdown** - Live countdown timer shows remaining Credits in real time (Factory App) - **Model selector expansion** - Model selector now includes a "show more" option to browse additional models (Factory App) ### Improvements - **Queued message redesign** - Queued messages now display with a cleaner, user-message-style layout - **/plugins and /skills menu refresh** - `/plugins` and `/skills` menus now match the updated `/sessions` and `/model` UI style - **Droid Shield placeholder detection** - Droid Shield now detects additional placeholder value patterns to better protect against accidental secret exposure ### Bug fixes - **Worker hang fix** - Fixed a regression where blocked workers could cause the CLI to hang - **Custom model streaming** - Fixed streaming errors when using custom models - **Long prompt handling** - Fixed an issue where long prompts could fail when passed to subagents - **Autonomy level cycling** - Fixed Ctrl+L not correctly cycling through autonomy levels - **Kitty terminal fix** - Fixed additional input handling issues in Kitty terminal ## CLI v0.93.0 · Desktop v0.48.0 (April 2, 2026) GPT-5.3-Codex fast mode, markdown links, automatic agents.md detection, and session retention policies ### New features - **GPT-5.3-Codex fast mode** - Added GPT-5.3-Codex as a fast mode model option for faster code generation - **Markdown link rendering** - Markdown links now render as clickable links in the CLI - **Automatic agents.md detection** - CLI now automatically detects and loads agents.md files in your project - **Session retention policies** - Configure auto-deletion policies for sessions to manage storage (Factory App) ### Improvements - **Plan update indicator** - A subtle "Plan updated" notification now appears when the plan changes - **JSON rendering in chat** - JSON data now renders with improved formatting in web and Factory App chat - **Commands near autonomy requests** - Suggested commands now appear closer to autonomy requests for easier access (Factory App) ### Bug fixes - **Review shows LGTM** - The review command now shows "LGTM" when no issues are found - **Long output truncation** - Long output lines in Execute tool previews are now truncated to fit terminal width - **Message ghosting** - Fixed ghost messages appearing when sending queued messages - **Improved error messages** - Model streaming errors now display user-friendly messages instead of raw errors - **Session loading** - Fixed various issues with loading and resuming sessions - **Onboarding link** - Fixed a broken link in the onboarding flow (Factory App) - **Mission Mode input** - Fixed input field behavior in Mission Mode (Factory App) ## CLI v0.92.0 · Desktop v0.47.0 (April 2, 2026) Unified Missions menu, settings and footer redesign, session forking, TUI charts, and SSH clipboard fixes ### New features - **Unified /missions menu** - Redesigned the `/missions` command with model tabs and improved navigation - **/settings redesign** - Refreshed `/settings` menu with an updated layout - **CLI footer redesign** - Redesigned CLI footer with a cleaner layout - **Per-worker token breakdown** - Mission Control now shows token usage breakdown per worker - **Session forking** - New `--fork` flag to create a new session by forking an existing one - **TUI chart rendering** - Charts and graphs now render directly in the terminal - **Exec mode system prompt flags** - New `--append-system-prompt` and `--append-system-prompt-file` flags for customizing system prompts in exec mode - **Cloud-synced read-only sessions** - Cloud-synced sessions are now viewable in read-only mode (Factory App) - **Configurable per-model context limits** - New setting to configure context token limits on a per-model basis ### Improvements - **Session loading performance** - Optimized session loading for faster startup and reduced load times - **Faster session search** - Session search is now significantly faster (Factory App) - **Faster Factory App session loading** - Sessions load directly from disk in the Factory App for improved speed - **Tool routing warnings** - Tool routing warnings now suggest model-specific alternatives - **Default model indicator** - `/model` command now highlights which model is the default - **SKILL.md support** - Built-in skills can now be loaded from SKILL.md files - **Sidebar polish** - Modernized sidebar user profile and dropdown animations (Factory App) ### Bug fixes - **Clipboard copy over SSH** - Clipboard copy now works reliably over SSH connections - **VS Code multiline input** - Fixed multiline input handling in VS Code terminal - **MCP authentication** - Restored MCP OAuth authentication flows - **PDF page limit error** - Added a clear error message when exceeding the 100-page PDF limit - **Terminal cleanup on exit** - Fixed terminal escape sequences persisting after CLI exit - **Slash command settings** - Fixed `/fast`, `/enter-mission`, `/exit-mission`, and `/new` commands not applying settings correctly - **Parallel tool rendering** - Fixed rendering issues with parallel tool executions - **Ghostty and xterm input** - Fixed keyboard input handling in Ghostty and xterm terminals - **Terminal multiplexer display** - Fixed screen clearing in terminal multiplexers like tmux and screen - **Terminal focus recovery** - Fixed terminal focus not being recovered after window switching - **AGENTS.md session loading** - AGENTS.md guidelines now load correctly when resuming existing sessions - **Task tool background panel** - Fixed streaming and display issues in Task tool background panels - **Fullscreen sidebar layout** - Fixed sidebar layout in fullscreen mode on macOS (Factory App) ## CLI v0.90.1 · Desktop v0.45.1 (March 30, 2026) Light theme readability fix ### Bug fixes - **Light theme readability** - Fixed user message colors and help bar shortcuts in the light theme for improved contrast and readability ## CLI v0.89.0 · Desktop v0.44.0 (March 27, 2026) Spec mode redesign, Missions worker view, fuzzy model search, and cleaner tool output ### New features - **Missions worker view** - New view for monitoring individual Missions worker progress and output ### Improvements - **Spec mode redesign** - Redesigned spec mode with a refreshed layout and improved workflow - **Fuzzy search in model selector** - Model selector now supports fuzzy search for easier model filtering - **Cleaner tool output** - Grep and Glob tool calls now display with the same clean formatting as Read tool calls - **Plan display polish** - Improved plan display formatting and cleaner worker output - **Changelog dismissal** - Changelog now auto-hides after first view and can be manually dismissed - **Billable Tokens renamed to Factory Standard Credits** - "Billable Tokens" renamed to "Factory Standard Credits" throughout the UI - **Auto-updating plugins** - Installed plugins now auto-update on CLI startup ### Bug fixes - **Missions worker stability** - Permission blocks in Missions workers no longer terminate the session - **/exit-mission autocomplete** - `/exit-mission` autocomplete now only appears when actively in a mission - **Empty message handling** - Fixed empty messages appearing in sessions ## CLI v0.87.0 · Desktop v0.42.0 (March 26, 2026) Spec approval handoff, shimmer animations for tool headers, and Missions reliability improvements ### New features - **Spec approval handoff** - Spec approvals can now be handed off to a new session, letting you continue work separately after approving a spec ### Improvements - **Shimmer animation for tool headers** - Tool call headers now display a shimmer animation for improved visual feedback ### Bug fixes - **Session storage reliability** - Fixed duplicate save errors and added automatic recovery from permission issues in session storage - **Sidebar navigation** - Fixed missing back arrow in sidebar menu titles (Factory App) - **Missions worker startup** - Improved Missions worker startup reliability with extended timeout - **Missions connection stability** - Increased socket connection timeout for more reliable Missions connections - **Droid Computer session creation** - Fixed errors when creating sessions while the Droid Computer is temporarily unavailable (Factory App) ## CLI v0.86.0 · Desktop v0.41.0 (March 25, 2026) /context slash command, automation sharing and forking, OS-level sandboxing, and terminal fixes ### New features - **/context slash command** - New `/context` command shows a breakdown of token usage in your current session - **OS-level sandboxing** - Added OS-level sandboxing for enhanced security when running CLI tools ### Improvements - **Ask question tool redesign** - Redesigned the Ask Question tool with improved text wrapping and layout ### Bug fixes - **Kitty terminal keypad keys** - Fixed keypad keys not working correctly in Kitty terminal - **Slack DM reliability** - Fixed bot message feedback loops in Slack DMs ## CLI v0.85.0 · Desktop v0.40.0 (March 24, 2026) Session settings override, SSE for remote MCP servers, sidebar animations, and reliability fixes ### New features - **Session settings override** - Project and folder-level settings can now override organization session defaults - **SSE support for remote MCP servers** - Remote MCP servers now support Server-Sent Events (SSE) transport ### Improvements - **Session group animations** - Session groups now animate smoothly when collapsing and expanding (Factory App) - **AutoWiki sidebar polish** - Improved AutoWiki sidebar design with smoother pan transitions (Factory App) - **Skills browser scrolling** - Improved scrollbar styling in the Skills browser (Factory App) - **Managed settings error messages** - Clearer error messages when loading managed settings ### Bug fixes - **Session titles** - Fixed session titles intermittently displaying incorrectly (Factory App) - **Missions worker BYOK settings** - Fixed Missions worker settings not persisting correctly when using custom API keys - **Session ID display** - Fixed session IDs being truncated (Factory App) - **Background process reliability** - Hardened background process and task store against data corruption - **Spec mode file edits** - Fixed file edits to .factory/specs being blocked during spec mode - **Chat message alignment** - Fixed user message text not aligning with droid answers - **Directory selector stability** - Fixed directory selector jiggling when switching directories (Factory App) - **Keyboard shortcuts** - Restored keyboard shortcuts in the app (Factory App) - **Session list display** - Fixed session list not showing enough sessions - **Droid Computer connection reliability** - Improved timeouts for initial Droid Computer connections (Factory App) - **Unknown model handling** - Fixed handling of unknown model IDs ## CLI v0.84.0 · Desktop v0.39.0 (March 23, 2026) Header redesign, Execute tool syntax highlighting, session type tags, and skill execution fixes ### New features - **Header redesign** - Refreshed CLI header with an updated layout for improved readability - **Execute tool syntax highlighting** - Execute tool calls now display with syntax highlighting for better readability - **Session type tags** - Sessions now display type tags for Missions and subagents for easier identification - **Usage Limits settings UI** - New interface for viewing and managing Usage Limits settings (Factory App) ### Improvements - **Create tool display** - Improved text wrapping and syntax highlighting for Create tool results - **CLI onboarding** - Improved onboarding flow for new CLI users ### Bug fixes - **Semantic diff rendering** - Fixed a random character appearing at the top of semantic diffs - **Skill execution** - Fixed skills not executing properly after being loaded - **Gemini BYOK thinking** - Fixed thinking signature preservation for Gemini models with custom API keys ## CLI v0.83.0 · Desktop v0.37.0 (March 20, 2026) Redesigned model selector and session list, active tool call spinner, and skill slash command fixes ### New features - **Redesigned model selector and session list** - Overhauled model selector and session list with improved navigation and layout - **Active tool call spinner** - A spinner now displays during active tool calls for better visibility into CLI activity - **Configurable subagent inactivity timeout** - New `subagentInactivityTimeout` setting to customize how long subagents wait before timing out due to inactivity ### Bug fixes - **Diff syntax highlighting** - Fixed syntax highlighting issues in diff display - **Skill slash command arguments** - Fixed user arguments not being passed to skills invoked via slash commands - **CLI status orb display** - Fixed random display issues with the CLI status orb - **Double loader on new session** - Fixed a duplicate loading indicator appearing when creating a new session (Factory App) - **Sidebar button accessibility** - Fixed sidebar buttons being covered by the draggable region (Factory App) - **Whitespace in working directory input** - Input for working directory now correctly trims whitespace (Factory App) - **Daemon connectivity** - Fixed daemon connection issues that could occur on some systems ## CLI v0.82.0 · Desktop v0.36.0 (March 19, 2026) Syntax highlighting for diffs, theme engine, file uploads, and session settings fixes ### New features - **Syntax highlighting for diffs** - Diffs displayed by the CLI now include syntax highlighting for improved readability - **Theme engine** - New customizable theme engine for personalizing CLI colors - **File upload support** - Upload CSV, Excel, and DOCX files directly in conversations (Factory App) - **MCP and skills controls in new sessions** - MCP servers and skills can now be configured when creating new sessions - **Sidebar redesign** - Revamped sidebar with improved navigation and layout (Factory App) - **Droid Computer info in session tooltip** - Session tooltips now display connected Droid Computer details (Factory App) ### Bug fixes - **Session settings preserved on fork** - Forking a session now preserves your session settings instead of resetting to global defaults - **Mission Control stability** - Fixed a memory leak in Mission Control - **Auto-update rollback** - Fixed rollback failures during CLI auto-updates - **Mission Control resize** - Fixed Mission Control display issues when resizing the terminal ## CLI v0.80.0 · Desktop v0.34.0 (March 18, 2026) Couchbase MCP server, AutoWiki dashboard links, and sub-agent reliability fixes ### New features - **Couchbase MCP server** - Added Couchbase to the MCP server registry for database integration - **AutoWiki dashboard link** - The Droid CLI now shows a direct link to the dashboard after AutoWiki uploads complete - **Touch device controls** - Submit and cancel buttons now appear on touch devices for easier interaction (Factory App) - **Droid Computer install commands** - The Connect to Droid Computer modal now shows curl and Windows install commands (Factory App) - **Dynamic loading status setting** - New setting to disable dynamic loading status messages (Factory App) ### Improvements - **SSH error messages** - Improved error messages for the SSH command - **BYO Droid Computer setup guidance** - Better user guidance for BYO Droid Computer setup commands - **Simplified Droid Computer setup** - Streamlined the Droid Computer setup flow ### Bug fixes - **Sub-agent spawning** - Fixed sub-agents failing to spawn correctly - **Esc key during tool execution** - Fixed Esc key during tool execution incorrectly interrupting the next user message - **Model error notifications** - Model API errors are now shown as user notifications instead of being hidden - **Daemon startup reliability** - Fixed daemon startup failures caused by insufficient timeout on slower connections - **Usage Limits display** - Fixed Usage Limits not displaying correctly for users (Factory App) - **Droid Computer card navigation** - Fixed click navigation on Droid Computer name cards (Factory App) - **Debug mode logging** - Fixed debug logs not appearing when running with the --debug flag ## CLI v0.77.0 · Desktop v0.31.0 (March 17, 2026) GPT-5.4 Mini model, faster @ search indexing, and sub-agent improvements ### New features - **GPT-5.4 Mini model** - New GPT-5.4 Mini model available in the model selector ### Improvements - **Faster @ search indexing** - File indexing for @ mentions now runs in a background worker for improved responsiveness - **Faster session listing** - The `/sessions` command is now optimized with an index for quicker loading - **Faster first prompt** - System info is now prefetched before the first prompt for quicker startup - **Sub-agent reliability** - Redesigned sub-agent architecture for improved task delegation reliability - **Skills preserved after compaction** - Skills invoked during a session are now preserved when context is compacted - **Droid Computer selection UI** - Improved visual design for the Droid Computer selection interface (Factory App) ### Bug fixes - **Slash command arguments** - Fixed slash command arguments being lost when submitting via the suggestion menu - **Missions model selection** - Missions now correctly uses the session model instead of the global settings model - **Worktree error messages** - Worktree setup failures now show descriptive error messages instead of generic errors - **Stale Missions warnings** - Fixed stale Missions runs incorrectly triggering concurrent run warnings - **Mode preservation on /new** - Fixed session mode being reset when creating a new session with `/new` - **Permission settings persistence** - Permission settings are now preserved across session respawns - **Chat composer undo** - Fixed undo not working after pasting text in the chat composer (Factory App) ## CLI v0.76.0 · Desktop v0.30.0 (March 16, 2026) Configurable LLM timeout, MCP server allowlist for Enterprise Controls, and Mission Control improvements ### New features - **Configurable LLM timeout** - LLM API timeout is now configurable via settings.json for customizing request timeout behavior - **MCP server allowlist** - Enterprise admins can now restrict which MCP servers are allowed via a new allowlist policy in Enterprise Controls (Factory App) - **Granola MCP server** - Added Granola to the MCP server registry for meeting notes integration ### Improvements - **Mission Control shortcut** - Mission Control now toggles with Ctrl+T for quicker access - **Rate Limit display** - Improved UI for Rate Limit notifications - **Droid Computer error messages** - Clearer error messages for Droid Computer commands - **Mission Control performance** - Reduced memory usage and latency in Mission Control - **Session ID tooltip** - Session ID now visible in the session tooltip for easier reference (Factory App) ### Bug fixes - **Custom model lookup** - Fixed BYOK custom model lookup failing after settings changes - **Model selection text color** - Fixed incorrect text color in the model selection dropdown (Factory App) - **Enterprise model defaults** - Cloud sessions now correctly respect enterprise model defaults (Factory App) - **Model policy overrides modal** - Fixed display issues in the model policy overrides modal (Factory App) - **Automations page layout** - Fixed automations page height in the Factory App - **MCP OAuth connection** - Improved retry handling for MCP OAuth callback connections - **Organization settings refresh** - Organization-managed settings now correctly refresh after login ## CLI v0.75.0 · Desktop v0.29.0 (March 13, 2026) Bug report performance fix and connection stability improvements ### Bug fixes - **Bug report performance** - Fixed bug reports hanging when processing large Missions runs - **Disappearing CLI line** - Fixed a line occasionally disappearing from CLI output - **Connection stability** - Improved handling of disconnection errors for more stable sessions ## CLI v0.74.0 · Desktop v0.28.0 (March 12, 2026) Configurable worktree directory, Grep multiline auto-detection, and Mission Control improvements ### New features - **Configurable worktree directory** - New setting to customize the directory used for git worktree operations ### Improvements - **Grep multiline auto-detection** - Grep tool now automatically enables multiline mode when search patterns contain newlines - **npm installation guidance** - CLI now provides clear instructions to add droid to your PATH after npm installation - **Retry reliability** - Improved retry handling for transient provider errors with automatic rotation - **System info truncation** - Large directory listings and git status output are now truncated to prevent system info overflow ### Bug fixes - **Thinking block preservation** - Invalid thinking blocks are now preserved as text instead of being silently dropped - **Git subdir plugin support** - Fixed support for git subdirectory plugins in project configurations - **Mission Control UX** - Fixed various Mission Control UI and interaction issues - **Spec mode subagents** - Subagents now correctly use read-only mode when the parent session is in spec mode - **MCP OAuth compatibility** - Improved MCP OAuth compatibility across servers - **Playwright MCP permissions** - Improved permission handling when using the Playwright MCP server - **Terminal tab title on resume** - Terminal tab title now correctly updates when resuming a session - **Concurrent session handling** - Improved detection and handling of concurrent CLI instances - **Code server connection** - Fixed code server connection issues in the Factory App ## CLI v0.73.0 · Desktop v0.27.0 (March 11, 2026) Emacs keybindings, remote access flag, faster response times, and reliability improvements ### New features - **Additional Emacs keybindings** - Added more Emacs-style keyboard shortcuts for text editing in the CLI - **`--remote-access` flag** - Remote session access now requires the explicit `--remote-access` flag for improved security ### Improvements - **Faster time to first token** - Reduced time to first response by ~8% through lazy loading and caching optimizations - **Session reconnection reliability** - Pending requests are now preserved and replayed when reconnecting after a disconnection (Factory App) ### Bug fixes - **Spec mode model handling** - Fixed spec mode using the wrong model for thinking block processing - **Auto-update reliability** - Fixed auto-update failures on read-only directories and Windows updates getting stuck in a pending state - **Mission Control layout** - Fixed layout stability issues in Mission Control - **Tool confirmation double-click** - Fixed tool confirmation buttons firing twice on a single click - **Tool format tolerance** - Improved tolerance for format variations in AskUser and TodoWrite tool responses - **Unicode character handling** - Fixed crashes caused by broken Unicode characters in messages - **MCP OAuth security** - Improved OAuth state verification for MCP server connections - **Multiple CLI instance isolation** - Fixed interference between multiple concurrent CLI instances - **Tool filtering in exec mode** - Fixed incorrect tool filtering when running in exec mode - **Model provider switching** - Fixed issues when switching between model providers mid-conversation - **Worktree branch handling** - Fixed error when using worktrees with an already checked-out branch - **Droid Computer connect modal** - Fixed VS Code config display in the Droid Computer connect modal (Factory App) ## CLI v0.72.0 · Desktop v0.26.0 (March 10, 2026) GPT-5.4 Fast model, /fast command, chat UI refresh, and MCP OAuth for remote sessions ### New features - **GPT-5.4 Fast model** - New GPT-5.4 Fast model available with a `/fast` slash command for quick model switching - **MCP OAuth in remote sessions** - MCP servers requiring OAuth authentication now work in remote Droid Computer sessions - **Chat UI refresh** - Refreshed chat interface with improved tool display, user message styling, and interrupt indicators - **Mission subagent streaming** - Real-time streaming updates from subagents in mission control ### Improvements - **Incremental rendering** - Enabled incremental rendering for smoother and more responsive CLI output - **Faster Droid Computer sessions** - Optimistic preconnection reduces wait time when starting Droid Computer sessions (Factory App) - **Working directory validation** - Droid Computer sessions now validate working directory configuration before starting (Factory App) ### Bug fixes - **Grep tool overflow** - Fixed Grep tool crashing on very large search results by returning partial results with an increased buffer - **Thinking indicator spacing** - Fixed spacing below the thinking indicator display - **Timeout warnings** - Fixed spurious timeout warnings appearing during normal operation - **Sentry integration reliability** - Improved reliability of the Sentry integration connection - **Slack integration reliability** - Improved error handling for Slack integration connections - **Mission Mode reasoning** - Fixed reasoning issues in Mission Mode ## CLI v0.71.0 · Desktop v0.25.0 (March 9, 2026) Autofix PR button, worktree naming, and Create tool syntax highlighting ### New features - **Autofix PR button** - Automatically fix issues found in pull requests directly from the app (Factory App) - **MCP OAuth configuration** - Added OAuth configuration support for customizable fields in MCP server settings ### Improvements - **Worktree naming** - The `--worktree` flag now accepts an optional name argument; the `--branch` flag has been removed - **Create tool syntax highlighting** - Create tool results now display with syntax highlighting instead of diff format ### Bug fixes - **Semantic diff rendering** - Fixed diff rendering in semantic diff view - **@ search prewarming** - Fixed @ search prewarming for improved file suggestion performance - **Factory App download button** - Fixed download Factory App button incorrectly showing in the Factory App - **Terminal creation flash** - Fixed "(exited)" text briefly flashing when creating new terminals (Factory App) ## CLI v0.70.0 · Desktop v0.24.0 (March 6, 2026) Git worktree support, native PDF reading, and runtime settings overlay ### New features - **Git worktree support** - New `--worktree` flag enables native git worktree support for working across multiple branches simultaneously - **Native PDF support** - The Read tool now natively supports reading PDF files across all model providers - **Runtime settings overlay** - New `--settings` flag allows applying settings at runtime with organization-precedence merge - **Skills UI for remote sessions** - Skills UI is now available in Droid Computer sessions (Factory App) ### Improvements - **Mission Control redesign** - Refreshed Mission Control interface with a cleaner layout ### Bug fixes - **Option+Arrow word jumping** - Restored Option+Arrow word jumping in ESC-prefix terminals - **Background process isolation** - Background process tracking is now isolated per CLI instance to prevent cross-instance interference - **@ search indexing** - Fixed file indexing in @ search suggestions - **Shell environment in daemon** - Daemon now correctly inherits shell environment variables - **Terminal compatibility** - Fixed Kitty keyboard protocol detection that could cause issues in certain terminals - **Streaming cancellation** - Improved reliability when cancelling streaming responses - **Daemon auto-updates** - Fixed an issue where daemon auto-updates could be incorrectly skipped ## CLI v0.69.0 · Desktop v0.23.0 (March 5, 2026) CJK internationalization, GPT-5.4 model updates, and --mission flag for droid exec ### New features - **CJK internationalization** - Full internationalization support for Japanese, Chinese, and Korean languages in the CLI - **GPT-5.4 model updates** - GPT-5.4 is now a default model option with expanded context limits - **`--mission` flag for droid exec** - New `--mission` flag enables non-interactive Missions orchestration via `droid exec` - **Personal analytics page** - View your personal usage analytics directly in the app (Factory App) - **Settings audit trail** - Track changes to settings values with a full audit trail (Factory App) - **Version history for Enterprise Controls** - View configuration change history in Enterprise Controls (Factory App) - **Combined billing summary** - Billing summary now available in both personal and organization settings views (Factory App) ### Improvements - **Git config updates** - Git configuration can now be updated when you explicitly request it - **BYOK Anthropic thinking** - Improved adaptive thinking support for custom Anthropic models - **Smarter retry handling** - Improved retry delays for provider capacity errors - **LLM timeout errors** - Clearer user-facing error messages when model requests time out - **MCP tool rendering** - Improved MCP tool result display in the Factory App ### Bug fixes - **VSCode keyboard shortcuts** - Fixed support for Cmd + key combinations in VSCode terminal - **Anthropic streaming stability** - Improved streaming reliability for Anthropic models - **Empty model responses** - Empty model responses are now handled gracefully instead of causing errors - **Thinking recovery** - Auto-recovery from thinking signature validation errors during reasoning - **Spec mode navigation** - Fixed back navigation from autonomy menu in spec mode - **Mission session cleanup** - Fresh session now starts correctly when exiting a proposed mission - **Gemini error handling** - Improved error handling for authentication and payment errors with Gemini models - **Profile display** - Fixed user profile parsing for accounts with incomplete name fields ## CLI v0.68.1 (March 5, 2026) VSCode Kitty protocol support and authentication fixes ### New features - **VSCode Kitty protocol support** - Added Kitty keyboard protocol support for VSCode terminal integration ### Bug fixes - **Authentication issues on login** - Fixed authentication issues that could occur during login ## CLI v0.68.0 · Desktop v0.22.0 (March 4, 2026) Subagent inactivity timer, Usage Limits settings, and streaming stability fixes ### New features - **Subagent inactivity timer** - Subagents now automatically stop after 3 minutes of inactivity to prevent idle resource usage - **Usage Limits settings** - Configure and view Usage Limits directly within sessions and settings (Factory App) ### Bug fixes - **VS Code auto-installation** - Fixed embedded VS Code auto-installation in the Factory App - **Factory App stability** - Fixed daemon restart issues in the Factory App - **Session resume** - Fixed reliability issues when resuming previous sessions - **Streaming stability** - Fixed a crash that could occur during streaming retries - **BYOK error handling** - Improved error messages for payment and authentication errors with custom API keys - **Context management** - Improved handling when conversations exceed context limits - **Image handling for text-only models** - Images are now properly handled when using text-only models - **GLM-4.7 error handling** - Fixed error handling for GLM-4.7 model errors - **MCP tool compatibility** - Improved MCP tool schema sanitization for better compatibility across model providers - **Custom model streaming** - Fixed streaming issues with custom models - **Reasoning level display** - Fixed reasoning level not displaying correctly after using /model in spec mode - **Authentication reliability** - Reliability improvements to authentication mechanisms ## CLI v0.67.0 · Desktop v0.21.0 (March 3, 2026) CLI banners, /stats command, skills search, and configurable image compression ### New features - **CLI banners** - Display customizable banners in the CLI for quick access to important information - **`/stats` command** - The `/wrapped` command has been renamed to `/stats` with date range filtering support - **Skills search bar** - Added search bar to the skills menu for quickly finding skills - **Configurable image compression** - New `image_quality` parameter in the Read tool for controlling image compression quality - **Electron app automation** - The agent-browser skill now supports automating Electron desktop applications ### Improvements - **Automation detail view** - Redesigned automation detail view to align with session UI patterns (Factory App) - **Windows ARM64 and Zed extension** - CLI distribution now includes Windows ARM64 builds and Zed editor extension ### Bug fixes - **Certificate loading order** - Fixed system certificate loading that could cause authentication failures on startup - **Gemini schema compatibility** - Fixed schema compatibility issues with Gemini models - **Payment redirect in Factory App** - Fixed payment page redirect handling in the Factory App - **Character encoding** - Fixed display of malformed UTF characters in messages - **Missions list formatting** - Fixed formatting of feature and worker list descriptions in Missions - **BYOK error handling** - Improved error messages for custom model (BYOK) configurations - **MCP server reconnection** - Fixed potential deadlock when MCP servers fail to reload during disconnection - **Execute tool file actions** - Fixed file operations not being applied correctly during command execution ## CLI v0.66.0 · Desktop v0.20.0 (March 2, 2026) /copy command, marketplace management, and Missions worker improvements ### New features - **`/copy` slash command** - Copy content to your clipboard directly from the CLI - **Marketplace management** - New UI for managing marketplace plugins under the Plugins section in enterprise controls (Factory App) - **Factory App version display** - Factory App now shows the current version number - **Grouped exec sessions** - Droid exec sessions are now grouped under a dedicated tab for cleaner organization (Factory App) ### Improvements - **Missions worker reliability** - Hardened Missions worker lifecycle with improved pause and resume handling ### Bug fixes - **Subagent model and permissions** - Fixed subagent model inheritance and permissions level - **BYOK compaction errors** - Upstream error details are now surfaced in BYOK compaction failures - **Cross-provider messages** - Fixed message conversion for OpenAI Responses API across providers - **Spec model display** - Corrected spec model display in `/model` selector and prevented global fallback on clear - **Session archiving** - Fixed session archive confirmation flow in the app - **Chat scrolling** - Fixed scroll-to-last-message behavior in session chat (Factory App) - **User Usage Limits modal** - Fixed scroll bug in the Individual User Usage Limits modal (Factory App) - **Diff generation** - Fixed potential crash with non-string inputs in diff generation ## CLI v0.65.0 · Desktop v0.19.0 (February 27, 2026) Diagnostics command, GitHub comments in diff viewer, and Mission Mode fixes ### New features - **/diagnostics command** - New `/diagnostics` command with overlay menu and footer status for troubleshooting CLI issues - **Unstaged changes in diff viewer** - View and manage unstaged changes directly in the diff viewer panel (Factory App) - **GitHub comments in diff viewer** - View GitHub PR comments alongside code changes in the diff viewer (Factory App) - **Session deep linking** - Share and navigate directly to specific sessions via deep links (Factory App) - **Diff preview for enterprise controls** - Preview configuration diffs when managing enterprise control settings (Factory App) ### Bug fixes - **Mission Mode model selection** - Fixed model selection issues when using Mission Mode - **Mission Mode validation** - Validation contracts now update correctly when Mission requirements change - **Bug report smart trimming** - Bug reports now smart-trim large files instead of failing - **Plugin auto-install** - Fixed duplicate core plugin auto-install failures being surfaced to users ## CLI v0.64.0 · Desktop v0.18.0 (February 26, 2026) Session archiving, inline rename, custom model reasoning effort, and Mission Mode improvements ### New features - **Session archiving** - Archive sessions from the CLI to keep your session list clean - **Inline session rename** - Rename sessions inline with cloud sync support - **Custom model reasoning effort** - Added reasoning effort support for custom models - **Concurrent Missions warnings** - Warns on Missions entry and confirms concurrent Missions runs - **Dry-run validation** - New dry-run validation approach during Missions planning - **/model scope clarity** - Clarified `/model` command scope and added set-as-default action ### Bug fixes - **Mission Mode commands** - Split Mission Mode enter and exit into separate commands for improved reliability - **Context management** - Fixed context management triggering too frequently for OpenAI models - **OpenAI reasoning effort** - Fixed maximum reasoning effort setting for OpenAI models - **Gemini rate limiting** - Improved Gemini 429 error handling with throttle-aware retry - **Bedrock reasoning effort** - Fixed reasoning effort settings for Bedrock models - **Empty end_turn responses** - No longer treats empty end_turn responses as errors ## CLI v0.63.0 · Desktop v0.17.0 (February 25, 2026) Model selector redesign, session default settings, and Droid Shield improvements ### New features - **Top 5 model selector** - Model selector now shows top 5 models with a "show more" option for cleaner UI - **Session default settings** - All session default settings (model, autonomy, reasoning) now exposed in `/settings` menu ### Bug fixes - **Droid Shield false positives** - Reduced false positives on example files and placeholders - **Glob tool patterns** - Fixed handling of non-array patterns input in Glob tool - **BYOK auth in exec mode** - Fixed authentication validation for BYOK custom models in exec mode - **Session model auth** - Check effective session model for BYOK auth skip, not just CLI flag - **Session loading** - Fixed session loading issues with long conversation histories ## CLI v0.62.0 · Desktop v0.16.0 (February 24, 2026) Bedrock/Vertex support, session-scoped settings, readiness-fix command, and Rate Limits ### New features - **Bedrock/Vertex support** - Added Bedrock and Vertex support for Claude Sonnet 4.6 and Opus 4.6 - **Session-scoped settings** - Model, autonomy, and reasoning settings are now session-scoped instead of global - **`/readiness-fix` command** - New slash command to automatically fix issues found in readiness reports - **Rate Limits UI** - Added UI handling for Rate Limit notifications - **README requirement for Missions** - Missions now require README updates before completion ### Bug fixes - **MCP reload prevention** - Avoided unnecessary MCP reloads on non-MCP settings changes - **Custom model migration** - Prevented no-op custom model migration writes - **Autonomy persistence** - Preserved autonomy level when spec sessions persist ## CLI v0.61.0 · Desktop v0.15.0 (February 23, 2026) Vertex/Bedrock Opus support, Mission Control polish, and interactive command prevention ### New features - **Vertex/Bedrock Opus support** - All latest Opus models now available via Vertex and Bedrock - **Online research in Missions** - Added online research and environment setup during Missions planning - **Recent local directories** - Quick access to recently used local directories ### Bug fixes - **Mission Control polish** - Visual improvements to Mission Control UI - **Interactive command prevention** - Prevent interactive command execution that could block the agent - **Missions latency** - Reduced `/enter-mission` latency and improved daemon error UX - **Spec/autonomy decoupling** - Decoupled spec and autonomy settings and permission outcomes ## CLI v0.60.0 · Desktop v0.14.0 (February 20, 2026) Faster search, Missions model selector, and Missions validation overhaul ### New features - **Missions model selector** - Choose specific models for Missions workers - **Prevent idle sleep** - CLI now prevents system idle sleep during Missions execution ### Improvements - **Faster search** - Sped up search functionality in large repositories with fuzzy search support ### Bug fixes - **Missions validation overhaul** - Complete overhaul of Missions validation for improved reliability - **Missions worker timeouts** - Improved Missions worker timeout handling - **Subagent model inheritance** - Fixed model inheritance for subagents - **Diff word highlights** - Fixed incorrect line wrapping in diff word-level highlights - **Gemini tool usage** - Improved Gemini tool usage reliability - **Readiness report sanitization** - Improved sanitization of sensitive data in readiness reports - **TUI daemon shutdown** - Fixed stopping TUI-spawned daemon on shutdown ## CLI v0.58.0 · Desktop v0.12.0 (February 19, 2026) Local settings override, environment variable references in custom models, and Missions UX improvements ### New features - **Local settings override** - New `settings.local.json` for local project/folder-level settings overrides - **Environment variable references** - Custom model API keys now support environment variable references - **Gemini 3.1 Pro model** - Added Gemini 3.1 Pro model support - **Missions planning pausing** - Allow pausing during planning phases with in-context resume in Mission Control view ### Bug fixes - **Anthropic provider reset** - Fixed stale Anthropic provider state on model switch - **MCP toggle effects** - MCP toggle changes now apply immediately without restart - **Browser MCP isolation** - Stopped auto-adding `--isolated` flag for browser MCP servers - **Windows PowerShell** - Fixed PowerShell execution issues on Windows - **GLM 4.7 routing** - Fixed GLM 4.7 model routing issues - **Ctrl+G spam prevention** - Blocked Ctrl+G while the agent is streaming to prevent notification spam ## CLI v0.57.16 · Desktop v0.11.15 (February 17, 2026) Claude Sonnet 4.6 model support ### New features - **Claude Sonnet 4.6** - Added Claude Sonnet 4.6 model support ## CLI v0.57.15 · Desktop v0.11.14 (February 16, 2026) Opus 4.6 Fast Mode pricing update ### Improvements - **Opus 4.6 Fast Mode pricing** - Removed promotional pricing for Opus 4.6 Fast Mode ## CLI v0.57.14 · Desktop v0.11.13 (February 13, 2026) Session tags, word-level diffs, new models, and built-in skills ### New features - **Session tags** - Categorize sessions with custom tags using the new `--tag` CLI flag for better organization - **Word-level diff highlighting** - GitHub diff mode now highlights only changed words within lines instead of entire lines - **GPT-5.3-Codex model** - Added GPT-5.3-Codex model with multi-turn phase support - **GLM-5 model** - Added GLM-5 model via Fireworks - **MiniMax M2.5 model** - Added MiniMax M2.5, replacing M2.1 - **Create-PR skill** - New built-in skill for guided pull request creation with local verification checklist - **Worker droid** - Basic worker droid now ships as a built-in droid for task delegation - **PR linking** - Sessions now display a status indicator showing linked PRs from git operations ### Bug fixes - **Stdin character corruption** - Replaced text input component to eliminate character corruption and broken special keys - **NODE_ENV injection** - Stopped injecting `NODE_ENV=production` into shell environment, fixing `npm` devDependencies issues - **MCP SDK alignment** - Updated to MCP SDK 1.26 with proper schema validation - **File watcher limits** - More granular file watching to avoid "too many open file descriptors" errors - **Factory App connection race** - Fixed daemon connection race condition in the Factory App with polling retry logic - **Sidebar sessions** - Fixed sessions disappearing from sidebar (Factory App) - **Open in editor** - Fixed "open in editor" functionality in the Factory App ## CLI v0.57.12 · Desktop v0.11.11 (February 11, 2026) Spec approval with comment and input stability ### New features - **Spec approval with comment** - Added a "Proceed with comment" option to spec approval, letting you approve a spec and attach a guiding comment in one action ### Bug fixes - **Fast model switching** - Fixed edge cases when switching models mid-conversation, including proper conversation state management - **Invalid MCP output schemas** - Tolerates invalid MCP tool output schemas instead of crashing - **Opus 4.5 max tokens** - Corrected incorrect Opus 4.5 max output tokens in model registry - **Ctrl+O during AskUser** - Allows Ctrl+O (open in editor) during the AskUser tool flow - **AskUser context** - Ensures context is clear when using the AskUser tool - **Session file table overflow** - Limits concurrency when fetching sessions in daemon to avoid "file table overflow" errors - **Duplicate plan steps** - Fixed duplicate plan steps in web UI when multiple TodoWrite calls share the same title (Factory App) ## CLI v0.57.11 · Desktop v0.11.10 (February 10, 2026) Subagent fixes and model switching improvements ### Bug fixes - **Subagent prompting** - Fixed subagent tool prompting and description handling - **AskUser tool** - Handles edge cases preventing stuck states when sessions are interrupted or messages arrive out of order (Factory App) - **Sidebar staleness** - Fixed missing last message and stale status indicators on the session sidebar (Factory App) ## CLI v0.57.10 · Desktop v0.11.9 (February 9, 2026) Skills UX overhaul, terminal tab titles, Factory App notifications, and mobile experience ### New features - **`/skills` UX overhaul** - Tabbed layout with pagination and a dedicated detail view for browsing and importing skills - **Terminal tab titles** - Terminal tab now displays the session name so you can identify Droid sessions across multiple tabs - **`droid update` command** - New command to manually update the CLI, plus ability to disable autoupdate via `/settings` - **Session end hooks** - Hooks now fire when running `/clear` or `/new`, enabling custom cleanup workflows - **Factory App notifications** - Native macOS and Windows notifications when sessions complete, with configurable preferences - **Factory App native view redesign** - Refreshed sidebar layout with smooth animations, new terminal editor icons, and keyboard shortcut formatting - **Mobile experience** - Overhauled mobile and responsive layout with touch-friendly settings and onboarding flows (Factory App) ### Bug fixes - **`/bug` command memory** - Optimized log loading with streaming I/O to prevent out-of-memory on large log files - **Session persistence** - Thinking and reasoning fields now preserved correctly when saving sessions - **MCP settings UI** - Filled toggle style, clickable tool rows, and capitalized server names (Factory App) - **Git timeout** - Increased git timeout to 60 seconds to prevent failures on slow connections or large repos ## CLI v0.57.9 · Desktop v0.11.8 (February 7, 2026) Opus 4.6 Fast Mode and pricing update ### Improvements - **Opus 4.6 Fast Mode** - Renamed from "Opus 4.6 Fast" to "Opus 4.6 Fast Mode" for clarity - **Opus 4.6 Fast Mode promo pricing** - Updated Model Multiplier from 12x to 6x promotional rate ## CLI v0.57.7 · Desktop v0.11.6 (February 6, 2026) Opus 4.6 Fast, /fork command, session archiving, and org-managed settings ### New features - **Claude Opus 4.6 Fast** - Added Opus 4.6 Fast as a new model option in the model selector - **`/fork` command** - Duplicate your current session with all messages preserved into a new session - **Org-managed settings** - Local settings that are set by your organization are now locked from editing - **Session archiving** - Archive sessions from the web app to keep your sidebar clean (Factory App) ### Bug fixes - **Shell environment loading** - Shell environment now loads consistently across Factory App, sandbox, and computer platforms - **Orphaned browser processes** - Agent-browser daemon processes are now cleaned up properly - **Task tool output** - Progress updates from Task tool are now truncated to prevent UI overflow ## CLI v0.57.5 · Desktop v0.11.4 (February 5, 2026) Opus 4.6, semantic diffs, npm package, plugin hooks, and startup optimizations ### New features - **Claude Opus 4.6** - Added Claude Opus 4.6 model support - **Semantic structured diffs** - Code change diffs now use semantic structure-aware rendering for better readability - **Plugin hooks** - Added support for plugin hooks with Claude hooks format compatibility - **npm package** - CLI is now installable via npm: `npm install -g droid` - **User-invocable skills as slash commands** - Custom skills marked as user-invocable are now registered as slash commands - **`FACTORY_LOG_FILE` and `FACTORY_DISABLE_KEYRING` env vars** - New environment variables for log file output and keyring control - **Enterprise Controls** - New settings page for managing agent behavior, command policy, model access, and security (Factory App) - **Skills UI** - Skills button and modal added to the chat input for easy skill access (Factory App) - **Ask User tool** - Ask User tool now available in the web and Factory App surfaces ### Improvements - **Startup latency** - Optimized CLI startup with sub-metric tracking for faster launch times ### Bug fixes - **Unicode crash** - Fixed CLI crash when sessions contain certain unicode characters - **Marketplace auto-update** - Auto-update now only runs in interactive mode, not during `droid exec` ## CLI v0.57.3 · Desktop v0.11.2 (February 3, 2026) Task tool streaming, image pasting, and skills improvements ### New features - **Task tool live streaming** - Real-time subagent tool progress display and Execute tool output streaming in web and Factory App surfaces - **Image pasting** - Paste images directly into the chat composer (Factory App) - **`.agent` skills folders** - Skills can now be loaded from `.agent` folders in your project for custom skill support ### Improvements - **Skills filtering** - Skills are now filtered by model invocability so only relevant skills appear in the Skill tool ### Bug fixes - **File suggestion UX** - Trailing space now added after file suggestion selection for smoother typing ## CLI v0.57.2 · Desktop v0.11.1 (January 30, 2026) Full text search, MiniMax M2.1, and AskUser improvements ### New features - **Full text search** - Search across your session history and codebase directly in the CLI - **MiniMax M2.1** - Added MiniMax M2.1 model support ### Bug fixes - **AskUser tool** - Improved questionnaire parsing and interaction reliability ## CLI v0.57.0 · Desktop v0.11.0 (January 29, 2026) Plugin hooks and ASCII mermaid diagrams ### New features - **ASCII mermaid diagrams** - Mermaid diagrams now render as ASCII art directly in the terminal ## CLI v0.56.1 · Desktop v0.10.2 (January 28, 2026) MCP tool management and Read tool fixes ### Improvements - **MCP tool management** - Improved MCP tool status tracking to prevent cross-contamination between servers ### Bug fixes - **Read tool images** - Fixed image support in the Read tool, along with rate limit retry and reasoning fixes - **Cache block limits** - Ensured no more than 3 cache blocks are sent to the LLM to prevent context issues ## CLI v0.56.0 · Desktop v0.10.2 (January 27, 2026) ACP daemon mode, lazy session loading, and model updates ### New features - **ACP daemon mode** - New daemon mode for Agent Control Protocol enabling persistent background sessions with improved session loading - **Lazy session loading** - Sessions now load messages lazily for faster startup when resuming large sessions - **Dynamic default settings** - Settings now support dynamic defaults for better configuration flexibility ### Improvements - **Parallel context gathering** - Missions workers now encouraged to read all context files in parallel for faster startup ### Bug fixes - **MCP tool schema isolation** - Broken tool schemas in one MCP server no longer affect other MCP servers - **Symlinked skill directories** - Fixed discovery of symlinked skill directories in `.factory/skills` - **Windows spec mode UI** - Fixed spec mode UI and autonomy mode text rendering on Windows - **AskUser parallel tool calling** - Fixed parallel tool calling with AskUser tool - **Skills menu rendering** - Fixed skills menu to always render correctly - **OpenAI custom model routing** - Fixed routing bug for OpenAI custom models - **Linear remote delegation** - Fixed Linear remote delegation issues - **Authentication atomic writes** - Auth writes now use atomic operations to prevent corruption - **ACP session loading** - Fixed multiple issues with ACP session loading ## CLI v0.55.2 · Desktop v0.10.1 (January 23, 2026) Marketplace auto-updates and background processes flag ### New features - **Marketplace auto-updates** - New management UI with auto-update setting and silent auto-installation of missing marketplaces on startup - **`--allow-background-processes` flag** - Added flag support for interactive and resumed sessions ### Bug fixes - **AskUser tool parsing** - Improved parsing logic for questionnaire formatting ## CLI v0.55.1 · Desktop v0.10.1 (January 22, 2026) OTEL tracing, bug reporting improvements, and review command fixes ### New features - **OpenTelemetry tracing** - Implemented OTEL tracing for the CLI to improve observability and debugging - **Mandatory bug report titles** - Added title requirement for `/bug` reporting to ensure clearer issue tracking ### Bug fixes - **`/review` default branch** - Fixed the `/review` command to correctly identify and use the repository's default branch - **Authentication** - Fixed several issues around authentication persistence and stability ## CLI v0.55.0 · Desktop v0.10.1 (January 21, 2026) Autonomy mode reskin and agent-browser integration ### New features - **Autonomy mode reskin** - Refreshed UI for autonomy mode indicators and transitions for better clarity - **Agent-browser integration** - CLI now uses the `agent-browser` CLI for browser automations, providing more reliable and token-efficient interactions ## CLI v0.54.0 · Desktop v0.10.0 (January 20, 2026) Readiness report improvements and AskUser fixes ### New features - **OTEL metrics for hooks and slash commands** - Customer metrics for `droid.hook.invocations` and `droid.slash_command.invocations` - **Optional `/readiness-report` model switching** - Users can now continue with their current model while seeing a recommendation, instead of being forced to switch ### Improvements - **Autonomy indicator format** - Updated format to show `Spec | Auto (X)` when spec mode is active; `Medium` shortened to `Med` ### Bug fixes - **AskUser tool improvements** - Questionnaire parsing now supports bullets/`1)` numbering, ignores headers/code fences, normalizes multi-word topics. AskUser can now run alongside TodoWrite with other tools deferred properly. Fixed TUI input handling via `KeypressProvider` subscription - **Windows UI** - Improved autonomy/spec mode status text rendering on Windows terminals ## CLI v0.53.0 · Desktop v0.10.0 (January 16, 2026) Readiness report in exec mode and settings stability ### New features - **`/readiness-report` in exec mode** - Support for `/readiness-report` slash command in droid exec mode with optional custom instructions ### Bug fixes - **Settings.json corruption prevention** - Settings file now uses atomic writes to prevent corruption - **Terminal info disabled** - Disabled getTerminalInfo to prevent issues - **Authentication stability** - Bundled keychain dependency in packaged builds to reduce flaky sign-in setups ## CLI v0.52.0 · Desktop v0.9.3 (January 15, 2026) Session sharing, ctrl+z suspend, and bug fixes ### New features - **`/share` command** - Share your current session with organization members, copies share URL to clipboard ### Improvements - **Ctrl+Z suspend** - Handle ctrl+z suspend like other TUIs - **Decoupled modes and autonomy** - Modes and autonomy levels are now independent settings in the TUI ### Bug fixes - **Task tool view** - Fixed Ctrl+O view for Task tool - **Settings file writes** - Avoid unnecessary file writes when updating unrelated settings - **Unified tool display** - Fixed tool result pending display issue - **Subagent version** - Fixed subagent spawning to use correct droid version ## CLI v0.49.0 · Desktop v0.9.2 (January 14, 2026) CLI Axiom skill, custom model fixes, and bug fixes ### New features - **CLI Axiom skill** - New skill for querying CLI metrics and logs with relevant datasets and fields ### Improvements - **Custom model requests** - Custom models now properly map to equivalent built-in models for correct request formatting ### Bug fixes - **iTerm scrollback clearing** - Fixed scrollback clearing for iTerm users with "Disable E3 scrollback clearing" preference enabled - **Memory leak fix** - Fixed memory leak by cleaning up Performance API entries and hinting Bun GC - **Session title generation** - Fixed session title generation for empty strings - **Warmup requests** - Warmup requests no longer consume unnecessary tokens by disabling thinking budget - **Light mode submenus** - Fixed light mode display issues for submenus - **Linux autoupdate** - Fixed autoupdate issues on Linux ## CLI v0.48.0 · Desktop v0.9.2 (January 13, 2026) ACP streaming, Windows autoupdate, and bug fixes ### Improvements - **Windows autoupdate** - Updates now work correctly on Windows using a deferred update strategy, applied on next startup - **Model selector** - Removed GLM 4.6 from the model selector ### Bug fixes - **Tool result rendering** - Internal placeholder markers no longer displayed in CLI output - **Session loading** - Fixed "Failed to load session" errors on cloud sessions - **Droid startup timeout** - Increased timeout for start droid to prevent timeouts ## CLI v0.47.0 · Desktop v0.9.1 (January 12, 2026) Bug fixes ### Bug fixes - **Read tool infinite loop** - Fixed infinite loop when Read tool returns empty string - **Session title generation** - Fixed race condition causing "Session must have at least 1 message" error - **Windows sound paths** - Fixed Bun virtual FS sound path extraction on Windows ## CLI v0.46.0 (January 9, 2026) Session rename, cloud sync, and bug fixes ### New features - **`/create-skill` command** - Create custom skills directly from the CLI with guided flow - **Skills in exec mode** - Skill tool now enabled when running `droid exec` - **`/rename` command** - Rename sessions with a new slash command, plus auto-generated session names based on conversation content ### Improvements - **PR overview in FetchUrl** - GitHub PR fetch results now include an overview section with files changed, comment counts, and table of contents - **Create-skill UX** - Improved create-skill flow with better guidance and examples - **Review branch fetch** - Git review now fetches current branch when opening instead of on mount ### Bug fixes - **Background process output** - Output from background processes now properly piped to droid - **Cloud sync settings** - Fixed settings cloud sync with droid status tracking - **ESC cancel duplication** - Fixed queued messages being duplicated when pressing ESC to cancel ## CLI v0.45.0 (January 8, 2026) Bug fixes and improvements ### Bug fixes - **Ctrl+O detailed view** - Now shows full accumulated command output instead of truncated view - **File autocomplete** - Menu now closes properly on Ctrl+C - **MCP modal** - Fixed modal and status notification display issues - **Autonomy level** - Fixed autonomy level not being applied correctly ## CLI v0.44.0 (January 7, 2026) Custom status line, bulk MCP tool management, and UX improvements ### New features - **`/statusline` command** - Configure custom status line, can import PS1 from shell config - **Bulk MCP tool management** - Manage multiple MCP tools at once via ToolsOverviewView - **Auto-include Skill tool** - Custom droids (subagents) now automatically include the Skill tool ### Improvements - **Expand diffs with Ctrl+O** - Tool confirmation prompts now support Ctrl+O to expand diffs ### Bug fixes - **Session resume system info** - Session resume now shows correct/up-to-date system info and local date - **Model display alignment** - Fixed model display alignment in UI - **MCP project server config** - MCP tool modifications now work correctly for project-defined servers ## CLI v0.43.0 (January 6, 2026) Terminal and chat UI fixes, improved spec mode ### Bug fixes - **Terminal panel display** - Fixed terminal being cut off when plan panel is shown - **Chat input width** - Chat input now extends to full terminal width - **Spec mode transitions** - `droid exec --use-spec` correctly exits spec mode after `ExitSpecMode` - **Default session settings** - Session defaults properly loaded from settings.json on each new session ## CLI v0.42.2 (January 5, 2026) GLM 4.7 support, Gemini improvements, and bug fixes ### Improvements - **GLM 4.7 support** - Updated GLM model deployments - **Gemini SDK upgrade** - Improved Gemini support - **Expandable tool results** - Tool results can now be expanded to show full details in web and Factory App surfaces ### Bug fixes - **Fixed light mode color themes in expanded mode** - **Fixed memory leak from unbounded tool executions** - **Fixed GitHub app installation OAuth flow** - **Fixed dotenv loading order with settings** - **Fixed token usage not preserved across manual compaction** - **Fixed onboarding error when user has existing org (Factory App)** - **Fixed users getting stuck in onboarding due to stage issues (Factory App)** - **Fixed mobile onboarding layout (Factory App)** ## CLI v0.41.0 (December 30, 2025) Year-in-review /wrapped command, PDF uploads, and faster model switching ### New features - **`/wrapped` command** - Year-in-review stats showing your Droid usage, badges earned, model preferences, and session metrics (cli) - **Fast model switching** - Faster provider/model switches without LLM compaction (cli) - **Batched tool permissions** - Parallel tool calls now show single permission prompt (cli) - **Custom BYOK models** - Custom models now appear in model selector (cli) - **Auto-sync repos** - GitHub/GitLab repos sync when opening template modal (Factory App) ### Improvements - **PDF & text file uploads** - Attach PDFs and text files to chat in web and Factory App surfaces - **Send button** - Visual send button in session chat input (Factory App) - **ASCII startup animation** - Animated Droid logo on CLI startup (configurable) (cli) - **Model Multiplier display** - Shows Factory model multipliers in model descriptions (cli) ### Bug fixes - **Fixed users getting stuck in broken onboarding state (Factory App)** - **Fixed CLI onboarding redirect for terminal-only users** - **Fixed missing repos in GitHub org management modal (Factory App)** - **Fixed multi-option spec mode selection (cli)** - **Fixed catch-22 bug preventing subscription start (Factory App)** ## CLI v0.40.0 (December 23, 2025) Custom models, parallel tool confirmations, and settings file watching ### New features - **Custom models from settings** - Load custom models directly from settings.json - **Parallel tool confirmations** - Single permission request for parallel tool calls instead of individual prompts - **Multi-option spec mode** - Spec mode can now present multiple implementation options for user selection - **Settings file watching** - Settings automatically reload when the settings file changes - **Token usage in exec mode** - Token usage is now displayed in exec streaming output - **Stop hook improvements** - Decision and reason support for Stop hook to match Claude code spec ### Bug fixes - **Opus 4.5 model fix** - Fixed Opus 4.5 to work as an effort model with thinking guard improvements - **Tool completion fix** - Fixed issue where tools weren't properly completed on new assistant messages - **Web search date fix** - Web search tool now includes dynamic date for more accurate results ## CLI v0.39.0 (December 19, 2025) Context utilization setting and Custom Droids enabled by default ### New features - **Context utilization setting** - New setting to display token usage indicator in the status bar - **Custom Droids enabled by default** - Custom Droids feature is now available to all users without requiring opt-in ### Bug fixes - **Grep tool fix** - Fixed pattern argument handling in the Grep tool - **BYOK Grok fix** - Fixed crash when using Grok models with thinking/reasoning streams - **Warmup improvements** - Added option to disable warmup requests and skip warmup for slash commands - **Non-git directory support** - Use current working directory as project directory when not in a git repository ## CLI v0.38.0 (December 18, 2025) GPT-5.2 improvements and .env loading fix ### Bug fixes - **GPT-5.2 improvements** - Fixed request parameters and reasoning effort options for better model performance - **Prevent .env auto-loading** - CLI no longer automatically loads `.env` files from the working directory in standalone builds, making behavior more predictable ## CLI v0.37.0 (December 17, 2025) Todo tool improvements, Chrome DevTools MCP, and model cleanup ### Improvements - **Todo tool improvements** - Simplified todo format and refreshed UI for better reliability - **Chrome DevTools MCP server** - Added Chrome DevTools Protocol to the MCP registry - **Model cleanup** - Removed obsolete models ### Bug fixes - **Fixed API request handling for Codex models** ## CLI v0.36.6 (December 17, 2025) Gemini 3 Flash, terminal links, and terminal reliability fixes ### New features - **Gemini 3 Flash model** - Added support for Gemini 3 Flash ### Improvements - **Agent readiness signals** - Expanded agent readiness signals in readiness reports ### Bug fixes - **Fixed overly strict filtering by allowing common read-only `git` commands** - **Fixed copy/paste issues on WSL** - **Fixed repository deduplication in `/readiness` reports** ## CLI v0.36.2 (December 13, 2025) GPT-5.2 model and MCP tool management ### New features - **GPT-5.2 model** - Added support for GPT-5.2 model - **MCP tool enable/disable** - Add ability to enable/disable individual MCP tools per server ## CLI v0.36.0 (December 10, 2025) MCP tools in droid creation and droid management fixes ### New features - **MCP tools in droid creation** - MCP tools are now displayed during custom droid creation and edit flows ### Bug fixes - **Fixed droid deletion when filename doesn't match metadata name** - **Standardized help text position and format in droid creation/edit flows** - **Add reasoning_effort to stream-json system init event** ## CLI v0.33.0 (December 8, 2025) Autonomy mode fixes and file handling improvements ### Bug fixes - **Fixed autonomy mode not being properly set on tool confirmations** - **Improved @-tagged file truncation by character count to prevent context overflow** - **Updated Opus 4.5 pricing with dismissible notice** - **Restored Figma MCP server to the registry** ## CLI v0.32.0 (December 5, 2025) Readiness report command, ESC key improvements, and input fixes ### Bug fixes - **Press ESC to close the expanded tool result view** - **Improved input box cursor up/down movement with wrapped lines** - **Fixed Windows pasting issues** - **Fixed `npm run format` conflict with legacy Windows format command in denylist** - **Fixed auth issues with Microsoft MCP servers** - **`/readiness` command is now enabled by default** ## CLI v0.31.0 (December 4, 2025) GPT-5.1-Codex-Max model, image compression, IDE auto-connect, and input fixes ### New features - **GPT-5.1-Codex-Max model** - Added support for GPT-5.1-Codex-Max model - **Image compression** - Images are now compressed before upload to reduce bandwidth - **IDE auto-connect setting** - Added setting to automatically connect to IDE from external terminals - **Disable hooks flag** - Added `--no-hooks` flag to disable hooks execution ### Bug fixes - **Improved Droid docs search** - **Long Execute tool commands now truncated in the UI header** - **Fixed on-disk images not being cleared after upload** - **Fixed thinking level changing mid-conversation by locking it per agent turn** - **Fixed review preset items using incorrect color for non-selected items** - **Fixed input box handling of multi-width characters like CJK and emoji** - **Fixed left/right arrow key navigation when @ suggestion menu is open** - **Bundled code-signed ripgrep binary for improved security and reliability** ## CLI v0.30.0 (December 3, 2025) Hidden file search, session search, image indicators, and ESC key fixes ### New features - **Search hidden files** - Grep and Glob tools now include hidden files in search results - **Session search** - Added search functionality to the session selector for quickly finding sessions - **Image indicator** - User messages now display an image indicator when images are attached ### Bug fixes - **Fixed EPERM permission errors on certain file operations - Windows file rename retry logic** - **Fixed ESC key navigation in `/review` and `/bg-process` commands for Ghostty terminal** - **MCP registry updates** ## CLI v0.28.1 (December 2, 2025) Certificate caching, image upload limits, and settings rendering fixes ### Improvements - **Certificate caching** - Certificates are now cached on startup for faster loading - **Image upload limits** - Conversation images are now limited to prevent 413 errors when uploading large images - **MCP registry helper** - Added note for STDIO servers in MCP registry explaining installation requirements ### Bug fixes - **Fixed invalid signature in thinking block after assistant message interrupt** - **Fixed settings menu not showing all options when loading asynchronously** ## CLI v0.28.0 (December 1, 2025) Session management, slash command navigation, thinking improvements, and Figma MCP support ### New features - **Show all sessions** - View and manage all your sessions across directories - **Windowed slash command navigation** - Improved navigation in the slash command menu with windowed scrolling - **Improved thinking display** - Refactored thinking block rendering for better clarity - **Figma MCP server** - Added support for Figma MCP server integration ### Bug fixes - **Fixed detection of Cerebras-style context length exceeded errors** - **Fixed duplicate spec modes and double empty line above input** - **Added promo price label for Opus model** - **Fixed `/install-github-app` command issues** ## CLI v0.27.4 (November 28, 2025) Interleaved thinking support, show thinking setting, and hooks permanently enabled ### New features - **Interleaved thinking support** - Display multiple thinking blocks during streaming, including redacted thinking blocks with safety messages - **Show thinking setting** - Show AI thinking/reasoning on the main view, turn on/off in `/settings` - **Hooks always enabled** - The hooks feature is now permanently enabled and available via the `/hooks` command ### Bug fixes - **Fixed ESC/Q navigation in `/model` menus - pressing ESC or Q now properly navigates back to previous menu level instead of closing all menus** - **Fixed garbage characters appearing in prompt input caused by terminal response bytes leaking into stdin** - **Fixed reasoning effort display in spec mode to correctly reflect the actual reasoning effort being used** - **Fixed OpenAI chat completions streaming for tool calls** - **Improved Gemini's usage of the todo tool with better prompting** - **Fixed model warmup for GLM-4.6** - **Fixed `/install-github-app` command issues** ## CLI v0.27.2 (November 25, 2025) Project-level MCP configs, IDE command, and custom model improvements ### New features - **Project-level MCP configs** - Configure MCP servers at the project level using `.factory/mcp.json` - **`/ide` command** - New command to manage VS Code, Cursor, and Windsurf IDE integrations. Shows current extension version or prompts to install - **Extra args in custom models** - Pass additional arguments to custom model configurations e.g. service_tier, temperature, top_p, etc ### Bug fixes - **Fixed task tool infinite wait when running subagents** - **Fixed expanding tilde (~) paths when loading sessions** - **Improved `@` search rankings - exact filename matches now appear at the top with VSCode-style two-column display** - **Removed automatic creation of `.factory/skills` folders** - **Show compacting state when triggered** - **Consolidate markdown rendering (fixes missing character in spec mode)** - **Parallelize loading certificates on Windows** ## CLI v0.27.1 (November 24, 2025) New model support, persistent session settings, and auto-update improvements ### New features - **Improved rewind functionality** - new `/rewind` command, plus speed and UX improvements. - **`/install-github-app` command** - New command to install the Factory GitHub app directly from the CLI. - **Images in MCP tool responses** - MCP tools can now return images that are properly sent to the LLM. ### Bug fixes - **Upgraded to bun 1.3.3** - **Safer mechanism to check if ripgrep is installed** - **Fixed gpt-5.1-codex reasoning** - **Shift+Backspace deletes single characters like Backspace** - **Fix GLM4.6 as a spec mode model** - **Fix exitSpecMode prompt** - **Fix custom models for subdroids** - **Prevent console logs during startup** - **Fix spec mode for Gemini** ## CLI v0.26.12 (November 22, 2025) Eval tooling upgrades, model catalog refresh, and reliability-focused TUI fixes ### New features - **`/review` command** - Provides an interactive code review workflow. Review code changes in different ways: comparing against a base branch, reviewing specific commits, examining uncommitted changes, or providing custom review instructions. ### Bug fixes - **Gemini 3 Pro now uses the updated 0.8× pricing multiplier - replacing introductory pricing** - **MCP navigator now fills the terminal** - **Update routing for /resume to filter correctly to /sessions** - **Position cursor at start when navigating history** - **Make all views in the mcp navigator full width** - **Fixed garbled character rendering on Windows** - **Unify LLM retry logic and do more retries in exec mode** - **Removed the deprecated Figma MCP server** ## CLI v0.26.10 (November 21, 2025) Expanded MCP registry and improved session reliability ### New features - **MCP Registry Expansion** - Added 30+ new MCP server integrations including development tools (Playwright, Braintrust, Honeycomb), databases (Supabase, MongoDB, Prisma, Neon), security scanning (Snyk, Semgrep), and more ### Bug fixes - **Fixed race conditions where model changes mid-turn** - **Fixed bug report command** - **Fixed support for duplicate model IDs in custom model configurations** - **Fix PowerShell command execution on Windows** ## CLI v0.26.8 (November 20, 2025) Background processes, MCP search, and TUI stability improvements ### New features - **Background Processes Support** - Added support for running and managing background processes in the CLI. Includes a new `bg-process` command to list, kill, and clean up processes, with persistent process state tracking. - **MCP Registry Search** - Added search functionality to the MCP registry list view (`/mcp`). Filter available MCP servers by typing directly in the list interface. ### Bug fixes - **Fixed race conditions and rendering issues when suspending/resuming the TUI (e.g., when opening an external editor).** - **Fixed circular dependency issues in grep tool logging.** ## CLI v0.26.7 (November 19, 2025) Directory-specific sessions, MCP image support, and Gemini fixes ### New features - **Directory-Specific Sessions** - Sessions are now stored per directory and automatically switch to the correct working directory when loaded. The `/sessions` command shows only sessions created in your current directory plus favorited sessions - **MCP Image Support** - MCP tools can now return images that are sent to the LLM for analysis - **Custom Models for Subdroids** - Subdroids can now use custom models instead of inheriting from parent session ### Bug fixes - **Fixed console output issues in exec mode** - **Fixed spec mode for Gemini models** - **Fixed exit spec mode prompt behavior** ## CLI v0.26.3 (November 18, 2025) Claude Code hooks migration, improved suggestions, and GLM support ### New features - **Skills Enabled by Default** - Skills command now enabled for all users - **Claude Code Hooks Auto-Migration** - Automatically detects and imports hooks from Claude Code CLI with interactive prompts. Smart translation converts `bash_tool` to `Execute` and `CLAUDE_CWD` to `DROID_CWD`, preventing duplicates and tracking migration state - **@ Suggestions Improvements** - Now includes folders in addition to files with better UI and performance ### Bug fixes - **Fixed GLM4.6 configuration issues preventing it from working in spec mode** - **Fixed warmup API calls** - **Shift+Backspace now deletes single character like Backspace** ## CLI v0.26.0 (November 14, 2025) Skills system, session favorites, and improved git integration ### New features - **Skills System** - Claude Code-compatible `.factory/skills` support for modular, prompt-based capabilities. Use the `/skills` command to manage skills and import from `.claude/skills` directories - **Session Favorites** - New `/favorite` command to pin/unpin sessions, keeping your active projects at the top of the session list - **Enhanced Bug Reporting** - The `/bug` command now zips session context, uploads it to Factory, and returns a shareable report ID automatically - **Cleaner Diff Viewer** - Improved UI with horizontal lines instead of borders ### Bug fixes - **Always prompt to accept or reject generated specs and pass the correct labels to the UI flow** - **`droid spec` now writes files to the expected default path without manual overrides** - **Fixed autonomy handling so MCP tools respect the configured confirmation level in every session** - **Better error handling in pre-update logic** - **Added Axiom MCP server to the registry** ## CLI v0.25.0 (November 13, 2025) Execute tool streaming, enhanced hooks system, and MCP autonomy controls ### New features - **Enhanced Hooks System** - Added 7 new hook types for complete lifecycle control: - UserPromptSubmit - Modify or validate prompts before they're sent to the agent - Stop - Execute custom logic when the agent completes (e.g., metrics collection) - SubagentStop - Track and log subagent task completion - PreCompact - Run custom logic before conversation compaction (can block compaction if needed) - SessionStart - Initialize sessions and inject environment variables - SessionEnd - Cleanup tasks when sessions end (triggered on logout, quit, Ctrl+C) - **MCP Tool Autonomy Levels** - Granular control over MCP tool confirmations: low autonomy requires confirmation for read-only MCP tools, high autonomy requires confirmation for all MCP tool calls - **Execute Tool Streaming** - Long-running commands now display the last two non-empty lines of output in real-time. No more wondering if Droid is stuck or slow. ### Bug fixes - **Fixed paste handling in hooks UI - no more escape sequence artifacts when pasting commands** - **Fixed message ID alignment after conversation compaction** - **Fixed unfinished tool uses being cleared when switching providers to retry** - **Fixed markdown prompt usage for GPT models** - **Fixed /mcp rendering for narrow terminal windows** - **Fixed bracketed paste handling when suspending TUI** - **Authentication errors now appear directly in the CLI with clear error messages** ## CLI v0.24.0 (November 12, 2025) Hooks system, interactive spec editing, and major performance improvements ### New features - **Hooks** - Introduced a powerful hooks system allowing you to run custom scripts before/after tool executions with configurable exit codes for different behaviors (success, warning, block, abort) - [Experimental - turn on in /settings] - **Interactive Spec Editing** - Added ability to interactively edit specs before execution - **Prompt Cache Warmup** - Added prompt-cache warmup while you're typing for faster responses ### Bug fixes - **Fixed model switching to correctly handle conversation compaction when needed** - **Fixed spec mode tab cycling for reasoning levels** - **Fixed bracketed paste when suspending TUI** - **Corrected line numbers in ApplyPatch tool diff display** - **Eliminated race condition in ripgrep path resolution** - **Various agent loop stability improvements** - **Tool execution now correctly respects autonomy mode** - **Applied workaround for Figma MCP server compatibility** ## CLI v0.23.0 (November 10, 2025) Markdown tables, pinned todos, and MCP improvements ### New features - **Markdown table rendering** - CLI now properly renders markdown tables for better display of structured data and documentation - **Pinned todo plans** - Pin important todo items to keep them visible and track tasks during long-running sessions - **MCP tool result formatting** - Improved formatting for MCP tool results with cleaner, more readable output ### Bug fixes - **Fixed thinking block support for chat completion APIs to enable extended reasoning in compatible models** - **Fixed issue with baseline versions reverting back to non-baseline builds on update** - **Fixed autonomy mode handling in `droid exec` command for more reliable execution** - **Restored Ctrl+T keyboard shortcut functionality on Windows** ## CLI v0.22.14 (November 7, 2025) Reasoning controls, bash output expansion, and MCP improvements ### New features - **Tab to cycle reasoning levels** - Navigate through different reasoning levels using the Tab key for better control - **Expanded bash mode outputs** - Transcript view now expands and displays bash command outputs for better visibility ### Bug fixes - **Fixed MCP tool discovery for droid exec so tools appear correctly in --list-tools and pass validation** - **Fixed diff view to line wrap for better readability of long lines** - **Fixed OAuth refresh for MCP servers** - **Fixed markdown rendering bug in transcript view** - **Fixed message cancellation handling in TUI** - **Fixed prompt cache key handling in completions** - **Brought back Ctrl+T keyboard shortcut on Windows** ## CLI v0.22.12 (November 6, 2025) Console improvements and enhanced reliability ### New features - **Quit/exit aliases** - type 'quit' or 'exit' commands for exiting sessions ### Bug fixes - **Collect all AGENTS.md and CLAUDE.md files for system reminders** - **Fixed MCP server argument parsing when adding servers through the UI** - **Improved system diagnostics utilities for better troubleshooting** - **Cleaning and dynamic rendering of the pending tools to make it clear Droid is not stuck** ## CLI v0.22.11 (November 5, 2025) Model support, tool permissions, and MCP improvements ### New features - **Context7 MCP server** - Added Context7 MCP server to the MCP registry ### Bug fixes - **Improved droid exec run type tracking** - **Improved MCP client info handling** ## CLI v0.22.10 (November 4, 2025) OAuth discovery, interrupt support, and stability improvements ### New features - **OAuth Discovery** - Automatically discover OAuth providers to streamline authentication setup for MCP - **OpenAI Reasoning Summaries** - Display OpenAI's extended reasoning summaries in expanded session view ### Bug fixes - **Fixed copy-pasting behavior in terminal sessions** - **Improved timeout handling for more reliable operations** - **Enhanced Windows compatibility** - **Enable custom droids by default** - **Improved error logging and diagnostics** ## CLI v0.22.9 (November 1, 2025) MCP Registry, tool permissions, session resume, and multiple bug fixes ### New features - **MCP Registry** - Added a curated selection of MCP servers for easy discovery and setup [MCP Registry interface showing curated MCP servers](/docs-assets/images/mcp-registry-2.mp4) - **Tool Permissions Support** - Implemented request tool permissions flow in stream-jsonrpc mode with support for multiple concurrent tool confirmations - **Session Resume** - Show resume command on exit (Ctrl+C or /quit) with new `-r, --resume` flag to continue sessions - **Thinking Traces Display** - Added UI component to show thinking blocks in expanded session view ### Bug fixes - **Fixed CLI completion sounds by extracting embedded sound files to host filesystem, making them accessible to external shell commands** - **Fixed tool use in user message after tool cancellation** - **Added baseline builds for x64 CPUs (pre-2013) that don't support AVX2 SIMD instructions** - **Updated static split logic to use text messages as stable checkpoints, reducing flickering during rendering** - **Fixed storage and sending of Gemini thinking content from delta.extra_content.google** - **Removed MultiEdit tool due to low success rate (`~80%`)** - **Added recovery from malformed optional arguments using optimistic JSON parsing** ## CLI v0.22.7 (October 31, 2025) Reasoning field support, Import from Claude, and bug fixes ### New features - **Reasoning field support in streaming responses** - Added support for displaying reasoning fields in chat completion streaming chunks for models that provide extended reasoning - **Import from Claude restored** - Re-enabled the "Import from Claude" option in the droids menu for easier conversation migration ### Bug fixes - **Fixed MCP status indicator to only show when MCP servers are actually configured** - **Fixed message list flickering** - **Added support for custom headers when adding MCP servers** - **Fixed sound files not being included in SEA binary releases** - **Fixed Ctrl+C interruption messages to only display to users and not appear in logs for cleaner output** - **Resolved issue where tool use would incorrectly appear in user messages after tool cancellation** ## CLI v0.22.6 (October 30, 2025) MCP Revamp and OpenAI organization field fix ### New features - **MCP Revamp** - Complete overhaul of Model Context Protocol implementation ([docs updated](/cli/configuration/mcp)) ### Bug fixes - **Fixed issue with OpenAI organization field** - **Restore Import (I) options in Droids (subagents) menu** ## CLI v0.22.5 (October 29, 2025) Default model behavior improvements ### Bug fixes - **Default Model Behavior** - Improved default model behavior ## CLI v0.22.4 (October 28, 2025) Report previews and help text fixes ### Bug fixes - **Restored Report Previews** - Re-enabled HTML preview for reliability reports so you can view them directly in the web - **Fixed --help text** - Running droid exec --help was previously giving incorrect info for using droid exec in streaming mode. This has now been updated to include the correct flag name with correct instructions. ## CLI v0.22.3 (October 21, 2025) Detailed transcript view, customizable sounds, terminal setup support, and model upgrades ### New features - **Detailed transcript view (Ctrl+O)** - Press Ctrl+O to view comprehensive tool execution details with complete breakdowns for all tool types including Execute, Edit, MultiEdit, Create, and more - **Customizable completion sounds** - Configure sound notifications when commands complete, with support for different sounds in focus mode and custom sound file paths - **Terminal setup support** - Added automated setup support for Warp, iTerm2, and macOS Terminal with automatic terminal indicators for improved integration - **PowerShell update** - Now uses `pwsh.exe` for better PowerShell compatibility and reliability on Windows systems - **Drag-and-drop image support** - Attach images to CLI sessions via drag-and-drop for easier multimodal interactions - **Custom models in subagents** - Added support for using custom models when executing subagents via Task tool - **Custom droid descriptions** - Moved custom droid descriptions to Task tool for better organization and usability - **`droid exec` custom model support** - Added support for custom models in droid exec sessions ### Bug fixes - **Resolved issues with prompt caching to improve performance and reduce costs** - **Added model ID logging on CLI tool execution for better observability** - **Limited messages rendered in Ctrl+O view to improve performance with large transcripts** - **Fixed sound files not being included in SEA binary releases** - **Fixed readonly value change issues that caused unexpected behavior** - **Fixed ExitSpecMode flickering issues for smoother user experience** - **Removed duplicate custom models header in settings** - **Improved session persistence to legacy sessions to maintain backward compatibility** - **Fixed infinite render issues on Windows systems** - **Fixed 400 error handling and validation** - **Added model validation error messages for better debugging** - **Fixed certificate loading timing to occur after logging initialization** ## CLI v0.21.3 (October 14, 2025) MCP OAuth, system certificates, fuzzy search, and session management improvements ### New features - **MCP OAuth support** - Added OAuth authentication support for Model Context Protocol in TUI for secure server connections - **System certificates support** - Added support for loading system certificates on Windows and macOS, with additional Windows system certificates support - **Fuzzy search implementation** - New fuzzy search for improved CLI command and option discovery - **Rewind fork workflow** - Added ability to rewind and fork from previous points in conversation sessions - **Bug report enhancements** - Added version, OS, and shell information to bug reports for better troubleshooting - **Live updates for subagent tool** - Added real-time progress updates when using the Task tool - **Enhanced custom droid prompts** - Improved prompts to encourage more proactive usage of custom droids - **Droid Shield optional setting** - Made Droid Shield an optional configurable setting with toggle support ### Bug fixes - **Added session_id to stream-json output format for better tracking and debugging** - **Enhanced authentication flow for smoother onboarding** - **Streamlined first-run experience by removing redundant onboarding steps** - **Fixed OAuth callback server port collision issues that prevented authentication** - **Fixed certificate loading timing issues causing TLS errors** - **Improved error metadata logging for better debugging** - **Fixed terminal setup issues across different shell environments** - **Better error messages when ripgrep (rg) isn't available in PATH** - **Improved handling of corrupted unicode normalization in file paths** - **Fixed impact level default value handling** ## CLI v0.19.8 (September 30, 2025) Droid exec enhancements, custom model support, streaming JSON, and stability improvements ### New features - **`droid exec` Slack integration** - Added `slack_post_message` tool to droid exec for posting updates to Slack channels - **`droid exec` streaming JSON input mode** - Added streaming JSON input mode for multi-turn exec sessions for better automation workflows - **`droid exec` tool configuration** - Added ability to enable/disable specific tools in droid exec sessions - **`droid exec` pre-created session IDs** - Added support for pre-created session IDs in droid exec for better session management - **Initial prompt support** - Added support for launching TUI with an initial prompt via command line - **GLM-4.6 support** - Added LLM proxy support for GLM-4.6 with Fireworks, Baseten, and DeepInfra providers - **Azure OpenAI for GPT-5** - Added Azure OpenAI support for GPT-5 Codex in TUI - **Custom droid generation** - Added auto-generation of custom droids using LLM for faster setup - **MCP streamable HTTP servers** - Added support for streamable HTTP MCP servers - **Model selector refresh** - Updated model selector UI with improved organization and image support warnings - **Tab key auto-complete** - Tab key now auto-completes custom commands without submitting for better UX ### Bug fixes - **`droid exec` session continuation fix** - Fixed droid exec to correctly continue an existing session instead of overwriting session data - **Dynamic reasoning label for Codex models in UI** - **AGENTS.md source paths now shown in settings for better transparency** - **Better error messages when custom model JSON configuration is broken** - **Manual update instructions when auto-update isn't available** - **Fixed task subagents to properly terminate on abort** - **Fixed file rename retry logic on Windows systems** - **Fixed backslash rendering issues on Windows** - **Major improvements to custom subagents reliability** - **Fixed task tool subagents prompt for better execution** - **Don't stop agent on cancelled file edits in spec mode** - **Validated working directory exists before executing ripgrep** - **Fixed subagents to inherit model from TUI parent session** - **Truncated large MCP tool results to prevent context overflow** - **Fixed /cost command with Fireworks cached input** - **Added user-agent header to LLM requests for better tracking** - **Locked down file system permissions for better security** - **Fixed rendering for failed edit calls in spec mode** - **Fixed compaction logic and retry without tool results** - **Applied output transforms before compaction/caching** - **Persisted session usage for accurate cost tracking** - **Fixed /new command to reset timer and sessionId properly** - **Fixed compaction retry without tool results** - **Fixed screenshot reading with unicode normalization fallback** # Feature Maturity How Factory labels feature maturity, and how releases and versioning work. Factory ships quickly, so some capabilities reach you before they are fully settled. A small set of maturity tags tells you how stable a feature is and what to expect from it. This page explains those tags and how our releases and versioning work. ## Maturity tags A maturity tag appears next to a page title and in the navigation sidebar. Most features are stable and carry no tag, so a tag is only present when a feature is still maturing or on its way out. marks a feature with limited or invite-only access that is already running in production. It is stable enough to rely on, though access and surface area are still expanding. marks a feature that is being phased out. Avoid it for new work and migrate to the recommended alternative when you can. **Generally available** features carry no tag. They are stable, supported, and safe to build on, and this is the default for everything in the docs unless a tag says otherwise. Generally available is the absence of a tag, not a label. "GA", "Beta", "Public Preview", and "Research Preview" are intentionally not part of the vocabulary, so you will never see them as tags. ## Releases and versioning The **[Full Changelog](/changelog/release-notes)** lists every notable change across the Factory App, Droid CLI, and other surfaces, newest first. Subscribe to it as a feed at `/changelog/rss.xml`. The Droid CLI is versioned, and changelog entries note the version a change shipped in (for example, `v1.10.0`). The Factory App and other hosted surfaces ship continuously, so their changes are dated rather than versioned. Watching for a capability to graduate? Follow the full changelog; a feature dropping its{' '} tag is announced there. # Agent Arena Agent Arena results and methodology for AI coding agents. # Legacy Bench Legacy-Bench results and methodology for AI agents working on legacy software. # NextJS Next.js Evals results and methodology for AI coding agents. # Review Benchmark Review Benchmark results and methodology for AI code review models. # Terminal Bench Terminal Bench results and methodology for AI coding agents. # Docs visual elements preview Draft gallery of new docs components for dense content pages, each shown against the bullet or card pattern it replaces. This draft page previews five new visual elements for dense docs pages, plus the existing components they pair with. Each section converts real content from a current page so the comparison is honest: same information, different surface. ## PropertyList Replaces `- **term** - description` bullet runs (the dominant pattern on [CLI settings](/droid-cli/settings), exit codes, and config enumerations). Today that content renders like this: {/* sweep-allow: term-bullets */} - **`off` / `none`**: no structured reasoning (fastest). - **`minimal`**, **`low`**, **`medium`**, **`high`**: progressively more deliberation. - **`extra high`**, **`max`**: the deepest reasoning, where a model offers it. The same content as a `PropertyList`, with room for meta chips the bullets cannot carry: How much structured reasoning the model applies before acting. `off` and `none` disable it entirely (fastest); `minimal` through `high` add progressively more deliberation; `extra high` and `max` request the deepest reasoning where a model offers it. How diffs render in the session. `github` is side-by-side and higher fidelity (recommended); `unified` is the traditional single-column format. Sync sessions to the web app. Set to `false` to keep sessions local only. Commands that can never run. Unlike the denylist there is no prompt and no way to approve them: the block applies even under full autonomy, auto-run, or `--skip-permissions-unsafe`. Disables all hooks globally when `true`. Prefer per-hook enablement. ## Badge An inline status pill for prose annotations: defaults, platform support, availability, risk levels. Generalizes the lifecycle pill already used next to page titles. Sandboxing is on by default on macOS only; Linux support is in preview. Commands on the blocklist are never run, denylisted commands require approval, and reasoning effort is model dependent. ## Kbd Inline keycaps for shortcuts that today render as plain inline code. Press ⌘K to open search, for the command palette, or Esc to dismiss. ## LabeledDivider The hairline rules zoning this page are the component itself: a mono caption plus rule that breaks a long page into scannable regions without spending a card or a heading. Left-aligned by default, centered below: ## Timeline A status-oriented milestone rail for lifecycle and rollout content, distinct from `Steps` (which stays reserved for procedural instructions). Converted from the feature-maturity narrative: Available to design partners behind an organization flag. Breaking changes may ship without notice. Open to all organizations. Interfaces are stable but SLAs are not yet enforced. Covered by production support and enterprise billing. ## ComparisonTable A decision surface between raw markdown tables and CardGroups: bordered frame, mono header row, optional recommended-column highlight. Converted from the "pick a surface" bullets on the exec overview: | Capability | Droid CLI | Droid Exec | | ------------------- | ---------------- | ---------- | | Interactive session | Yes | No | | CI/CD pipelines | Partial | Yes | | Structured output | Text | Text, JSON | | Autonomy control | Prompted | Flag-based | | Best for | Development work | Automation | ## Callouts The five callout variants, unchanged. They stay the tool for asides; the new components exist so callouts stop carrying enumerations. Sessions sync to the web app unless `syncSessions` is `false`. Add low-risk utilities to the command allowlist to skip prompts. Blocklisted commands never run, even under full autonomy. Reasoning effort trades latency for deliberation. Your CLI is authenticated and ready. ## Inline vars and lifecycle Shared values render through vars components, so Droid and links like [support@factory.ai](mailto:support@factory.ai) stay canonical, and pills annotate feature maturity inline and next to page titles. ## Steps Numbered `Steps` stay reserved for procedural instructions the reader performs (`Timeline` above covers status). Unchanged: Run the install command for your platform. Start `droid` and sign in via the browser prompt. Describe the change; review the diff before applying. ## Lists Top-level bullets keep the orange square marker; nested levels now get a 1px indent guide and a hollow marker so depth stays legible: - Deterministic controls run first - Command allowlist skips prompts for trusted utilities - Command blocklist never runs, even with approval - Model-led review runs second - The risk model catches misses - The downgrade model clears false alarms ## Cards `Card`/`CardGroup` are retained for navigation and genuine feature summaries, where the blueprint surface and brackets earn their weight: The agent in your terminal: sessions, autonomy levels, and IDE hooks. Package procedures Droid loads on demand across your team. ## Tabs `Tabs` for parallel platform or option branches currently written as bullets: Sandboxing uses the system sandbox profile and is on by default. Sandboxing uses namespaces and is available in preview builds. Sandboxing is not yet available; auto-run policies still apply. ## Code surfaces `CodeGroup` switches between equivalent snippets, and `Terminal` gives shell commands window chrome: ```bash title="macOS / Linux" curl -fsSL https://app.factory.ai/cli | sh ``` ```powershell title="Windows" irm https://app.factory.ai/cli/windows | iex ``` ## Stat strip The Software Factory dashboard kit (`StatStrip`, `Stat`, and the `Radar` and `Gauge` instruments) for numeric claims on overview pages: ## Checklist `Checklist` for requirement and support matrices: Runs read-only commands without prompts Respects the command allowlist Runs blocklisted commands (never, even with approval) Network isolation depends on platform ## FAQ `FAQ` and `Troubleshooting` rows carry schema.org FAQPage markup, the one sanctioned disclosure surface: Yes. Blocklisted commands never run; there is no prompt and no override. Yes. Denylisted commands always ask first but run once you approve. ## Mermaid Architecture and flow diagrams use `mermaid` fences with the Factory theme init block, never one-off SVG components: ```mermaid %%{init: {"theme": "base", "themeVariables": {"fontFamily": "Geist Mono, monospace", "fontSize": "13px", "primaryColor": "#161413", "primaryBorderColor": "#342F2D", "primaryTextColor": "#FAFAFA", "lineColor": "#4D4947", "textColor": "#D6D3D2"}}}%% flowchart LR a["Authored MDX"] --> v["Velite build"] v --> r["Registry components"] r --> d["docs.factory.ai"] ``` Not demoed here: the media family (markdown images, `FrameGroup`, `DocsVideo`) needs real assets; the install surfaces (`InstallPanel`, `CliInstallCommands`, `DesktopDownloadCards`) and the hand-authored API kit (`Endpoint`, `ApiField`, ...) are specialized page kits. The general-purpose disclosure family (`Accordion`, `Expandable`) is deliberately banned by `docs:check:components`; compress with `PropertyList` and `ComparisonTable` instead of hiding content. ## RelatedLinks The `RelatedLinks` component is the single sanctioned docs page footer: a compact row of curated cross-group links under the standardized "Related resources" mono caption (omit the `label` attribute; the validator rejects overrides). Every docs page closes with at most one `` block, and it must be the final element of the page body. Package procedures Droid loads on demand. Configure models, tools, and behavior. Run Droid headless in scripts and CI. Isolate execution for riskier work.