This is the full developer documentation for Firewatch
# Get started with Firewatch
> Configure your profile, on-call coverage, escalation policy, and first alert.
Firewatch needs four connected pieces before it can page anyone:
1. an active **service**;
2. an **on-call schedule** or service owners;
3. an **escalation policy** assigned to that service; and
4. at least one available notification channel for each responder.
Organization access and on-call eligibility are separate. Administrators assign people as **Responders** when they can be scheduled or receive service-owner escalation. **Collaborators** can view and collaborate without being eligible for paging.
## Join or create an organization
[Section titled “Join or create an organization”](#join-or-create-an-organization)
Firewatch Cloud is designed for future self-service team signup. During the private beta, join the access list from the Firewatch website or accept the invitation sent by your organization. Paid checkout and automatic team provisioning are not enabled yet.
On a new self-hosted installation, the first person to register becomes the organization administrator. The default `Bootstrap` registration mode closes registration after that account is created. Additional users join by administrator invitation.
Firewatch Cloud charges only for responders. Collaborators do not consume a responder seat. Assigning responder access at any point counts for that entire billing month, even if the person is changed back to a collaborator later.
## 1. Check your profile
[Section titled “1. Check your profile”](#1-check-your-profile)
Open **Settings → Profile** and confirm:
* your name;
* your personal display time zone, or that you want to inherit the organization time zone; and
* your verified phone number when SMS or voice delivery is available.
Your profile time zone controls how Firewatch displays times and renders alert templates for you. It does not change a schedule’s time zone or handoff rules.
Open **Settings → Authentication** to add a passkey, set or change a password, and configure authenticator-app two-factor authentication when local authentication is available.
## 2. Create a service
[Section titled “2. Create a service”](#2-create-a-service)
Open **Services** and create the system you want to monitor. The creator becomes its first primary owner. Add other primary or secondary owners as needed.
A service can have one escalation policy. Alerts for a service without an assigned policy still create incidents, but they do not enqueue responder notifications.
## 3. Define coverage
[Section titled “3. Define coverage”](#3-define-coverage)
Open **On-call → Schedules** and create a rotation:
* choose the schedule time zone and local handoff time;
* add responders in rotation order;
* use a fixed duration or an alternating coverage pattern; and
* include only the weekdays that should have coverage.
Review the result in **Calendar**. Excluded weekdays are real coverage gaps. Fill an intentional exception with an override; fix an accidental gap before using the schedule for production.
## 4. Build an escalation policy
[Section titled “4. Build an escalation policy”](#4-build-an-escalation-policy)
Open **On-call → Escalation policies**. The first step is immediate. Later steps wait for their configured delay after the preceding step.
Each step can target:
* the responder currently resolved from an on-call schedule;
* every primary owner of the affected service; or
* all owners of the affected service.
Assign the policy to the service under **Services**.
## 5. Connect alert delivery
[Section titled “5. Connect alert delivery”](#5-connect-alert-delivery)
Firewatch attempts every deployment-enabled channel available for the targeted responder:
* email when email delivery is enabled;
* Slack when the organization workspace is connected and the responder is mapped;
* SMS and optional voice when telephony is enabled and the responder has a verified phone number; and
* each enabled organization outbound webhook.
Self-hosted administrators configure provider credentials under **Admin settings**. Firewatch Cloud operates the providers, while organization administrators still connect their Slack workspace and responders still verify their contact methods.
## 6. Send an alert
[Section titled “6. Send an alert”](#6-send-an-alert)
Open **Alert sources**, select a service, and choose an integration. Firewatch can accept its normalized alert format directly or generate an AWS SNS/SQS Lambda adapter.
Create an organization API key under **Settings → API keys** for external alert senders. Store the full secret immediately—it is shown only once.
Before relying on Firewatch:
1. trigger a non-production alert;
2. confirm it creates an incident for the expected service;
3. verify the expected responder and channels;
4. acknowledge it and confirm later escalation steps are skipped; and
5. send the corresponding recovery event and confirm the incident resolves.
Before relying on Firewatch for production paging, test failure, retry, backup, and recovery paths for your deployment.
# Profile, contact methods, and sign-in
> Manage time zones, verified phone numbers, passkeys, passwords, and two-factor authentication.
## Profile and time zone
[Section titled “Profile and time zone”](#profile-and-time-zone)
Open **Settings → Profile** to update your name when the deployment allows it. Managed identity providers can make names read-only.
You can inherit the organization time zone or select a personal one. Your personal selection controls display and notification-template timestamps; it does not alter on-call schedule rules.
## Account email
[Section titled “Account email”](#account-email)
Firewatch uses the account email for invitations, account-security mail, email alerts, and exact-email Slack mapping.
The self-hosted bootstrap administrator does not need an email-verification round trip. Other registration modes and managed deployments can require verification.
## Phone number
[Section titled “Phone number”](#phone-number)
When Twilio is enabled, add an E.164 phone number such as `+12025550123`. Firewatch sends a verification code before the number becomes eligible for incident SMS or voice delivery.
Only a confirmed number is used. Removing it disables both phone channels for your account. Voice delivery also depends on the deployment-wide incident voice setting and may be disabled even when SMS is available.
## Passkeys
[Section titled “Passkeys”](#passkeys)
Passkeys provide phishing-resistant passwordless sign-in through WebAuthn. Firewatch supports up to ten named passkeys per account. You can add, rename, and remove them under **Settings → Authentication**.
A passkey-only account cannot remove its final passkey. Keep at least two on separate devices when possible. Removing a passkey from Firewatch does not remove the corresponding credential from the device or password manager.
Passkeys require HTTPS except for the browser’s special `localhost` development case. Changing a self-hosted installation’s WebAuthn RP ID makes passkeys registered under the old ID unusable.
## Passwords and two-factor authentication
[Section titled “Passwords and two-factor authentication”](#passwords-and-two-factor-authentication)
Accounts allowed to use local authentication can set or change a password. Firewatch supports authenticator-app TOTP as a second factor and issues recovery codes when TOTP is enabled.
Store recovery codes outside the device that holds the authenticator. Each code is single-use. Disabling TOTP requires the current password.
Enterprise SSO policy can disable local passwords, passkeys, TOTP, and recovery codes for an organization.
## Recovery limitation
[Section titled “Recovery limitation”](#recovery-limitation)
General email account recovery is not implemented yet. Do not assume that an SMTP configuration provides a forgotten-password flow. Keep multiple passkeys, recovery codes, or an administrator-controlled identity-provider recovery path appropriate to your deployment.
# Incidents and alert response
> Declare, acknowledge, investigate, resolve, reopen, and export Firewatch incidents.
## Incident identity
[Section titled “Incident identity”](#incident-identity)
Every incident belongs to one service and receives a stable reference such as `INC-000142`. Its severity is `SEV1`, `SEV2`, or `SEV3`, and its status moves through:
```text
triggered → acknowledged → resolved
```
A resolved incident can be reopened. Reopening starts another response cycle without removing the earlier history.
Incidents can be declared manually or created from alert ingestion. Alert events retain their provider source, identity, labels, source link, occurrence time, and ingestion outcome.
## Find an incident
[Section titled “Find an incident”](#find-an-incident)
The incident list supports server-side search, pagination, and filters for:
* status;
* severity; and
* service.
The dashboard summarizes open, unacknowledged, and open-SEV1 counts along with recent activity.
## Respond
[Section titled “Respond”](#respond)
Open an incident to:
* acknowledge ownership;
* change severity;
* resolve it with an optional reason and actual resolution time;
* reopen a resolved incident; and
* add responder notes.
A responder can edit their own note. Firewatch appends a new revision rather than replacing history, and the UI can display prior revisions.
Acknowledgement or resolution causes any not-yet-delivered escalation work to be recorded as skipped.
## Timeline and delivery audit
[Section titled “Timeline and delivery audit”](#timeline-and-delivery-audit)
The lifecycle records incident creation, status and severity transitions, correlated alerts, escalation stages, delivery successes and failures, skipped notifications, and responder-note revisions.
Delivery attempts store the channel, recipient, outcome, and error. Provider failure does not erase the incident or its other channel attempts.
## Suggested runbooks
[Section titled “Suggested runbooks”](#suggested-runbooks)
Firewatch scores active runbooks using their service associations and matching words from the incident title, description, and service name. Suggestions are assistive, not an automated safety decision; responders should verify that a procedure applies before following it.
## Export
[Section titled “Export”](#export)
An incident can be exported as:
* a PDF operational record; or
* CSV containing its lifecycle and note revisions.
Exports reflect the current stored record. Protect them like other incident data because they can contain alert descriptions, responder identities, and notes.
# On-call schedules and escalation
> Configure rotations, coverage gaps, overrides, shift trades, calendars, and escalation targets.
## Schedule time zones
[Section titled “Schedule time zones”](#schedule-time-zones)
Every schedule owns an IANA time zone and a local rotation start. A 09:00 handoff remains 09:00 in that schedule through daylight-saving changes.
Organization and personal time zones affect defaults and display. They do not silently change an existing schedule.
Only organization members with **Responder** on-call access can be added to a rotation, override, or shift trade. Collaborators retain product access but are not eligible for alert routing.
## Coverage patterns
[Section titled “Coverage patterns”](#coverage-patterns)
A coverage pattern says how many covered days each responder owns before the rotation advances. Participants are evaluated in the displayed order, and the pattern repeats for as long as the schedule runs.
The schedule editor provides these presets:
| Pattern | Stored as | Example with Avery, Blair, and Casey |
| ----------------- | ------------- | ---------------------------------------------------------------------------------------------------------------------- |
| Weekly | `7` | Avery covers seven days, Blair the next seven, then Casey. |
| Daily | `1` | Avery covers Monday, Blair Tuesday, Casey Wednesday, then Avery Thursday. |
| Weekday / weekend | `5 / 2` | Avery covers Monday–Friday, Blair covers Saturday–Sunday, and Casey gets the next Monday–Friday block. |
| Four / three | `4 / 3` | Avery covers Monday–Thursday, Blair covers Friday–Sunday, Casey gets the next four-day block, and the cycle continues. |
| Custom | Your sequence | `2 / 2 / 3` assigns two days to Avery, two to Blair, three to Casey, then repeats with the next participant. |
The numbers describe consecutive **coverage days**, not a repeating assignment to one named person. After every segment, Firewatch advances to the next active participant. For example, `4 / 3` with three responders produces:
1. Avery for four coverage days;
2. Blair for three coverage days;
3. Casey for four coverage days;
4. Avery for three coverage days; and
5. Blair for the next four coverage days.
Use **Weekly** for a conventional weekly handoff, **Daily** when ownership changes every day, and **Weekday / weekend** when weekday and weekend blocks should go to different people. Use **Custom** only when the handoff cadence cannot be expressed by a preset.
Each segment must be between 1 and 28 coverage days. A pattern can contain up to 14 segments and a schedule can contain up to 100 unique participants.
## Weekday exclusions and gaps
[Section titled “Weekday exclusions and gaps”](#weekday-exclusions-and-gaps)
Choose the weekdays on which the rotation provides coverage. Excluded weekdays:
* have no rotation responder;
* do not advance the alternating pattern; and
* appear as gaps in the calendar.
The rotation start must fall on an included weekday. An exact schedule override can cover a normally excluded day.
Use this for schedules that deliberately operate only on certain weekdays. Do not exclude weekends from a 24×7 production schedule: an excluded day is intentionally **uncovered**, not silently assigned to Friday’s responder.
## Scheduled participant changes
[Section titled “Scheduled participant changes”](#scheduled-participant-changes)
Each participant can have an optional **joins at** and **leaves at** instant. Firewatch recalculates the active participant list at those instants, allowing a future addition or removal without a precisely timed manual edit.
Review the future calendar after changing participant dates. Removing someone from the active roster changes subsequent rotation positions.
## Overrides
[Section titled “Overrides”](#overrides)
An override assigns one organization member for an exact start and end time and supersedes the normal rotation. Overrides on the same schedule cannot overlap.
Use overrides for leave, temporary cover, or one-off gaps. They are different from changing the rotation because they do not rewrite future cadence.
## Shift coverage and trades
[Section titled “Shift coverage and trades”](#shift-coverage-and-trades)
Any member can request:
* **coverage**, where another member takes the requester’s shift; or
* a **trade**, where each member takes one of the other’s non-overlapping shifts.
The requester must own the entire offered rotation window. For a trade, the target must own the entire return window. Only the target can accept or decline, and only the requester can cancel while the request is pending.
Acceptance writes the corresponding schedule overrides atomically. Outstanding requests appear as a count next to **Shift trades**.
## Month and week calendars
[Section titled “Month and week calendars”](#month-and-week-calendars)
The **Calendar** page has directly linkable month and week views. Select a shift to inspect its schedule, responder, time range, source, and related services.
Holiday calendars add organization closures and early-close annotations to the calendar. They are informational overlays today: they do not remove or replace on-call coverage automatically. Use weekday exclusions or overrides when the rotation itself should change.
Built-in holiday calendars cannot be edited. Custom calendars can contain full closures or early-close events in their own time zone.
## Escalation policies
[Section titled “Escalation policies”](#escalation-policies)
A policy contains up to ten ordered steps:
* the first step is immediate;
* every later delay is relative to the preceding step; and
* a step can target a schedule, primary service owners, or all service owners.
When an incident is created, Firewatch writes every delayed step to its durable outbox. A schedule target is resolved when that step becomes due, so current overrides and effective participant changes are honored.
Acknowledging or resolving the incident prevents remaining steps from sending. The skipped work remains visible in the incident lifecycle instead of being silently deleted.
At each due step, Firewatch attempts all enabled delivery channels available for each targeted responder. It does not send email, then Slack, then phone as three separate policy steps unless you model separate responder targets yourself.
# Runbooks
> Create internal Markdown procedures or link external runbooks to services and incidents.
Firewatch supports two runbook types:
* **Internal** runbooks store Markdown in Firewatch.
* **External** runbooks link to an existing procedure at an absolute URL.
Both types can be associated with one or more services and suggested from an incident.
## Internal runbooks
[Section titled “Internal runbooks”](#internal-runbooks)
The editor supports Markdown including headings, lists, links, tables, and images. You can paste, drop, or select PNG, JPEG, GIF, and WebP images up to the limit shown in the editor.
Image bytes use the configured asset provider:
* self-hosted Compose stores them in the persistent `firewatch-runbooks` volume; and
* Firewatch Cloud operates durable object storage for uploaded images.
The database stores the asset metadata and runbook content. Back up both the database and filesystem image volume in a self-hosted deployment.
## External runbooks
[Section titled “External runbooks”](#external-runbooks)
Use an external runbook when another system is the authoritative source. The link opens outside Firewatch, so its availability and access control remain the responsibility of that system.
Do not put temporary signed URLs or credentials in the runbook URL.
## Service associations and suggestions
[Section titled “Service associations and suggestions”](#service-associations-and-suggestions)
Associate a runbook with the services it can help recover. Firewatch ranks suggestions using that association and matching terms from the incident and runbook content.
Associations improve discovery but do not grant access outside Firewatch and do not cause a procedure to execute automatically.
## Archiving
[Section titled “Archiving”](#archiving)
Archiving removes a runbook from active use without deleting its historical record. The runbook list can show active, archived, or all runbooks.
# Services and alert sources
> Model service ownership, assign escalation policies, and ingest alerts safely.
## Services
[Section titled “Services”](#services)
A service is the organization-scoped system affected by an incident. It owns:
* a unique name and optional description;
* one or more primary or secondary owners;
* an optional escalation policy; and
* its incident and alert history.
The member who creates a service becomes its first primary owner. Primary and secondary are routing labels: both owner types can manage the service, while an escalation step can specifically target primary owners.
Only responders can create a service or be selected as an owner because owner targets participate in escalation. Service owners and organization administrators can edit ownership and routing. Firewatch prevents removal of the final owner. Archiving a service preserves its history but prevents new incidents and alert ingestion for that service.
## Assign an escalation policy
[Section titled “Assign an escalation policy”](#assign-an-escalation-policy)
Create the policy under **On-call → Escalation policies**, then select it from the service settings. The assignment is read when a new incident is created. Existing incidents retain the escalation work already written to the durable outbox.
Without a policy, Firewatch records the incident but does not page a responder.
## Normalized alert ingestion
[Section titled “Normalized alert ingestion”](#normalized-alert-ingestion)
External systems submit normalized events to:
```text
POST /api/v1/alerts
X-API-Key: fw_.
Content-Type: application/json
```
The organization API key authenticates the sender. The inbound normalized-alert endpoint does **not** currently add a second HMAC signature scheme. If a vendor requires signature validation, validate the vendor request in an adapter before submitting the normalized event with the Firewatch API key.
Important fields have separate jobs:
| Field | Purpose |
| ------------------ | ----------------------------------------------------------------- |
| `eventId` | Stable identity for one provider transition and its retries |
| `deduplicationKey` | Groups repeated firing events into one active incident |
| `correlationKey` | Connects trigger and recovery events when supplied |
| `action` | `trigger` creates/correlates; `resolve` closes active correlation |
| `source` | Stable adapter/provider identifier |
Firewatch serializes concurrent events for the same identity and correlation, stores the normalized event, and returns the original result for an idempotent replay.
See [Alert adapter integration](/docs/integrations/alert-adapters/) for the complete payload and AWS envelope behavior.
## AWS SNS and SQS alerts
[Section titled “AWS SNS and SQS alerts”](#aws-sns-and-sqs-alerts)
The **AWS SNS / SQS** source generates a Lambda adapter that forwards an SNS or SQS event batch to Firewatch. It understands:
* direct SNS notifications;
* SQS messages;
* SNS notifications delivered through SQS;
* CloudWatch alarm state changes; and
* already normalized Firewatch alert objects.
`ALARM` triggers an incident and `OK` resolves it. The Lambda uses a Firewatch organization API key; Firewatch does not store the alert-source AWS credentials.
This alert source is independent of Firewatch’s background-work queue. Self-hosters can ingest AWS alerts while continuing to use PostgreSQL for background work.
## Source and recovery links
[Section titled “Source and recovery links”](#source-and-recovery-links)
When provided, `sourceUrl` must be an absolute HTTP or HTTPS URL and is displayed with the incident’s alert event. Treat it as responder-visible data; do not put credentials or sensitive query parameters in it.
# Alert lifecycle
> How Firewatch turns an incoming signal into an owned incident response.
Firewatch separates the signal that something happened from the decisions needed to reach and coordinate the right responder.
## 1. Ingest
[Section titled “1. Ingest”](#1-ingest)
An alert source sends a versioned request with an organization API key. Firewatch applies authentication, idempotency, and correlation before the alert event changes response state. The normalized endpoint does not add a second inbound HMAC scheme: an adapter must validate any provider-specific signature before it submits the normalized event.
## 2. Resolve ownership
[Section titled “2. Resolve ownership”](#2-resolve-ownership)
The affected service selects an escalation policy. A step can resolve the current responder from an on-call schedule, all primary service owners, or all service owners. Schedule targets honor the schedule time zone, participant effective dates, weekday coverage, and active overrides when the step becomes due.
## 3. Deliver
[Section titled “3. Deliver”](#3-deliver)
At each due step, Firewatch attempts every deployment-enabled channel available for the target: email, Slack DM, SMS, optional voice, and each enabled organization outbound webhook. Each attempt remains independently observable, so one provider failure does not erase the incident or prevent other channels from being attempted.
## 4. Escalate or acknowledge
[Section titled “4. Escalate or acknowledge”](#4-escalate-or-acknowledge)
Acknowledgement assigns visible ownership. If nobody acknowledges before the next relative delay expires, the policy advances without rewriting history. Acknowledging or resolving records later queued steps as skipped.
## 5. Coordinate and learn
[Section titled “5. Coordinate and learn”](#5-coordinate-and-learn)
The incident timeline, responders, runbooks, and delivery attempts preserve a single operational record. Templates can make delivery channel-specific while keeping the underlying incident facts consistent.
# Firewatch architecture
> Understand Firewatch's processes, persistence, queueing, authentication, and deployment boundaries.
## Status and guiding constraint
[Section titled “Status and guiding constraint”](#status-and-guiding-constraint)
Firewatch is a modular on-call and incident-alerting product. Schedules, escalation execution, alert ingestion, incidents, runbooks, safe message templates, and email, Twilio, Slack DM, and signed-webhook delivery are implemented across independently deployable processes.
Firewatch has one product implementation with multiple independently deployable processes. The public repository is the source of truth for the domain, application behavior, HTTP API, worker, web application, and portable deployment artifacts. Self-hosted and managed installations use the same product behavior.
flowchart TB accTitle: Firewatch runtime architecture accDescr: The React web app calls the API, which persists work to PostgreSQL for the worker to deliver through provider abstractions. Web\["React / Vite Web"] -->|"HTTP /api/v1"| Api\["ASP.NET Core API"] Mobile\["Future native clients"] -.->|"same public API"| Api Api --> Db\["PostgreSQL\
state + outbox"] Db --> Worker\[".NET Worker"] Worker --> Delivery\["Delivery-provider\
abstractions"] Delivery --> Channels\["Email · SMS · voice\
Slack · signed webhooks"] Worker --> Queue\["Queue-provider\
abstraction"] Queue --> PgQueue\["PostgreSQL queue\
self-hosted default"] Queue -.-> Sqs\["AWS SQS\
optional provider"] Api --> Slack\["Slack OAuth\
member mapping"] Slack -.-> Channels
Firewatch keeps API, worker, and provider boundaries consistent across deployment models.
Enable JavaScript to view this architecture diagram.
The arrows describe the durable asynchronous path, not separate product implementations. The API and worker share libraries and a database schema but run, scale, restart, and deploy independently.
## Repository boundaries
[Section titled “Repository boundaries”](#repository-boundaries)
| Project | Responsibility | May depend on |
| --------------------------------- | --------------------------------------------------------------------- | ---------------------------------------------------------- |
| `Firewatch.Domain` | Pure business types and rules | Nothing outside the base class library |
| `Firewatch.Contracts` | Versioned HTTP and message DTOs | Base class library only |
| `Firewatch.Application` | Use cases, ports, and orchestration | Domain and contracts |
| `Firewatch.Infrastructure` | EF Core context, PostgreSQL persistence, outbox, repositories, DI | Application, contracts, queue abstractions, EF Core/Npgsql |
| `Firewatch.Queueing.Abstractions` | Small publishing/consuming contracts and message identity conventions | Base class library and stable contracts as needed |
| `Firewatch.Queueing.Postgres` | Portable PostgreSQL queue/outbox processing | Queue abstractions and PostgreSQL infrastructure |
| `Firewatch.Queueing.Sqs` | SQS publishing, long polling, acknowledgement, and visibility leases | Queue abstractions and AWS SDK for SQS |
| `Firewatch.Storage.S3` | Optional runbook-image object storage | Application asset-store port and AWS SDK for S3 |
| `Firewatch.Email.Ses` | Optional transactional email through the SESv2 API | Application email port, infrastructure options, AWS SDK |
| `Firewatch.Api` | REST endpoints, HTTP concerns, health, OpenAPI, composition root | Application and concrete runtime adapters |
| `Firewatch.Worker` | Hosted consumers, retries, outbox work, composition root | Application and concrete runtime adapters |
The API never exposes EF Core entities. Versioned types from `Firewatch.Contracts` form the wire boundary. Endpoint registration is kept out of `Program.cs` so the composition root stays readable.
This is a modular monolith: modules share a release and relational data store, while dependencies point inward toward domain and application code. It avoids network boundaries inside the product until there is evidence that one is needed.
## Deployable processes
[Section titled “Deployable processes”](#deployable-processes)
### API
[Section titled “API”](#api)
`Firewatch.Api` owns synchronous HTTP handling under `/api/v1`, OpenAPI, exception translation, request/correlation IDs, structured JSON logs, and liveness and readiness endpoints. It is stateless apart from PostgreSQL and can be scaled horizontally after migrations have been run.
### Worker
[Section titled “Worker”](#worker)
`Firewatch.Worker` owns asynchronous execution. It polls or consumes through the selected queue provider, dispatches durable outbox work, records completion, and honors cancellation during shutdown. API replicas do not run worker loops, and worker replicas do not serve the public API.
### Web
[Section titled “Web”](#web)
`apps/web` is a standalone React/Vite application. It communicates only through the public HTTP API by using `packages/api-client`; it does not reference .NET projects or persistence types. The production web container serves static files and proxies same-origin `/api`, `/health`, and `/openapi` requests to the API upstream selected at deployment time.
## Authentication and tenant authorization
[Section titled “Authentication and tenant authorization”](#authentication-and-tenant-authorization)
ASP.NET Core Identity owns local user credentials and account-security behavior. A user has a first name, last name, and normalized unique email address. Users do not own tenant data directly: an organization membership connects a user to an organization and records whether that member is an organization administrator. Organization-scoped use cases resolve the active tenant from an authenticated membership and enforce that boundary in application and persistence queries; they do not trust an arbitrary organization identifier supplied by a client.
The browser signs in with an HTTP-only, Secure authentication cookie. Requests that change state also require a CSRF token, including requests made through the same-origin web proxy. Cookie contents are protected by the ASP.NET Core Data Protection key ring stored in the shared PostgreSQL database, so horizontally scaled API replicas and replacement containers share the same keys. Production operators remain responsible for protecting that key ring at rest and rotating its deployment-managed protecting key.
Organization API keys are tenant-scoped machine credentials rather than user sessions. An organization administrator can create and revoke them. The full secret is returned only at creation, while Firewatch stores a non-reversible hash for constant-time verification. Authorization uses the key’s organization and active/expiry state and never treats an API key as a global administrator credential. Automated organization provisioning uses a separate workload-token policy and does not reuse tenant API keys.
## Services and incidents
[Section titled “Services and incidents”](#services-and-incidents)
A monitored service is an organization-scoped thing that can degrade or fail. Its normalized name is unique within the organization. A service always begins with at least one owner: the member who created it as a primary owner. Owners carry a persisted `primary` or `secondary` label; multiple owners may use either label, API responses order primary owners first, and both labels grant the same service-management permission. Service owners and organization administrators can edit the service, add or remove active members as owners, and archive it. The domain rejects removal of the final owner, and organization-member removal is blocked when it would orphan an active service. Archival is non-destructive: archived services retain their owners, incidents, alert events, and runbook associations but cannot receive new incidents or be edited. A service may select one escalation policy for newly created incidents. Alert sources and runbooks reference this same service identifier rather than introducing separate copies of the service concept.
An incident belongs to exactly one service and organization. It has a global, human-readable `INC-000001` reference, severity (`SEV1`, `SEV2`, or `SEV3`), status (`triggered`, `acknowledged`, or `resolved`), and a source kind. Manual declaration accepts the actual incident start time rather than assuming the API request time. Resolution likewise records its selected time and an optional reason. Manual declarations and normalized alert ingestion create the same incident aggregate; alert-created incidents retain their source events and correlation identity.
Every incident owns a chronological timeline. Creation, acknowledgement, resolution (including its optional reason), reopening, responder notes, severity changes, escalation stages, notification deliveries, failed attempts, and skipped notifications append new immutable entries containing the actor, status transition, optional severity transition, message, and timestamp. Reopening begins a new response cycle without removing the prior cycle’s events. The domain exposes no mutation or removal operation for entries, and the EF context rejects modified or deleted timeline rows. Incident listing performs tenant-scoped server-side search, filtering, and pagination; dashboard counts and recent incidents are computed from the same persisted data. Tenant identity is repeated on timeline rows and enforced through composite foreign keys so an incident, service, and timeline cannot cross organization boundaries.
## Time zones, schedules, and escalations
[Section titled “Time zones, schedules, and escalations”](#time-zones-schedules-and-escalations)
An organization stores a default IANA timezone for new configuration, while every user stores their own display timezone. Neither value silently overrides a schedule: each on-call schedule explicitly owns the IANA timezone that defines its wall-clock rotation start and handoff cadence. This keeps a 09:00 handoff at 09:00 through daylight-saving transitions. Notification content converts dates to the recipient’s personal timezone.
A schedule contains an ordered rotation of organization members. A rotation can use one fixed segment length or a repeating pattern such as `4,3`, allowing one schedule to alternate Sun–Wed and Thu–Sat coverage. Schedules also store the weekdays that provide coverage; excluded weekdays create explicit gaps and do not advance the pattern. Each participant can have optional effective-from and effective-until instants, so planned hires, leave, and removals are reflected in future calendars without a precisely timed manual edit. Membership changes are treated as schedule transitions and the active roster is recalculated at the effective instant.
Non-overlapping overrides use exact start and end instants and temporarily supersede both the rotation and weekday exclusions. Organization member removal is blocked while that member remains referenced by a rotation.
An escalation policy contains up to ten ordered steps. Step one is immediate; each later step stores a delay relative to the preceding step. A step targets an on-call schedule, the service’s primary owners, or all service owners. When an incident is created for a service with a policy, the API snapshots every step into a delayed outbox message in the same transaction. The worker resolves the current schedule responder or service owners when each step becomes due, so rotation changes and active overrides are honored at delivery time. Acknowledgement or resolution causes remaining stages to be audited as skipped.
Registration is deployment-configurable. Self-hosted Compose selects `Bootstrap`, so exactly the first administrator can create an organization and choose a passkey (recommended) or password without needing an outbound mail server. Registration then closes automatically and additional members join by administrator invitation. `Open` deliberately keeps public registration available and may require email ownership verification according to deployment configuration. `Provisioned` removes public registration while allowing an authenticated provisioning service to create an organization and its initial administrator. It can support a self-service checkout flow without exposing an unauthenticated organization-creation endpoint. `Disabled` is the fail-closed product default. The same web build reads the registration capability from the API rather than embedding a deployment-edition switch.
ASP.NET Core Identity 10 owns passkey attestation, assertion, replay counters, and credential storage. Public passkeys are discoverable and require local user verification, enabling username-less sign-in. The explicit RP ID is inferred from the canonical public URL unless an operator deliberately overrides it. An authenticated user can keep up to ten named passkeys. Credential removal is serialized per user, and a passkey-only account cannot remove its final passkey.
Local Identity credentials are one authentication method, not a permanent identity-provider boundary. External login records are modeled by stable provider and provider-subject identifiers. The public API exposes a trusted, explicitly configured extension loader plus a provider-neutral federated sign-in service. That service owns account linking, organization-boundary enforcement, optional JIT membership creation, invitation completion, and the normal Identity cookie. Protocol adapters validate their upstream response and submit only the resulting identity to this service.
Trusted authentication extensions can validate an external identity, veto local credential login, add product routes, and register a namespaced PostgreSQL context whose migrations participate in the one-off migration mode. A missing, invalid, or duplicate configured extension fails API startup. Self-hosted deployments load no extension by default.
## Durable work and delivery semantics
[Section titled “Durable work and delivery semantics”](#durable-work-and-delivery-semantics)
The PostgreSQL outbox closes the failure window between changing application state and requesting asynchronous work:
1. An application use case stores its state change and a uniquely identified outbox record in one database transaction.
2. A worker claims pending work without allowing another worker to own the same lease concurrently.
3. The selected provider publishes or handles the message.
4. Completion is recorded durably. Notification deliveries store every attempt, recipient, channel, outcome, and error. Failures use exponential backoff and become dead letters after the configured maximum.
Delivery is **at least once**. A stable message identifier travels with every message; it is not regenerated during retry. Consumers must use that identifier as an idempotency key and make their state change plus processed-message marker atomic wherever possible. A handler must assume it can run again after a timeout, lost acknowledgement, process crash, or queue visibility expiry.
The abstractions intentionally cover only publishing and consuming Firewatch messages. They are not a general-purpose event bus. PostgreSQL is the default and requires no additional broker beyond the database Firewatch already needs. Its `FOR UPDATE SKIP LOCKED` claim path supports multiple worker replicas. When SQS is selected, committed outbox rows are relayed to SQS; workers long-poll and use a durable processed-message record plus a PostgreSQL advisory lock for idempotent redelivery. Email delivery uses `IEmailSender`; phone verification and delivery use `IPhoneSender`. SMTP, SES, and optional Twilio adapters do not change incident behavior. Slack delivery uses the immutable user ID captured by verified-email mapping. Each delivered Slack alert stores its bot-message reference. Signed Block Kit actions mutate the incident in the API transaction, and the resulting outbox event makes the worker update the original message and append a lifecycle thread reply. Outbound webhooks use a stable delivery ID plus timestamped HMAC signature. Organization template records are shared through PostgreSQL, avoiding divergent template files across worker replicas.
## Architectural smoke flow
[Section titled “Architectural smoke flow”](#architectural-smoke-flow)
The asynchronous processing smoke flow is authenticated system-test infrastructure:
1. `POST /api/v1/system/test-jobs` creates a test-job row and its outbox message in one PostgreSQL transaction.
2. The API returns `202 Accepted`, including the stable job identifier and a status location.
3. The separately running worker claims and processes the outbox item, logs its message identifier, and marks the test job complete.
4. A duplicate delivery observes the processed identifier and does not perform duplicate completion.
5. `GET /api/v1/system/test-jobs/{id}` reports the state to the web application or a developer.
This flow validates the architecture; it is not a model for alerts, incidents, or escalation policies.
## Health and observability
[Section titled “Health and observability”](#health-and-observability)
* `GET /health/live` reports that the API process can serve requests. It does not fail merely because PostgreSQL is unavailable.
* `GET /health/ready` verifies required dependencies, including PostgreSQL, and is the deployment readiness gate.
* API and worker logs are structured JSON and include message or request context plus correlation/request identifiers where available. Deployments keep the process streams distinct and may add service-name enrichment in their log collection layer.
* OpenTelemetry registration is an extension point for traces and metrics. An exporter or full observability backend is deliberately not bundled.
Liveness should restart a stuck process. Readiness should remove an unhealthy API from traffic. Neither endpoint should leak credentials or detailed exception data.
## One product, two deployment models
[Section titled “One product, two deployment models”](#one-product-two-deployment-models)
For self-hosting, Compose runs PostgreSQL, API, worker, and web, with the PostgreSQL queue provider. A future public Helm chart may package the same containers.
Managed installations configure the same provider interfaces and product contracts while operating infrastructure, secrets, backups, and delivery providers for customers. Private infrastructure does not introduce a second copy of Firewatch’s domain behavior.
## Future iOS client
[Section titled “Future iOS client”](#future-ios-client)
A native iOS application will be another client of the same versioned REST API; the web app is not a backend-for-frontend. The checked OpenAPI document can be fed to a maintained Swift OpenAPI generator (for example Apple’s Swift OpenAPI Generator) when iOS work begins. Authentication and API version rules live at the public API boundary rather than in a web-only endpoint; future compatibility changes must preserve that boundary.
# Firewatch documentation
> Use, administer, integrate, and operate Firewatch on-call alerting.
Learn how Firewatch routes alerts, manages on-call coverage, escalates missed pages, and keeps responders informed.
[Start reading](/docs/getting-started/)
[Browse integrations](/docs/integrations/)
## Machine-readable documentation
[Section titled “Machine-readable documentation”](#machine-readable-documentation)
Agents and tooling can use the documentation files generated from this same source. Start with the concise index, then choose the abridged or complete content set for the context window you have available.
* [Documentation index](/llms.txt)
* [Abridged documentation](/llms-small.txt)
* [Complete documentation](/llms-full.txt)
## Choose a guide
[Section titled “Choose a guide”](#choose-a-guide)
Get started
Configure your profile, service, schedule, escalation policy, and first alert. [Follow the setup path](/docs/getting-started/).
Understand alert routing
Learn how services, schedules, escalation policies, and notification channels determine who is contacted. [Read the alert lifecycle](/docs/concepts/alert-lifecycle/).
Connect integrations
Configure alert sources and outbound email, Slack, SMS, voice, and webhook delivery. [Review integrations](/docs/integrations/).
Build with the API
Discover the versioned REST API and its generated OpenAPI contract. [Open the API guide](/docs/api/).
Operate a self-hosted installation
Install the complete product with Docker Compose and configure the services you operate. [Open the self-hosting guide](/docs/self-hosting/).
## Set up your on-call workflow
[Section titled “Set up your on-call workflow”](#set-up-your-on-call-workflow)
1. **Understand the alert lifecycle** from ingestion through acknowledgement, escalation, and resolution.
2. **Create a service, schedule, and escalation policy** that reflect your actual responder coverage.
3. **Connect an alert source** and configure the delivery channels you operate.
4. **Send a test alert** and verify acknowledgement, escalation, and recovery before relying on the system for production paging.
Before relying on any paging system in production, validate alert delivery, escalation, backup, and recovery behavior against your organization’s availability, security, and compliance requirements.
# Organization administration
> Manage members, duplicate users, routing integrations, templates, and API keys.
## Access roles and responder seats
[Section titled “Access roles and responder seats”](#access-roles-and-responder-seats)
Organization administrators can invite members, promote or demote administrators, and remove members. Invitations are single-use and expire.
Firewatch stores two independent membership settings:
* **Access role:** an Administrator can manage organization configuration; a Member has normal product access.
* **On-call access:** a Responder can own services, join schedules, receive overrides, and participate in shift trades; a Collaborator can view and collaborate without being eligible for paging.
An administrator can be either a responder or collaborator. This keeps a billing or security administrator free when they never participate in on-call. Firewatch Cloud counts responder memberships as paid seats; collaborator memberships are free.
For Firewatch Cloud, assigning responder access at any point counts as one paid responder seat for the full billing month. Changing that person back to a collaborator does not remove the charge for that month. This prevents rapid role changes from bypassing seat billing.
Members can invite free collaborators. Only administrators can add a responder, invite another administrator, or change on-call access.
Before changing a responder to a collaborator, remove their service ownership, schedule participation, overrides, and pending shift trades. Firewatch rejects the change while any routing reference remains.
Firewatch blocks removals that would leave active product records without a valid owner or responder. Resolve those references before trying again.
Only administrators can change roles, merge users, configure organization integrations, edit alert templates, or manage API keys.
## Merge duplicate users
[Section titled “Merge duplicate users”](#merge-duplicate-users)
Use **Organization → Members → Merge duplicate** when two active accounts represent the same person after an identity-provider change or duplicate registration.
Choose:
* the **source** account that will be retired; and
* the **target** account whose profile and email will remain.
Enter the target email exactly to confirm. Firewatch then:
* preserves the stronger administrator role;
* transfers non-duplicate service ownership and schedule participation;
* transfers schedule overrides, external-login records, passkeys, and claims;
* transfers the Slack mapping when it does not conflict;
* adopts the source’s verified phone only when the target lacks one;
* retains responder access when either merged account was a responder;
* redirects existing merged-user references to the target; and
* disables and audits the source account.
Both accounts must be active, email-confirmed members of only the current organization. The source cannot be the administrator’s current session. Resolve pending invitations and source-account shift trades first, and keep the combined passkey count at ten or fewer.
Merging is not a reversible UI operation. Confirm the retained account and email carefully.
## Slack workspace
[Section titled “Slack workspace”](#slack-workspace)
When the deployment has a Slack app configured, open **Organization → Integrations** and select **Add to Slack**. A Slack administrator must approve the requested bot scopes.
Synchronization compares current, verified Firewatch member emails with Slack and stores immutable Slack user IDs. Firewatch:
* does not import the Slack directory as Firewatch users;
* does not guess using display names; and
* sends alerts by direct message, not by mentioning a person in a shared channel.
Unmatched members continue to receive their other available channels. Re-run synchronization after membership or Slack-email changes.
Slack interaction requests are validated with Slack’s timestamped signature. Acknowledge and resolve buttons update Firewatch, then the worker updates the original alert message and adds lifecycle context.
## Alert templates
[Section titled “Alert templates”](#alert-templates)
Open **Organization → Alert templates** to customize plain-text email subject and body, SMS, voice, Slack, and webhook wording.
Templates use only the variables shown in the page. They cannot execute code, read secrets, or emit arbitrary email HTML. Firewatch validates length and syntax before saving and safely creates the email HTML alternative.
Use **Restore defaults** to remove the organization customization.
## Outbound webhooks
[Section titled “Outbound webhooks”](#outbound-webhooks)
An organization can configure up to ten signed delivery endpoints. The signing secret appears only at creation or rotation. Pausing an endpoint retains its configuration without sending new deliveries.
Rotation invalidates the previous secret immediately. See [Outbound alert webhooks](/docs/integrations/outbound-webhooks/) for signature verification and destination restrictions.
## API keys
[Section titled “API keys”](#api-keys)
Organization API keys authenticate machine integrations such as alert adapters. The full secret is shown only when created; Firewatch stores a non-reversible validation value.
Keep keys in a secret manager, use separate keys for separate senders, and revoke a key when its owner or purpose changes. An organization API key is not a global Firewatch administrator credential.
## Self-hosted provider settings
[Section titled “Self-hosted provider settings”](#self-hosted-provider-settings)
Self-hosted administrators can revisit **Admin settings** after the initial setup. The page detects email, Twilio, and Slack status and generates the environment changes selected by the administrator.
The UI does not write the host’s `.env` file. Copy the generated values into deployment secret storage, run `docker compose up -d`, then re-check the status. Firewatch Cloud does not expose provider credentials to customer administrators.
# API overview
> Discover and use Firewatch’s versioned REST API.
Firewatch exposes a versioned REST API and generates an OpenAPI document from the same contracts used by the application.
## Discovery
[Section titled “Discovery”](#discovery)
The marketing domain publishes an [RFC 9727 API catalog](/.well-known/api-catalog) with links to the application API, OpenAPI description, and these human-readable docs.
The application OpenAPI document is available at:
```text
https://app.onfirewatch.com/openapi/v1.json
```
Self-hosted installations expose the same document at `/openapi/v1.json` on the API origin.
## Authentication
[Section titled “Authentication”](#authentication)
Browser sessions use same-origin secure cookies and antiforgery protection. Organization API keys are intended for service integrations and alert ingestion. Keep API keys in secret storage and rotate them when their scope or owner changes.
## Stability
[Section titled “Stability”](#stability)
The `/api/v1` prefix is the compatibility boundary. Additive fields may appear within a version; breaking contract changes require a new API version.
# Development guide
> Set up the Firewatch toolchain and run the product, tests, and API generation locally.
## Toolchain
[Section titled “Toolchain”](#toolchain)
Install the following tools on the development host:
* .NET 10 SDK;
* Node.js 24 or another version satisfying the root `engines` field, with the npm version recorded by `packageManager`;
* Docker Engine with the Compose v2 plugin;
* Git; and
* optionally `psql` for database inspection.
## Initial setup
[Section titled “Initial setup”](#initial-setup)
From the repository root, run the reproducible setup command:
```bash
dotnet tool restore && dotnet restore Firewatch.sln && npm ci
```
Re-run the setup command whenever the lock files, tool manifest, or project graph changes.
## Run the complete stack
[Section titled “Run the complete stack”](#run-the-complete-stack)
```bash
docker compose up --build
```
Open . The API is also published at . Compose waits for PostgreSQL readiness, starts the API and its local migration step, waits for API readiness, then starts the independent worker and web processes.
The base stack bootstraps the first self-hosted administrator without an email round trip and leaves email delivery disabled. To test verification and SMTP delivery with Mailpit at , start the explicit development override:
```bash
docker compose -f docker-compose.yml -f docker-compose.mailpit.yml up --build
```
Useful commands:
```bash
docker compose ps
docker compose logs --follow api worker
docker compose down
```
`docker compose down` preserves the PostgreSQL named volume. Use `docker compose down --volumes` only when you intentionally want to erase local Firewatch data.
## Run processes directly
[Section titled “Run processes directly”](#run-processes-directly)
Keep infrastructure in Docker:
```bash
docker compose -f docker-compose.yml -f docker-compose.mailpit.yml up -d postgres mailpit
```
In one terminal, configure a host-reachable database and run the API:
```bash
export ConnectionStrings__Firewatch='Host=localhost;Port=5432;Database=firewatch;Username=firewatch;Password=change-me-for-non-local-use'
export Database__MigrateOnStartup=true
export Auth__PublicBaseUrl=http://localhost:5173
export Email__Enabled=true
export Email__Smtp__Host=localhost
export Email__Smtp__Port=1025
export Email__Smtp__Security=None
dotnet run --project src/Firewatch.Api
```
In a second terminal, export the same connection string, select the local queue provider, then run:
```bash
export ConnectionStrings__Firewatch='Host=localhost;Port=5432;Database=firewatch;Username=firewatch;Password=change-me-for-non-local-use'
export Queueing__Provider=Postgres
dotnet run --project src/Firewatch.Worker
```
In a third terminal, start Vite:
```bash
npm run dev
```
Open for the direct Vite workflow.
Vite is the only direct development server. The React application still uses the HTTP API through its same-origin development proxy; it never imports backend implementation code. `VITE_DEV_API_PROXY_TARGET` changes the proxy target from its `http://localhost:8080` default while preserving same-origin browser calls. Set `VITE_API_BASE_URL` only when intentionally bypassing that proxy for a different public API origin.
## Exercise the smoke flow
[Section titled “Exercise the smoke flow”](#exercise-the-smoke-flow)
With the API and worker running:
```bash
curl --include \
--request POST \
--header 'Content-Type: application/json' \
--data '{"name":"documentation-smoke-test"}' \
http://localhost:8080/api/v1/system/test-jobs
```
The response is `202 Accepted` and includes an identifier and status URL. Replace `` below with the returned identifier:
```bash
curl http://localhost:8080/api/v1/system/test-jobs/
docker compose logs worker
```
The status eventually becomes `Completed`, and the worker log includes the stable message identifier. Retrying delivery does not create a second completion.
## Tests and quality checks
[Section titled “Tests and quality checks”](#tests-and-quality-checks)
Run backend checks:
```bash
dotnet format Firewatch.sln --verify-no-changes --no-restore
dotnet build Firewatch.sln --configuration Release
dotnet test Firewatch.sln --configuration Release --no-build
```
Integration tests use Testcontainers with real PostgreSQL behavior. Docker must be running; SQLite is intentionally not used as a substitute.
Run frontend checks:
```bash
npm run frontend:check
```
Or run the constituent commands:
```bash
npm run lint
npm run format:check
npm run typecheck
npm test
npm run build
npm run api:check
```
## OpenAPI and TypeScript client generation
[Section titled “OpenAPI and TypeScript client generation”](#openapi-and-typescript-client-generation)
The committed `openapi/firewatch.v1.json` document is the language-neutral API contract. After changing API routes or contracts, first regenerate the document from the ASP.NET Core project using the repository’s deterministic OpenAPI generation command, then generate the TypeScript types:
```bash
dotnet build src/Firewatch.Api/Firewatch.Api.csproj --configuration Release
npm run api:generate
npm run api:check
```
`openapi-typescript` is used because it is maintained, produces small static types without a proprietary runtime, and composes with the maintained `openapi-fetch` client. Generated files under `packages/api-client/src/generated` must not be manually edited. The handwritten package wrapper is the stable import surface consumed by `apps/web`.
`npm run api:check` regenerates to a temporary file and compares bytes, so CI fails when the checked-in client does not match the checked-in OpenAPI document. Both files must be committed with an API change.
A future iOS repository can feed the same OpenAPI file to Apple’s Swift OpenAPI Generator (with versions pinned by that repository). No iOS application or Swift-specific API is created here.
## EF Core migrations
[Section titled “EF Core migrations”](#ef-core-migrations)
Restore the pinned local `dotnet-ef` tool before migration work:
The design-time fallback matches the unchanged `.env.example` database. Export `ConnectionStrings__Firewatch` first when using different credentials or a different database.
```bash
dotnet tool restore
```
Create a migration after changing the EF model:
```bash
dotnet ef migrations add \
--project src/Firewatch.Infrastructure \
--startup-project src/Firewatch.Api \
--output-dir Persistence/Migrations
```
Apply it to the configured database:
```bash
dotnet ef database update \
--project src/Firewatch.Infrastructure \
--startup-project src/Firewatch.Api
```
Verify that the model and migration snapshot agree, as backend CI does:
```bash
dotnet ef migrations has-pending-model-changes \
--project src/Firewatch.Infrastructure \
--startup-project src/Firewatch.Api
```
Containerized deployments should use the API image’s one-off migration mode:
```bash
docker compose run --rm api --migrate
```
Do not enable automatic startup migrations for a multi-replica deployment. Use forward-compatible expand/contract changes and run each migration once before replacing application replicas.
## Common problems
[Section titled “Common problems”](#common-problems)
* **Readiness returns 503:** inspect `docker compose logs postgres api` and verify `ConnectionStrings__Firewatch` and the PostgreSQL health check.
* **Integration tests cannot start:** verify `docker info` works from the same terminal and that the Docker daemon is running.
* **The web page loads but API calls fail:** for Vite, check its `/api` proxy or an explicitly set `VITE_API_BASE_URL`; for the web container, check `FIREWATCH_API_UPSTREAM` and API readiness.
* **Registration says email setup is required:** the API is configured to require verification. Start the Mailpit override or configure `Email__Enabled`, the SMTP settings, and `Auth__PublicBaseUrl`.
* **Generated client check fails:** regenerate with `npm run api:generate`, review the API diff, and commit the generated file together with the schema change.
# Local email testing
> Test SMTP and account-email flows locally with the opt-in Mailpit development override.
Mailpit is an opt-in development-only SMTP inbox. It is not included in the default self-hosted stack, which bootstraps the first administrator without email verification and leaves email alerts disabled.
Start Firewatch with the local email override:
```bash
docker compose -f docker-compose.yml -f docker-compose.mailpit.yml up -d --build
```
The override enables email verification and starts Mailpit. Omit it when testing the normal self-hosted first-run experience.
Open Firewatch at `http://localhost:3000` and Mailpit at `http://localhost:8025`. Mailpit captures messages sent through its SMTP port; it does not deliver them to the public internet.
The override enables email in both the API and worker and connects them to `mailpit:1025` without authentication or TLS. The worker sends incident alerts; the API sends account and invitation mail. Port `1025` is also bound to loopback so a process launched directly on the host can use `localhost:1025`. It also permits the Vite development origin at `http://localhost:5173` for passkey ceremonies while the canonical Docker web origin remains `http://localhost:3000`.
To stop only the development inbox:
```bash
docker compose -f docker-compose.yml -f docker-compose.mailpit.yml stop mailpit
```
Do not rely on Mailpit in production. Configure the standard `Email__*` settings against an authenticated, TLS-enabled SMTP service instead.
# Integrations
> Connect alert sources and notification providers to Firewatch.
Firewatch integrations fall into two groups.
## Alert sources
[Section titled “Alert sources”](#alert-sources)
Alert sources create or update Firewatch alert events. Firewatch accepts its versioned alert-ingestion format directly and accepts AWS SNS or SQS event batches through the generated Lambda adapter. API keys scope each source to an organization, while idempotency and correlation keys prevent repeated provider delivery from creating duplicate response state.
The normalized inbound endpoint does not currently verify a generic HMAC signature in addition to the Firewatch API key. When a monitoring provider signs its webhooks, validate that provider signature and freshness in the adapter before forwarding the normalized event.
Use the [alert adapter guide](/docs/integrations/alert-adapters/) when a monitoring provider needs a small normalization layer before it calls Firewatch.
## Notification providers
[Section titled “Notification providers”](#notification-providers)
Notification providers reach responders or downstream systems:
* SMTP for incident and account-security email;
* Slack for direct responder alerts and incident context;
* Twilio for verified SMS and optional voice delivery;
* outgoing signed webhooks for customer-owned automation.
Self-hosters bring their own provider credentials. Firewatch Cloud manages provider infrastructure according to the subscribed plan.
Provider secrets belong in deployment secret storage, never in source control or browser-visible configuration.
Outbound customer automation should use [signed outbound webhooks](/docs/integrations/outbound-webhooks/). Slack interaction requests use Slack’s timestamped signing contract; that is separate from alert-source authentication and outbound webhook signatures.
# Alert adapter integration
> Normalize monitoring-provider events and submit them through Firewatch's alert ingestion API.
Alert-source adapters normalize vendor webhooks and submit them to:
```text
POST /api/v1/alerts
X-API-Key: fw_.
Content-Type: application/json
```
Create the API key under **Settings → API keys** and configure a service before sending alerts. The standalone Apache-2.0 `@onfirewatch/sdk` package contains the typed client, adapter framework, and Prometheus Alertmanager adapter.
The dashboard’s **Alert sources** page generates the configuration for a selected service and adapter. Its test action sends a real SEV3 event through the ingestion, incident, escalation, and notification-outbox pipeline.
## Normalized event
[Section titled “Normalized event”](#normalized-event)
```json
{
"serviceId": "0e48d1d1-a2d6-460b-93a2-cd7a4a86ca0a",
"source": "prometheus-alertmanager",
"eventId": "abc123:trigger:2026-07-20T00:00:00.000Z",
"action": "trigger",
"deduplicationKey": "abc123",
"correlationKey": "abc123",
"title": "Checkout error rate is high",
"description": "More than 10% of checkout requests are failing.",
"severity": "SEV1",
"occurredAt": "2026-07-20T00:00:00Z",
"labels": {
"environment": "production"
},
"sourceUrl": "https://prometheus.example.test/graph"
}
```
`action` is `trigger` or `resolve`. `eventId` must identify one provider state transition and remain stable across retries.
## Delivery guarantees
[Section titled “Delivery guarantees”](#delivery-guarantees)
* **Idempotency:** an organization, source, and event ID can be processed once. A replay returns the original alert-event and incident identifiers.
* **Deduplication:** firing events with the same source, service, and deduplication key attach to the existing active incident.
* **Correlation:** a correlation key connects related trigger and resolution events. Resolution events close the correlated active incident.
* **Concurrency:** PostgreSQL transaction-scoped advisory locks serialize both event IDs and correlation keys, preventing simultaneous deliveries from opening duplicate incidents.
* **Auditability:** normalized alert events are retained in `alert_events`, and correlation and resolution changes are appended to the incident timeline.
* **Escalation:** a newly created alert incident uses the service’s escalation policy and writes notification requests to the same durable outbox used by manually created incidents.
The endpoint returns `201 Created` when it creates an incident and `200 OK` for replays, deduplicated events, resolutions, and unmatched resolutions.
## AWS SNS and SQS
[Section titled “AWS SNS and SQS”](#aws-sns-and-sqs)
Firewatch also accepts the standard AWS Lambda event envelopes produced by SNS triggers and SQS event-source mappings:
```text
POST /api/v1/integrations/aws/alerts/
X-API-Key: fw_.
Content-Type: application/json
```
Choose **AWS SNS / SQS** on the dashboard’s **Alert sources** page to generate a Node.js Lambda handler. Attach an SNS trigger or an SQS event-source mapping to that Lambda and store the Firewatch URL, API key, and service ID as Lambda environment variables.
The adapter accepts batches of up to ten records. It unwraps SNS notifications, SQS bodies, SNS-over-SQS messages, CloudWatch alarm payloads, and normalized Firewatch alert objects. CloudWatch `ALARM` creates or correlates an incident; `OK` resolves it. AWS message IDs provide retry-safe event identity, while alarm ARNs or names provide the deduplication and correlation key.
This is an alert-source integration only. It does not change `Queueing__Provider`: the API and worker continue to use the PostgreSQL outbox unless a deployment separately selects another internal queue implementation.
# Outbound alert webhooks
> Configure signed incident-delivery webhooks and validate Firewatch request signatures.
Organization administrators can configure up to ten outbound endpoints under **Organization → Integrations**. Firewatch sends an `incident.notification.requested` event for each escalation notification. The payload contains structured incident, service, recipient, and escalation data, plus the organization’s rendered webhook message template.
## Request headers
[Section titled “Request headers”](#request-headers)
Every request includes:
| Header | Meaning |
| ----------------------- | ---------------------------------------------------------- |
| `X-Firewatch-Event` | `incident.notification.requested` |
| `X-Firewatch-Delivery` | Stable UUID for this notification delivery and its retries |
| `X-Firewatch-Timestamp` | Unix timestamp in seconds |
| `X-Firewatch-Signature` | `v1=` followed by a lowercase HMAC-SHA256 hex digest |
The signature base is the exact UTF-8 string:
```text
..
```
Compute HMAC-SHA256 with the endpoint’s `fwhsec_...` signing secret and compare the complete `v1=` value in constant time. Read and verify the raw body before JSON parsing; re-serializing JSON changes the signed bytes.
Receivers should also:
1. reject timestamps more than five minutes away from their current time;
2. persist processed `X-Firewatch-Delivery` values and ignore duplicates;
3. return a `2xx` response only after durably accepting the event;
4. keep the signing secret in a secret manager and rotate it from the UI if it may have been exposed.
The signing secret is displayed only when an endpoint is created or rotated. Firewatch stores only its encrypted form. Rotation invalidates the previous secret immediately.
## Node.js verification example
[Section titled “Node.js verification example”](#nodejs-verification-example)
```js
import { createHmac, timingSafeEqual } from "node:crypto";
export function verifyFirewatchWebhook({
secret,
rawBody,
deliveryId,
timestamp,
signature,
now = Date.now(),
}) {
const timestampNumber = Number(timestamp);
if (!Number.isSafeInteger(timestampNumber)) return false;
if (Math.abs(now / 1000 - timestampNumber) > 300) return false;
const base = `${timestamp}.${deliveryId}.${rawBody}`;
const expected = `v1=${createHmac("sha256", secret)
.update(base, "utf8")
.digest("hex")}`;
const expectedBytes = Buffer.from(expected, "ascii");
const actualBytes = Buffer.from(signature, "ascii");
return (
expectedBytes.length === actualBytes.length &&
timingSafeEqual(expectedBytes, actualBytes)
);
}
```
## Destination safety
[Section titled “Destination safety”](#destination-safety)
Redirects are not followed. HTTPS is required, and localhost, private, link-local, and other unsafe network destinations are blocked by default to reduce SSRF risk. A self-hoster can deliberately opt into an internal receiver with `FIREWATCH_WEBHOOK_ALLOW_PRIVATE_NETWORKS=true`; plain HTTP additionally requires `FIREWATCH_WEBHOOK_ALLOW_HTTP=true`. Keep both disabled for an internet-facing deployment unless there is a reviewed operational need.
# Self-hosting Firewatch
> Install, configure, secure, back up, and operate Firewatch with Docker Compose.
## Production readiness
[Section titled “Production readiness”](#production-readiness)
Firewatch includes the complete alert, schedule, escalation, incident, and notification paths needed for self-hosted paging. Before putting the deployment in a critical path, verify its notification providers, escalation behavior, backups, upgrades, and recovery procedures against your availability, security, and compliance requirements.
## Quick start
[Section titled “Quick start”](#quick-start)
Prerequisites are Docker Engine and Docker Compose v2. Clone the public repository, then run:
```bash
docker compose up --build
```
The services are:
| Service | Default host address | Purpose |
| ---------- | ----------------------- | ------------------------------------------------- |
| `web` | `http://localhost:3000` | React application and same-origin API proxy |
| `api` | `http://localhost:8080` | Versioned REST API, OpenAPI, health |
| `worker` | None | Outbox/queue processing in an independent process |
| `postgres` | `127.0.0.1:5432` | Durable state and default queue/outbox provider |
Confirm readiness:
```bash
docker compose ps
curl --fail http://localhost:8080/health/live
curl --fail http://localhost:8080/health/ready
```
The first Compose boot sets `Database__MigrateOnStartup=true` on the single API container. This is a local/simple-hosting convenience so `docker compose up --build` works against a clean volume. It is not the managed or multi-replica migration strategy.
## Configuration and secrets
[Section titled “Configuration and secrets”](#configuration-and-secrets)
The initial boot needs no mail server. Self-hosted administrator registration does not send a verification email, and the email notification channel remains off until SMTP is configured. Before using the stack outside a disposable local machine, copy `.env.example` to the ignored `.env` and change at least `POSTGRES_PASSWORD`. Do not commit `.env`, cloud credentials, provider tokens, connection strings, or database dumps.
The local HTTP stack uses `FIREWATCH_ENVIRONMENT=Development` so ASP.NET can issue session and antiforgery cookies on `http://localhost`. Any deployment reachable by other machines must terminate HTTPS, set `FIREWATCH_ENVIRONMENT=Production`, and set `FIREWATCH_PUBLIC_BASE_URL` to its public `https://` origin. Production mode deliberately refuses to issue these cookies over plain HTTP. It also refuses to start with the documented development database password or an empty database password, and requires a canonical HTTPS public origin.
The default `FIREWATCH_QUEUE_PROVIDER=Postgres` requires no AWS account. API and worker use the same `ConnectionStrings__Firewatch` value assembled by Compose. The SQS adapter is intended for a managed configuration and is not required for self-hosting.
AWS SNS/SQS alert ingestion is independent of the internal queue provider. Self-hosters can keep `FIREWATCH_QUEUE_PROVIDER=Postgres` and use the **Alert sources → AWS SNS / SQS** setup generated in the UI. The generated Lambda forwards SNS or SQS event batches to Firewatch using an organization API key; no AWS credentials are stored by Firewatch.
### Database UI
[Section titled “Database UI”](#database-ui)
The default self-hosted stack does not include a database administration UI. For temporary troubleshooting, start the optional Adminer override:
```bash
docker compose -f docker-compose.yml -f docker-compose.adminer.yml up -d adminer
```
Open `http://localhost:8081` and sign in with:
| Field | Value |
| -------- | ------------------------------------ |
| System | `PostgreSQL` |
| Server | `postgres` |
| Username | the `.env` `POSTGRES_USER` value |
| Password | the `.env` `POSTGRES_PASSWORD` value |
| Database | the `.env` `POSTGRES_DB` value |
The override’s `ADMINER_BIND_ADDRESS` defaults to `127.0.0.1`. Keep that loopback binding: Adminer has direct database access and must not be exposed to the public internet or routed through the product UI. For a remote self-hosted server, leave it on loopback and use an SSH tunnel when access is needed:
```bash
ssh -L 8081:127.0.0.1:8081 your-server
```
Then browse to `http://localhost:8081` on the administrator’s machine. Change `ADMINER_PORT` if port 8081 is already in use. Pin updates by changing `ADMINER_VERSION`, reviewing the release, and recreating the service.
Compose defaults to `FIREWATCH_REGISTRATION_MODE=Bootstrap`, which permits exactly the first self-hosted administrator to create an organization and then closes registration automatically. No ownership email is sent in this mode. The administrator chooses a discoverable passkey (recommended) or a password of at least 12 characters, then follows the in-product setup guide to set the organization time zone, confirm the public URL, and review optional delivery providers. Additional users join through administrator invitations. `Open` deliberately enables ongoing public registration, while `Disabled` closes it; existing users can still sign in. `Provisioned` is reserved for authenticated organization automation and is not a self-hosting bootstrap mechanism.
`FIREWATCH_SELF_REGISTRATION_ENABLED` is a separate kill switch that takes precedence over the selected mode. It defaults to `true` so a new installation can create its first administrator. Set it to `false` after bootstrap when you want to guarantee invitation-only access. Turning it off does not affect sign-in or acceptance of administrator invitations.
`FIREWATCH_USER_PROFILE_EDITING_ENABLED=true` allows signed-in users to update their own first and last names and is the self-hosted default. Set it to `false` when those identity attributes are managed centrally; the API will advertise the restriction to the web app and reject profile-update requests directly.
Firewatch sends incident email and account-security notifications through a configurable SMTP server. Both are disabled by default. The setup guide reports whether SMTP is active and generates the required `.env` block. Configure the `FIREWATCH_SMTP_*` values and use a sender domain your SMTP service is authorized to send as. `StartTls` is the normal setting for port 587; `SslOnConnect` supports implicit TLS, commonly on port 465. Store SMTP passwords as deployment secrets rather than committing them.
After changing deployment variables, run `docker compose up -d` so the API and worker receive the same configuration. During first-run setup, use **Re-check** on the relevant step. After setup, verify provider status under the applicable organization settings page. Mailpit remains available only as an explicit development override; it is not part of the self-hosted stack.
### Phone delivery with Twilio
[Section titled “Phone delivery with Twilio”](#phone-delivery-with-twilio)
Set `FIREWATCH_TWILIO_ENABLED=true` and supply the account SID, auth token, and an SMS/voice-capable E.164 sender in `FIREWATCH_TWILIO_*`. Users can then verify one primary E.164 phone number under **Settings → Profile**. Verified numbers receive incident SMS by default. Automated voice calls are disabled unless `FIREWATCH_TWILIO_INCIDENT_VOICE_ENABLED=true` because calls are more intrusive and can have materially different provider costs. These credentials are global deployment secrets and are never exposed to users.
### Slack app
[Section titled “Slack app”](#slack-app)
Set `FIREWATCH_SLACK_ENABLED=true` with a Slack app client ID, client secret, and signing secret to make **Organization → Integrations → Add to Slack** available. Register `/api/v1/organization/integrations/slack/callback` as the OAuth redirect URL. Enable **Interactivity & Shortcuts** and register `/api/v1/integrations/slack/interactions` as its request URL. Firewatch requests `chat:write`, `im:write`, `users:read`, and `users:read.email`; synchronization looks up only existing, verified Firewatch member emails and stores their immutable Slack user IDs. The worker opens a bot DM and sends each incident escalation to the mapped on-call responder. Responders can acknowledge or resolve from the DM. Firewatch verifies Slack’s timestamped HMAC signature, updates every delivered alert when status changes, and posts the lifecycle event in the message thread. It never guesses from a display name and does not import the Slack directory. Unmatched members continue to use their other verified notification methods.
### Alert templates
[Section titled “Alert templates”](#alert-templates)
Organization administrators can customize email subject/body, SMS, voice, Slack, and webhook wording under **Organization → Alert templates**. Templates use a small allow-listed `{{variable.name}}` syntax; they cannot execute code, read configuration or secrets, or emit raw HTML. Firewatch validates every template before saving and keeps built-in defaults available through **Restore defaults**. The settings are stored in PostgreSQL, so every API and worker replica sees the same version without custom container images.
### Signed outbound webhooks
[Section titled “Signed outbound webhooks”](#signed-outbound-webhooks)
Configure outbound alert endpoints under **Organization → Integrations**. Firewatch creates a separate HMAC-SHA256 signing secret for each endpoint, encrypts it at rest with the shared Data Protection key ring, and displays it only on creation or rotation. Receivers validate the timestamp, delivery ID, raw body, and `X-Firewatch-Signature`; see [Outbound alert webhooks](/docs/integrations/outbound-webhooks/) for the exact payload and a verification example.
Outbound destinations require HTTPS and public network addresses by default. For a reviewed internal self-hosted receiver, set `FIREWATCH_WEBHOOK_ALLOW_PRIVATE_NETWORKS=true`; set `FIREWATCH_WEBHOOK_ALLOW_HTTP=true` only when plain HTTP is also explicitly required. These values must be the same on the API (configuration validation) and worker (delivery), which the Compose file handles.
Set `FIREWATCH_PUBLIC_BASE_URL` to the external origin through which users reach the web application, including `https://` in a real deployment and no trailing path. Firewatch uses this value for verification and invitation links and infers the passkey Relying Party ID from its host. Normally leave `FIREWATCH_PASSKEY_SERVER_DOMAIN` empty. Set it only when deliberately sharing the RP ID across trusted subdomains, and never serve untrusted content below that domain. Do not set it to the API container name or rely on an arbitrary incoming `Host` header.
WebAuthn requires HTTPS outside the browser’s special `localhost` development case. Production must keep the public URL, TLS origin, reverse-proxy host, and RP ID stable; changing the RP ID makes previously registered passkeys unusable.
Users can register up to ten named passkeys from **Settings → Authentication**. Firewatch prevents a passkey-only account from deleting its final credential; operators should encourage at least two credentials on separate devices. Email account recovery is not enabled yet. The intended self-hosted policy is a short-lived verified-email link that permits setting a password, followed by sign-in and passkey re-enrollment. Managed deployments may disable that route in favor of administrator or enterprise identity-provider recovery.
Account passwords, invitation tokens, recovery tokens, external-provider client secrets, and Data Protection key-protection material are secrets. Do not place them in `.env` under source control. Firewatch stores its shared Data Protection key ring in PostgreSQL so logins survive API restarts and multiple replicas; a production operator is responsible for database encryption, access control, backups, and any configured certificate or key used to protect that key ring.
The web image sends same-origin API calls through its Nginx proxy. Compose sets `FIREWATCH_API_UPSTREAM=http://api:8080`; operators can change that internal URL without rebuilding the static assets.
### Runbook image storage
[Section titled “Runbook image storage”](#runbook-image-storage)
Images uploaded to Firewatch-authored runbooks are stored by the API through a provider-neutral asset-store interface. The open-source deployment uses durable filesystem storage at `Runbooks__StoragePath`; Compose mounts the `firewatch-runbooks` named volume at `/app/data/runbooks`. This keeps local and self-hosted installations lightweight—MinIO or another object-storage service is not required—and survives ordinary container replacement.
Compose runs a short-lived `runbook-storage-init` container before the API to give the API’s non-root account ownership of this volume. The official image uses UID/GID `1654`; `FIREWATCH_CONTAINER_UID` and `FIREWATCH_CONTAINER_GID` exist only for operators who deliberately build a custom image with a different runtime account.
Do not point this setting at a temporary container directory. A non-Compose installation should mount a persistent filesystem at the configured path. The public `Firewatch.Storage.S3` adapter lets an operator select S3 without changing the runbook API, editor, or database model. Self-hosting keeps the filesystem provider by default so the one-command install does not require an AWS account.
## Data and backups
[Section titled “Data and backups”](#data-and-backups)
PostgreSQL data lives in the `firewatch-postgres` named volume. Runbook image objects live in the separate `firewatch-runbooks` named volume. Routine container replacement does not remove either volume. Treat the database as the authoritative metadata state, including outbox and idempotency records, and back up the image volume alongside it so Markdown image references remain usable.
Create a logical backup to the host (the output file may contain sensitive data):
```bash
docker compose exec -T postgres \
pg_dump --format=custom --username=firewatch --dbname=firewatch \
> firewatch-backup.dump
```
Production operators should additionally automate encrypted backups, point-in-time recovery, retention, off-host copies, and restore drills. A backup that has not been restored successfully is not a recovery plan.
Never run `docker compose down --volumes` unless permanent deletion of the local database and uploaded runbook images is intended.
## Migrations and upgrades
[Section titled “Migrations and upgrades”](#migrations-and-upgrades)
For a controlled deployment, run the target API image’s migrations exactly once before starting new API or worker replicas:
```bash
docker compose run --rm api --migrate
docker compose up --detach --no-build
```
For a source checkout, rebuild with `docker compose up --build`. For eventual published releases, set `FIREWATCH_VERSION` to a versioned image tag and pin an image digest in higher-assurance environments.
Before upgrading:
1. read release notes and minimum PostgreSQL requirements;
2. take and verify a backup;
3. run the one-off migration command;
4. start API and worker at matching versions;
5. wait for readiness and exercise the test-job smoke flow;
6. review worker retries and structured logs.
Rolling an image back does not reverse its database migration. Follow the expand/contract compatibility policy in [Firewatch architecture](/docs/concepts/architecture/).
## Network and process boundaries
[Section titled “Network and process boundaries”](#network-and-process-boundaries)
Only the web/API ports need ingress. PostgreSQL is published on loopback by the development Compose file for local tools; remove that host-port mapping on a remote server. The worker has no public port.
Before any real deployment exists, place web/API traffic behind a TLS-terminating reverse proxy, restrict database access, use a secrets manager, and configure resource limits and log retention. Forward request IDs and only trust forwarded headers from the known proxy.
Apply per-client rate limiting at that trusted ingress for authentication routes. Firewatch enforces per-account/invitation throttles and password lockout itself, but it does not trust unconfigured forwarded IP headers or impose one shared proxy-IP bucket that could lock out every user behind the web container.
API and worker can be scaled independently after migrations are separated from startup. Graceful shutdown windows should be at least 30 seconds for API and 45 seconds for worker. At-least-once delivery means a worker termination can cause a message retry; idempotent handlers make that safe.
## Health and troubleshooting
[Section titled “Health and troubleshooting”](#health-and-troubleshooting)
* `/health/live` answers whether the API process is alive.
* `/health/ready` includes PostgreSQL and gates dependent Compose services.
* `/web-health/ready` checks only the web container.
Inspect structured logs and service state:
```bash
docker compose ps
docker compose logs --follow --tail=200 api worker postgres
```
If the API is live but not ready, start with PostgreSQL health and the connection string. If jobs stay pending, verify the worker is healthy, uses the same database and provider, and can claim outbox work. Stable message IDs in worker logs are the primary correlation key for retries.
## What belongs elsewhere
[Section titled “What belongs elsewhere”](#what-belongs-elsewhere)
Self-hosting runs the complete public Firewatch product. Managed-service infrastructure and operations do not change the portable product behavior described in these docs.
There is no Helm chart yet; the [Helm packaging status](https://github.com/onfirewatch/firewatch/blob/main/deploy/helm/README.md) records that deliberate placeholder.