Identities, addresses and domains
Binding for implementation. This page defines the lifecycles of identities and addresses, username
validation, the domain kinds and their onboarding, the DomainMonitor health state machine and domain
fallback. The connection methods that keep a domain’s DNS outside Cloudflare (dns_records,
send_only, smtp_relay), and the nameservers and delegated_subdomain methods, are specified in
Domains on any DNS host; this page links to it where they differ.
| Requirements | FR-IDN-1 … FR-IDN-4, FR-ADR-1 … FR-ADR-7, FR-DOM-1 … FR-DOM-6 (FR-DOM-7 … FR-DOM-12 in Domains on any DNS host), FR-TEN-3 |
| Edge cases | A1, A3–A5, A11–A14, C3, D2, G7, H1–H7; N1–N30 in Domains on any DNS host |
| Code | crates/worker/src/handlers/{identities.rs, addresses.rs, domains.rs}, db/{identities.rs, addresses.rs, domains.rs}, crons/retire.rs, domains/{mod.rs, cloudflare_api.rs, ses_api.rs, records.rs, monitor.rs, fallback.rs}; crates/core/src/{address.rs, dns.rs, domain_fsm.rs} |
| ADR | 0003 Addressing with catch-all and a directory, 0005 State machines |
Cloudflare API paths and behaviour on this page were read on 2026-10-09 from the Cloudflare API reference and the Email Service documentation; SES paths from the SES v2 API reference the same day. Items the documentation does not confirm are marked “verify at build time” with the spike that settles them.
Identities
Lifecycle
| State | Event | Guard | Action | Next |
|---|---|---|---|---|
| – | POST …/identities | Valid body; username free; client_id new | Insert identity and addresses (D1 batch); MailboxRequest::Init; identity.created | active |
| – | POST …/identities with a known client_id | Same client_fingerprint | Return 200 with the existing identity; call Init again (idempotent, repairs a lost first call) | unchanged |
| – | POST …/identities with a known client_id | Different fingerprint | 409 client_id_conflict | – |
active | PATCH status: paused | identities:write | pause_reason = 'manual'; identity.paused | paused |
active | Abuse threshold (Outbound) | – | pause_reason = 'abuse_threshold'; identity.paused with metrics | paused |
active | Tenant suspended | – | pause_reason = 'tenant_suspended'; identity.paused | paused |
paused | PATCH status: active | reason manual: identities:write; reason abuse_threshold: platform, partner or tenant key, audit-logged, and only a platform key on a tenant a partner’s key created (J17); reason tenant_suspended: refused (409 identity_paused) | pause_reason = NULL; identity.resumed | active |
paused (tenant_suspended) | Tenant resumed | – | identity.resumed | active |
active, paused | DELETE | identities:write and erasure:manage | Tombstone and remove every address; create the identity-scope erasure (Privacy) | deleting |
deleting | Erasure completed with no holds left | – | identity.deleted, once, from the erasure job’s outbox with identity_id set (Privacy § 6.5) | deleted |
A paused identity still receives and stores mail (FR-IDN-3). Every send needs an accountable human
(owner_name and owner_email, FR-IDN-2); the check is in the send policy pipeline.
The system identity
The deployment’s own mail (console sign-in links and codes, invitations, and the other mail the designs
send from PM_SYSTEM_FROM) goes out through one reserved identity, the system identity, through the
normal outbound pipeline.
- Created by setup (CLI and setup §6.3, step 22) on the default tenant, at the
address of
PM_SYSTEM_FROM(defaultPylota Mail <no-reply@{PM_PLATFORM_DOMAIN}>): anidentitiesrow withis_system = 1,username= the address’s local part,display_name= its display name,owner_name = 'Operator'andowner_email= setup’s--owner-email(elsepostmaster@{PM_PLATFORM_DOMAIN}), andsend_policy.daily_cap= 50,000; plus oneactiveprimary address on the platform domain. Setup writes both rows withmailbox_do_id = '', and the every-minute cron mints the mailbox and sendsMailboxRequest::Init, as the monitor hook does for domains. At most one row hasis_system = 1(a partial unique index). - Exempt from username validation. Its local part may be a reserved name (
no-replyis), because only setup writes its addresses, at creation and whenPM_SYSTEM_FROMchanges, and setup writes them through its internal path (the D1 query API), never the public routes, so the reserved-name and role-name steps of Username validation never run for it. Setup still checks that the local part is ASCII, at most 64 octets, with a valid address syntax. - Not listed to tenants. List endpoints (
GET /v1/identities,GET /v1/tenants/{t}/identities,mail_list_identities, the console) never return it, and a tenant or identity key that names it gets404 identity_not_found. Only a platform key reads or changes it, by ID. It is not counted againstinboxes. - Mail sent to it is stored in its mailbox like any identity’s (bounces and replies to sign-in mail), readable only with a platform key.
- Exempt from the tenant daily cap and from abuse auto-pause. Its sends are not counted against the
default tenant’s
tenant_daily_send_cap; its ownsend_policy.daily_cap(50,000) still applies (Outbound › Policy pipeline, step 18). The delivery consumer records its outcomes but never pauses it (Outbound › Abuse auto-pause): one person’s bounce must not stop everyone’s sign-in mail. When a notification send through it is still refused (429 daily_cap_reached,409 identity_pausedafter a manual pause, or409 domain_not_ready), the Notifier keeps the item, retries it and raisessystem_mail_blocked(Notifications § 7). - Fixed retention. Its mailbox keeps messages 30 days and raw MIME 7 days, whatever the default
tenant’s
retentionpolicy says: the tenant retention job uses these cutoffs for the identity withis_system = 1(Privacy §5.2). Deleting a person also erases the system mail sent to them (Privacy §6.9). - Changing
PM_SYSTEM_FROM. A re-run of setup inserts the new address as anactivealias on the platform domain through its internal path, in one D1 query API batch as at creation (no reserved-name or role-name check, sonoreplyis accepted), then promotes it through the API with the bootstrap platform key; the old address retires as usual. This is the one exception to the rule that every identity keeps exactly one platform-domain address for life (Format): only setup’s internal write creates a second platform-domain address, for the system identity only, while the publicPOST …/addressesnever does (and would refuse a reserved name with400 address_reserved). Promoting it makes the previous primaryretiringwith the usual grace, as on any other domain, instead of anactivealias. The promoted address is the system identity’s platform address from then on. Every other identity’s platform address still cannot be retired or deleted. - Without the console (
PM_CONSOLE=off) it still exists and still sends invitations, becausePM_SYSTEM_FROMandPM_CONSOLE_HOSTare top-level settings, not console settings (Rust workspace §6.1).
Create
POST /v1/tenants/{tenant_id}/identities:
- Validate the body:
username(Username validation),display_name(1–78 characters, Unicode allowed, no CR or LF),owner.email(RFC 5321),signature.html(sanitised with the inbound policy before storage),metadata(≤ 16 keys, ≤ 512 bytes per value). For any key but a platform key,send_policy.daily_capmay not exceed the tenant’s effectiveidentity_daily_send_cap(403 scope_denied,details.field = "send_policy.daily_cap"); the same check runs onPATCH /v1/identities/{identity_id}(Configuration › Who may change a field). client_id(FR-IDN-1):client_fingerprint = hex(sha256(canonical_json(body)))(canonical JSON as in Outbound).SELECT id, client_fingerprint FROM identities WHERE tenant_id = ?1 AND client_id = ?2decides between create, replay and409 client_id_conflict.- Domain. Without
domain_idthe primary address is on the platform domain. Withdomain_idit must name a domain of this tenant in statehealthyordegraded, else409 domain_not_ready. - Addresses. The identity always gets its platform address
{username}{tenant.address_suffix}@{PM_PLATFORM_DOMAIN}(FR-DOM-1),role = 'primary'withoutdomain_id, elserole = 'alias'. Withdomain_idit also gets{username}@{domain}as the primary,activewhen the domain can route it (Routing an address), elsepending. Each address is checked againstaddresses_address(409 address_taken) and againstaddress_tombstones(Tombstones). mailbox_do_id = objects.new_object_id(Mailbox).- One D1
batch:INSERT INTO identities …,INSERT INTO addresses …(one or two rows),INSERT INTO audit_log …(identity.create). MailboxRequest::Init { tenant_id, identity_id, created_at }, which writes the owner intometaand emitsidentity.createdonce. If it fails, the request returns503 unavailable; the client retries with the sameclient_id, and the replay path callsInitagain.201with the Identity object.
Updates, pause and tenant suspension
PATCH /v1/identities/{id} updates D1, then sends MailboxRequest::EmitEvent with identity.updated
(changed = field names), identity.paused or identity.resumed. Suspending a tenant
(PATCH /v1/tenants/{id} with status: suspended) sets tenants.suspended_at and suspended_by
(platform or partner, from the calling key’s level; a partner key cannot resume a tenant whose
suspended_by is platform, 403 scope_denied) and, in the same D1 batch, pauses every active identity with pause_reason = 'tenant_suspended'; resuming reverses only
those. Each identity then gets its event.
Delete (FR-IDN-4, A13)
One D1 batch:
INSERT OR IGNORE INTO address_tombstones (address_hash, identity_id, reason)
SELECT ?2 /* computed per row in Rust: HMAC-SHA256(PM_HASH_KEY, address) */, identity_id, 'deleted'
FROM addresses WHERE identity_id = ?1; -- executed once per address with its hash
DELETE FROM addresses WHERE identity_id = ?1;
UPDATE identities SET status = 'deleting', updated_at = ?3 WHERE id = ?1;
INSERT INTO jobs …; INSERT INTO erasure_requests …; -- scope identity, see Privacy and erasure
Mail to any of its addresses is refused with 550 5.1.1 from the moment the directory cache expires
(≤ 60 seconds), including replies to in-flight threads (A13). The address rows are
gone by the time the erasure job runs, so the batch stores the deleted rows’ (zone_id, routing_rule_id)
pairs, and the ses_bounce_rule of each retired address on an SES domain, in the job’s params_json. The
job deletes those literal routing rules, removes those addresses from their pm-retired-{n} rules (so an
erased address is dropped like an unknown one instead of bouncing with 5.1.6), wipes the mailbox and sets
deleted, which frees the username (identities_username excludes deleted rows)
(Privacy › Identity scope).
Username validation
core::address::validate_username(input, domain_class, max_len, tenant_suffix, existing_usernames) -> Result<String, AddressError>
runs these steps in order (domain_class selects the reserved set below: Platform for a username of
the default tenant, Tenant for any other username and for a local part on a tenant domain; max_len is
24 for a username and 40 for an alias local part on a tenant domain, POST …/addresses) (A1, A3, A4,
A12, FR-ADR-6, FR-ADR-7):
| # | Step | Error |
|---|---|---|
| 1 | fold(input) (below) equals the fold of a name reserved on this domain class | 400 address_reserved |
| 2 | Input contains a non-ASCII character and is mixed-script (its resolved script set is empty), or contains a strong right-to-left character | 400 address_reserved |
| 3 | Input contains any other non-ASCII character (SMTPUTF8 local parts cannot be routed by Email Routing) | 400 address_unsupported |
| 4 | Lower-case (ASCII). Must match ^[a-z0-9][a-z0-9._-]{0,N}$ with N = max_len − 1 ({0,23} for a username, {0,39} for an alias local part), must not contain .., must not end in . | 400 address_invalid |
| 5 | Exact match of a name or pattern reserved on this domain class | 400 address_reserved |
| 6 | len(username) + len(tenant_suffix) > 40 (room for a thread token in 64 octets) | 400 local_part_too_long |
| 7 | Another non-deleted identity in the tenant has the same username, or a different username with the same fold | 409 username_taken |
Display names are Unicode and never refused for script reasons (FR-ADR-7).
Request validation of username and local_part runs through this function, never through a schema
pattern: openapi.yaml leaves both request fields unconstrained and documents the stored form, so a
non-ASCII or confusable input gets address_unsupported or address_reserved from steps 1–3, not a
generic 400 invalid_request. Every route that writes a username or an address runs it, for every
identity. The only writes that skip it are setup’s writes of the system identity’s username and
addresses, through the D1 query API rather than a route (The system identity),
which check ASCII, length and address syntax only.
Reserved names (FR-ADR-6):
| Set | Names | Platform domain | Tenant domain |
|---|---|---|---|
| RFC 2142 role names | info, marketing, sales, support, noc, security, hostmaster, usenet, news, webmaster, www, uucp, ftp | reserved | allowed |
| RFC 2142 operational names that must reach a person | postmaster, abuse | reserved | reserved |
| Service and mail-system names | noreply, no-reply, donotreply, do-not-reply, mailer-daemon, mailerdaemon, mail-daemon, bounce, bounces, root, admin, administrator, sysadmin, system, daemon, nobody, null, devnull, listserv, majordomo, unsubscribe, dmarc, journal (the hidden journal address of Outbound) | reserved | reserved |
| Patterns | a prefix owner-, a suffix -request (RFC 3834 responders skip these), a prefix pm- (reserved for the service) | reserved | reserved |
The shared platform domain is one namespace for every tenant, so a role name there would let one
tenant’s agent receive mail meant for the operator; on a tenant’s own domain the tenant decides who
answers support@ or sales@. The platform set applies to local parts that stand alone on the
platform domain: the usernames of the default tenant, whose suffix is empty, so its platform addresses
are {username}@{PM_PLATFORM_DOMAIN}. Every other tenant’s platform addresses carry its suffix
(support.acme@agents.example is not the role address support@), so its usernames, and aliases on
tenant domains (POST …/addresses), are checked against the tenant set.
Role mail on a tenant domain. Mail to postmaster@ or abuse@ a tenant domain that reaches the
Worker (a catch-all apex) goes to the tenant’s owner contact: the email of the member with role
owner. forward() only reaches verified Email Routing destinations, so the email() handler instead
sends the owner a new message from postmaster@{PM_PLATFORM_DOMAIN} through Email Sending, with the
original attached as message/rfc822 (its headers only when it is over 4 MiB), and accepts the original.
Nothing is stored in a mailbox. A tenant without an owner falls back to PM_SECURITY_CONTACT, else
550 5.1.1 (Inbound › Steps). Mail for PM_SECURITY_CONTACT (an email address,
bare or mailto:) is sent the same way, as a new message, and never with forward(): forward()
reaches only verified Email Routing destination addresses
(email handler,
read 2026-10-09), and setup registers none.
Confusable detection
fold(s) implements the UTS #39 skeleton (version 18.0.0, read 2026-10-09) closely enough to compare
identifiers with ASCII reserved names, and is reused for look-alike domains and display names
(D2):
- Apply Unicode full case folding (for ASCII input: lower-casing).
internalSkeleton, per UTS #39: (a) NFD; (b) remove everyDefault_Ignorable_Code_Point; (c) replace each character by its prototype fromconfusables.txt; (d) NFD again.- Lower-case the result and repeat step 2 until it no longer changes (at most 3 rounds), because prototypes can be upper case (for example a digit zero maps to a capital O).
The bidirectional wrapper of the full skeleton is omitted: step 2 of validation already refuses any
input with a strong right-to-left character. Two strings are confusable when their folds are equal; the
mapping handles multi-character prototypes (for example m maps to rn, so rnailer-daemon and
mailer-daemon fold to the same string).
Data. cargo xtask gen-unicode generates Rust tables from the pinned files confusables.txt
(UTS #39 18.0.0), DerivedCoreProperties.txt (Default_Ignorable_Code_Point) and
ScriptExtensions.txt/Scripts.txt, checked into crates/core/data/. To bound the bundle, the
confusable table keeps only entries whose prototype is entirely ASCII; inputs that would need other
entries are non-ASCII and are refused at step 3 anyway.
Mixed script. The resolved script set is the intersection over all characters of their augmented
Script_Extensions sets, with Common and Inherited counting as all scripts (UTS #39 §5.1). An empty
set means mixed-script.
Addresses
Format
| Domain | Address | Notes |
|---|---|---|
| Platform | {name}{tenant.address_suffix}@{PM_PLATFORM_DOMAIN}, for example bookings.acme@agents.example | {name} is validated as a username. Only the default tenant (made by setup) has an empty suffix |
Tenant zone, delegated or external | {local_part}@{domain}, for example bookings@mail.acmecarhire.example | local_part validated as a username for a tenant domain, with a maximum of 40 characters |
Addresses are stored lower case with an A-label domain; dots are significant (A1).
At most 20 addresses per identity in any state, and one pending address per identity and domain.
Every identity keeps exactly one platform-domain address for its whole life (the system identity, whose
address setup may change, is the one exception: The system identity). It is the
fallback address (Fallback behaviour, FR-DOM-6), so it is never retired automatically,
cannot be retired or deleted through the API (409 address_in_use), and on promotion away from it
becomes an active alias rather than retiring.
Lifecycle
| State | Event | Guard | Action | Next |
|---|---|---|---|---|
| – | POST …/addresses | Domain of this tenant, not removing/removed; ≤ 20 addresses; address free and not tombstoned | Delete an older pending address of this identity on the same domain (and its literal rule) (A11); insert role = 'alias'; route it; identity.address_added | active if routable now, else pending |
pending | Domain reaches healthy/degraded and the address is routed | – | identity.address_activated | active |
pending | Literal rule creation fails (H6) | – | Stays pending; retried by the domain’s monitor (1, 5, 15, 60 minutes, then hourly); issue routing_rule_failed on the domain’s health | pending |
pending | DELETE …/addresses/{id} | Never received mail | Delete the row and its rule | – |
active alias | POST …/promote | Domain healthy or degraded (else 409 domain_not_ready) | In one batch: this address becomes primary; the previous primary becomes an alias, retiring with retire_at = now + retire_previous_after_days (default 90, range 0–365; 0 means retired now), except the platform address, which becomes an active alias (the system identity’s previous platform address retires instead, The system identity). When the promoted address is itself the platform address, this is a rollback as in the next row: the current primary becomes an active alias, not retiring (A14); identity.address_promoted with previous_primary | active primary |
retiring alias | POST …/promote (rollback, FR-ADR-4) | Domain healthy or degraded | This address becomes primary, retire_at = NULL; the current primary becomes an active alias (the change is undone, not mirrored); identity.address_promoted | active primary |
active alias | POST …/retire | Not the primary (409 address_is_primary); not the platform address (409 address_in_use) | after_days > 0: retiring, retire_at = now + after_days; 0: retired now with identity.address_retired | retiring / retired |
retiring | POST …/retire | – | retire_at updated (0 retires now) | retiring / retired |
retiring | retire_at reached (Retirement) | – | retired_at = now; identity.address_retired | retired |
retired | – | – | Terminal. The row is kept forever so the address is never reassigned; inbound gets 550 5.1.6 (FR-ADR-3) | – |
| any | Identity deleted | – | Tombstoned and row removed | – |
A retiring address keeps receiving mail into the same identity, and replies on threads where the counterparty wrote to it are sent from it until it retires (C3, Threading). New threads send from the primary (FR-ADR-2).
Routing an address
Domain routing_mode | Routable when |
|---|---|
catch_all (platform domain, zone apex, delegated, and inbound = ses) | Immediately: the catch-all (or, for inbound = ses, the receipt rule pm-deliver, which has no recipient condition) sends every address to the Worker |
literal (zone subdomain) | A literal rule exists. Created synchronously in the request: POST /zones/{zone_id}/email/routing/rules with { "matchers": [{ "type": "literal", "field": "to", "value": "{address}" }], "actions": [{ "type": "worker", "value": ["pylota-mail"] }], "enabled": true, "name": "pylota-mail {adr_id}", "priority": 0 }; the returned rule ID is stored in routing_rule_id. A matcher value is at most 90 characters, so a longer address is refused with 400 address_invalid. At most 200 rules per domain; the 201st address is refused with 409 domain_in_use and details.reason = "routing_rule_limit". That the worker action’s value is the Worker’s script name is not stated in the reference; verify at build time (S9) |
forward (inbound = forward) | Immediately (mail arrives at the identity’s platform address, forwarded by the domain’s own mail system) |
Retirement
The every-minute cron (crons/retire.rs):
SELECT id, identity_id, tenant_id FROM addresses
WHERE status = 'retiring' AND retire_at <= ?1 ORDER BY retire_at LIMIT 100;
UPDATE addresses SET status = 'retired', retired_at = ?1, updated_at = ?1
WHERE id = ?2 AND status = 'retiring' AND retire_at <= ?1; -- per row; skip if changes = 0
For each changed row it sends EmitEvent identity.address_retired to the identity’s mailbox. The
directory cache in email() expires within 60 seconds, after which inbound gets 550 5.1.6. Literal
routing rules of retired addresses are kept, so that the Worker (not Cloudflare) answers and can return
5.1.6; when a domain reaches 190 rules, the oldest retired addresses’ rules are deleted (those
addresses then get Cloudflare’s own unknown-recipient rejection).
Tombstones (A5)
address_tombstones holds hex(HMAC-SHA256(PM_HASH_KEY, address)), never the clear address. Creating an
address whose hash is tombstoned is allowed only when identity_id equals the tombstone’s identity_id
and that identity is active or paused; otherwise 409 address_taken, across all tenants
(A5, FR-ADR-5). Tombstones are written when an identity is deleted or erased, so in
practice a tombstoned address is never reused.
Domains
Kinds
A domain’s connection method (method), chosen when it is added, fixes its kind, inbound and
transport (FR-DOM-7, Domains on any DNS host §2).
The kinds are platform, zone, delegated and external:
| Kind | method | is_apex | routing_mode | inbound | transport | reply_token | Inbound | Outbound |
|---|---|---|---|---|---|---|---|---|
platform (one per deployment, tenant_id = NULL) | platform (written by setup) | 1 (required) | catch_all | routing | cloudflare | subaddress | Catch-all to the Worker | Email Sending |
zone, apex | cloudflare_zone, nameservers | 1 | catch_all | routing | cloudflare | subaddress | Catch-all to the Worker | Email Sending |
zone, subdomain | cloudflare_zone | 0 | literal | routing | cloudflare | subaddress | One literal rule per address (≤ 200) | Email Sending |
delegated | delegated_subdomain | 1 (the apex of its child zone) | catch_all | routing | cloudflare | subaddress | Catch-all to the Worker on the child zone | Email Sending |
external | dns_records | 0 or 1 | catch_all | ses | ses | subaddress | SES receipt rule pm-deliver → S3 → SNS push and SQS backstop → Worker | SES with Easy DKIM and a custom MAIL FROM |
external | send_only | 0 or 1 | forward | forward | ses | none | The domain’s own mail system forwards to the identity’s platform address | SES with Easy DKIM and a custom MAIL FROM |
external | smtp_relay | 0 or 1 | forward; catch_all with inbound: ses | forward or ses | smtp | none; subaddress with inbound: ses | As send_only, or as dns_records | The customer’s own SMTP relay, after a passing alignment probe |
reply_token = 'subaddress' on inbound = ses depends on spike S11 showing that user+tag@ reaches the
Worker; if it does not, those domains use none. The dns_records, send_only, smtp_relay,
nameservers and delegated_subdomain rows are specified in
Domains on any DNS host; this page keeps the shared steps.
The platform domain must be a zone apex because catch-all rules exist only on the apex. A zone holds at most 30 mail domains (routing and sending together, apex included) (limits, read 2026-10-09).
Adding a domain
POST /v1/tenants/{tenant_id}/domains with domains:write. Common checks: the name is a valid DNS name,
lower case, A-label; 409 domain_exists if a row with that name exists and is not removed (a removed
row is reused: same ID, new tenant, state pending). Every onboarding step is idempotent: it reads
first (for example lists rules or sending subdomains by name) and creates only what is missing, so a
failed request can be repeated. A failed provider call returns 502 upstream_error with
details.step, and no D1 row is written until every step has succeeded.
The request names a method. When it is absent, the old kind is mapped (zone → cloudflare_zone,
external → send_only), and kind: zone with "create_zone": true is the old spelling of
nameservers. Each method has its own onboarding:
method | Onboarding | The deployment needs (refusal without it) |
|---|---|---|
cloudflare_zone | Kind zone | PM_CF_API_TOKEN (422 cf_token_required) |
nameservers | Creating a zone, then Kind zone at the apex | PM_CF_API_TOKEN that can create zones (422 cf_token_required); for a tenant or partner key, the policy domains.allow_create_zone: true (422 transport_unavailable, details.reason = "zone_creation_not_allowed") |
delegated_subdomain | Domains on any DNS host §3.3 | PM_CF_API_TOKEN (422 cf_token_required) and PM_CF_SUBDOMAIN_SETUP=on (422 transport_unavailable, subdomain_setup_disabled) |
dns_records | §4.3 | SES with receiving (422 transport_unavailable, ses_not_configured or ses_receiving_not_configured) |
send_only | Kind external and §4.4 | SES (422 transport_unavailable, ses_not_configured) |
smtp_relay | §5 | Relay credentials that pass a one-off connection (400 smtp_port_not_allowed, 422 smtp_tls_required, 422 smtp_auth_failed); with inbound: ses, SES with receiving |
A method that needs an SES identity (dns_records, send_only, smtp_relay with inbound: ses) is
refused with 422 transport_unavailable, details.reason = "ses_identity_limit", once the region holds
10,000 identities.
The Cloudflare token. PM_CF_API_TOKEN must be set on the Worker for cloudflare_zone,
nameservers and delegated_subdomain; without it, creating such a domain returns
422 cf_token_required. dns_records, send_only and smtp_relay make no Cloudflare call. The one
exception is an apex cloudflare_zone domain (catch-all, no literal rules):
pmail domains add --local-token can onboard it with the operator’s local CLOUDFLARE_API_TOKEN and
insert the row itself
(CLI and setup §18.1), and the Worker’s cron completes
the row as for the platform domain. A subdomain, nameservers and
delegated_subdomain cannot be added that way, because the Worker has to keep calling Cloudflare over the
domain’s life: literal rules per address, onboarding once a new zone is active, delegation checks.
Creating an address that needs a literal rule without the token also returns 422 cf_token_required.
Records (FR-DOM-3) come from the provider at request time for every method: Cloudflare’s routing and
sending DNS endpoints for zones (including nameservers and delegated_subdomain once active), the
name_servers returned when a zone is created, and SES GetEmailIdentity (DKIM tokens,
SigningHostedZone, MAIL FROM status) for dns_records, send_only and smtp_relay with
inbound: ses. The Worker composes only its own values: the ownership TXT, the pm-bounce MAIL FROM
records and the SES endpoint hosts of ses_region
(§4.1). GET …/records
re-reads them; they are never copied from documentation or templates.
Zone permission
cloudflare_zone, nameservers and delegated_subdomain work inside the deployment’s own Cloudflare
account, which holds every tenant’s zones and the zones of the deployment’s own hosts. A tenant or
partner key may therefore use only zones its tenant is entitled to (H8). Platform
keys skip this check. Two sources grant a zone to a tenant:
- Claimed zones.
zone_claims(Data model) records each zone this deployment created for a tenant (nameservers,delegated_subdomain): the row is inserted in the D1 batch that inserts the domain row, andzone_nameis unique, so a zone is claimed by one tenant at most. The claim is deleted by thedelete_zonestep of Domain removal and when the zone expires (zone_expired, Creating a zone step 5), together with the zone. - Listed zones. The tenant’s policy
domains.cloudflare_zones, an array of zone names that only a platform key can write (Configuration › Who may change a field), for zones of the account that an operator assigns to the tenant.
The check runs in two places, both before anything is written, and refuses with 403 scope_denied,
details.reason = "zone_not_allowed":
- By name, before any Cloudflare call. Let the deployment zones be the registrable domains (public
suffix list) of
PM_PLATFORM_DOMAIN,PM_API_HOSTandPM_CONSOLE_HOST. The name is refused when it equals or is under a deployment zone, or under a zone another tenant claimed. Forcloudflare_zone(and with itreplace_mx, which deletes MX records only inside that zone), the name must also equal or be under a zone this tenant claimed or a zone in itsdomains.cloudflare_zones. A listed zone grants names strictly under it: its apex, andreplace_mxthere, stay platform-only, so listing an operator’s zone (for examplepylota.io, to allownotify.pylota.io) never hands over that zone’s own mail.nameserversanddelegated_subdomaincreate their own zone and need no grant, only the two refusals. - On the zone found (Kind
zonestep 1, which takes the most specific zone of the account containing the name). The found zone must be claimed by this tenant (zone_claims.zone_idwith itstenant_id), or listed in itsdomains.cloudflare_zonesby name and claimed by no other tenant. A zone claimed by another tenant is refused even when the policy lists it, so a listed parent zone never reaches a more specific zone another tenant owns.
Both refusals have the same body, whether or not a zone of that name exists in the account, so the
answer does not reveal other tenants’ zones. pmail domains add --local-token runs with the operator’s
own Cloudflare token and inserts the row itself; it is a platform operation and not checked.
Kind zone
The cloudflare_zone method. Needs PM_CF_API_TOKEN (422 cf_token_required without it, as above) and
PM_CF_ACCOUNT_ID, which setup writes.
-
Find the zone. List zones by name for the account, trying the domain and then each parent label up to the registrable domain (
GET /zones?name={name}; verify the query parameters at build time). Not found:404 domain_not_found. Thenameserversmethod creates the zone instead (Creating a zone). For a tenant or partner key, the found zone then passes the second zone permission check, or the request gets403 scope_denied(zone_not_allowed) before anything is changed. -
Existing mail at an apex (H5). Query MX at the apex on both DoH resolvers. If it has MX records other than the hosts Email Routing expects (taken from step 6, never hard-coded) and the request lacks
"replace_mx": true, refuse with409 existing_mxand a fix saying that existing mail would stop. Withreplace_mx, delete those MX records through the DNS records API before enabling routing. -
SPF preflight (H2). If the apex already publishes SPF, count the DNS lookups of the record Email Routing will need merged with the existing one (SPF lookup count). More than 10, or more than 2 void lookups: refuse with
400 spf_lookup_limit,details.lookups, and a fix naming the includes to flatten. -
Ownership record. Generate
ownership_token(16 random bytes, Crockford base32) and create TXT_pylota-mail.{domain}=pm-verify={token}through the DNS records API. -
Receiving (when
receiving):POST /zones/{zone_id}/email/routing/dnswith{ "name": "{domain}" }(“Add and lock the necessary MX and SPF records”). For a subdomain the reference does not confirm this call enables routing on the subdomain; S9 verifies it.PATCH /zones/{zone_id}/email/routingwith{ "support_subaddress": true }, souser+token@matchesuser@and the+tokenstays inmessage.to.- Apex:
PUT /zones/{zone_id}/email/routing/rules/catch_allwith{ "actions": [{ "type": "worker", "value": ["pylota-mail"] }], "matchers": [{ "type": "all" }], "enabled": true, "name": "pylota-mail" }. - Subdomain: literal rules are created per address (Routing an address).
-
Sending (when
sending):POST /zones/{zone_id}/email/sending/subdomainswith{ "name": "{domain}" }(the response holdstag,dkim_selectorandreturn_path_domain; whether an apex can be onboarded through this endpoint is verified by S9), thenPATCH /zones/{zone_id}/email/sending/subdomains/{tag}with{ "drop_suppressed_recipients": false, "preview_enabled": false }(Outbound › G4; Email preview keeps a copy of each sent message for about seven days and is on by default for new sending domains, Privacy). Both fields are in the Cloudflare API reference for this endpoint (read 2026-10-09).SES identity for the failover (optional; only when
sending, the SES transport is configured, and the domain is a tenant domain, never the platform domain). This prepares the Email Sending failover of J5, so thatPATCHtosesworks later without any DNS change: create the SES identity as in step 2 of Kindexternal(CreateEmailIdentity, orGetEmailIdentityonAlreadyExistsException, through the SES token bucket of Domains on any DNS host §4.8), publish its three Easy DKIM CNAMEs{token}._domainkey.{domain}→{token}.{SigningHostedZone}through the DNS records API, and setses_identity= the domain andses_region=PM_SES_REGION. The CNAMEs joinrecords_jsonwithpurpose: "dkim"andrequired: false. No custom MAIL FROM is set up: during a failover SES uses its own MAIL FROM domain, so SPF does not align but Easy DKIM does, and DMARC passes on DKIM. This step never refuses the domain: when the token bucket answersBusy, or an SES call fails, the domain is created without it and its monitor runs the step again in the background (a background caller, at most once an hour); when the region already holds 10,000 identities it is skipped. Until it has run, aPATCHtosesgets422 transport_unavailable. A domain added before SES was configured has no SES identity, and neither has one inserted bypmail domains add --local-token, because the Worker has no Cloudflare token to publish the CNAMEs with. -
Event subscription (when
sending): find the queue ID ofpm-delivery-eventsby listing the account’s queues, thenPOST /accounts/{account_id}/event_subscriptions/subscriptionswith:{ "name": "pylota-mail {domain}", "enabled": true, "source": { "type": "email.sending", "zone_id": "{zone_id}", "domain": "{domain}" }, "destination": { "type": "queues.queue", "queue_id": "{queue_id}" }, "events": ["message.delivered", "message.deferred", "message.bounced", "message.failed", "message.rejected", "message.complained"] }The returned ID is stored in
event_subscription_id. Theemail.sendingsource shape is from Wrangler’s source, not yet the API reference; S9 verifies it.Spike S9 fallback: manual delivery events. When the subscription cannot be created at runtime (the create call answers
401,403,404,405or501: the API or the token cannot do it), the domain is still created. The row is inserted withevent_subscription_id = NULL, and the domain object reportsdelivery_events: "manual"anddetails.action = "run pmail domains subscribe {domain}"in the201response and in every later read, next to its records.429and5xxanswers are not this case: the create request fails with502 upstream_erroras for any other onboarding call, and is safe to retry. Delivery events for the domain start only oncepmail domains subscribe {domain}has run (CLI and setup §18.4): it creates the subscription with the operator’s localCLOUDFLARE_API_TOKENand records its ID. Until then sends work, delivery statuses stay atsubmitted(the transport’s acceptance), uncertain sends are not reconciled, andpmail doctorfailssending.event_subscriptionswith that command as the fix. The same applies to anameserversdomain, whose step 7 runs in the monitor once the zone is active. If S9 also shows that the Worker cannot delete a subscription, domain removal keeps going andpmail doctorlists the left-over subscription with thewrangler queues subscription deletecommand. Tests:it::domains::s9_manual_delivery_events(cloudflare_zone) andit::domains::s9_manual_delivery_events_nameservers.delivery_eventsis derived, not stored:activewhen the domain sends through Cloudflare and has anevent_subscription_id, or sends through SES or SMTP (their events arrive through SNS or DSNs);manualwhen it sends through Cloudflare without one;nonewhensendingis false. -
Read the records back (FR-DOM-3).
GET /zones/{zone_id}/email/routing/dnsandGET /zones/{zone_id}/email/sending/subdomains/{tag}/dns; normalise each to{ type, name, value, priority, purpose, required }withpurposeone ofmx,spf,dkim,return_path,dmarc,ownership,ns; add the ownership TXT; store asrecords_json. These are the records shown to users. They are never copied from documentation or templates. -
Insert the row (
state = 'pending',monitor_do_id), callDomainRequest::Init, which starts verification at once, and emitdomain.created.
Creating a zone
This is the nameservers method (old spelling: kind: zone with "create_zone": true), for a domain
used only for mail (Domains on any DNS host §3.2). Platform keys
may always use it; tenant and partner keys only when the tenant’s policy has domains.allow_create_zone: true (see
Adding a domain).
- Dedicated-domain check (N21). Before creating anything, query both DoH
resolvers for
A,AAAAandMXat the name and forCNAME/Aatwww.{name}. If any exist and the request lacks"confirm_dedicated": true, refuse with409 domain_not_dedicated;details.recordslists what was found, and the fix says the website or mail on the domain would stop. - Create the zone:
POST /zoneswith{ "account": { "id": "{account_id}" }, "name": "{domain}", "type": "full" }. Cloudflare error1105becomes429 upstream_rate_limitedwithRetry-Afteranddetails.retry_afterof 10800 seconds (N22); a zone hold becomes409 zone_hold. - The zone is created in a pending state and the response’s
name_serversare returned inrecordsasNSrecords (purpose: "ns") to set at the registrar.expected_ns_jsonis set toname_servers. The D1 batch that inserts the domain row also inserts itszone_claimsrow (zone ID, zone name, the tenant, the domain), so no other tenant can use the zone throughcloudflare_zone(Zone permission). The first zone permission check ran before step 1. - Steps 2–8 of Kind
zonerun once the zone is active: the monitor polls the zone (GET /zones/{zone_id},status = "active"; verify the field at build time) on each check while the domain ispending, and runs onboarding then, at the apex (catch-all).confirm_dedicatedstands in forreplace_mxat step 2, because the user has already accepted that existing mail stops. - Expiry (N23). Cloudflare deletes a Free-plan zone that is not activated within
28 days. The monitor sends a final
domain.reminderat day 21. If the zone disappears, the domain moves toremovedwithstate_reason = zone_expired, itszone_claimsrow is deleted, anddomain.removedcarriesreason: "zone_expired"; the user can add it again.
Kind external
The send_only method. dns_records and smtp_relay are external too; their onboarding is in
Domains on any DNS host §4.3 and
§5. Needs the SES transport
(PM_SES_REGION and both SES secrets); without it 422 transport_unavailable,
details.reason = "ses_not_configured".
- Ownership record as above; the user publishes it.
- SES identity:
POST /v2/email/identitieswith{ "EmailIdentity": "{domain}", "ConfigurationSetName": "pylota-mail" }(SigV4).AlreadyExistsException→GET /v2/email/identities/{domain}. The response’sDkimAttributes.Tokens(three) andSigningHostedZonegive three CNAME records{token}._domainkey.{domain}→{token}.{SigningHostedZone}(built from the returned zone, which differs by region).ses_identity= the domain. - Custom MAIL FROM
pm-bounce.{domain}, as in step 5 of §4.3. - Records = the three DKIM CNAMEs, the MAIL FROM MX and TXT at
pm-bounce.{domain}, and the ownership TXT. There is no routing record: the user configures their mail system to forward each address to the identity’s platform address, shown per address in the domain response. - Insert the row (
kind = 'external',method = 'send_only',inbound = 'forward',transport = 'ses',routing_mode = 'forward',reply_token = 'none',ses_region,mail_from_domain) and start monitoring, as above.
Each address on such a domain carries forwarding, which stays unverified until a forwarding test
(POST …/addresses/{address_id}/test-forwarding) or a real forwarded message arrives
(§4.4).
For a domain with inbound = forward (send_only, and smtp_relay with inbound: forward), inbound
mail arrives with the platform address as envelope recipient. When a
To/Cc address of the message belongs to the same identity on that domain, the inbound
pipeline treats the message as delivered to that address (delivered_to = the external address,
is_bcc = 0), so replies are sent from it.
The platform domain
pmail setup onboards the platform domain with the operator’s local token (steps 4–8 of kind zone,
apex) and inserts its row (kind = 'platform', method = 'platform', tenant_id = NULL,
state = 'pending') through the D1 query API
(CLI and setup §6.6). Only the Worker can mint a DomainMonitor
ID (IDs are bound to the jurisdiction), so setup writes the row with monitor_do_id = ''.
Monitor hook. The every-minute cron (* * * * *) completes such rows:
SELECT id, kind FROM domains WHERE monitor_do_id = '' LIMIT 20;
-- per row, after minting a DomainMonitor ID in PM_JURISDICTION:
UPDATE domains SET monitor_do_id = ?1, updated_at = ?2 WHERE id = ?3 AND monitor_do_id = '';
Only when the UPDATE changed the row does the cron send DomainRequest::Init, which starts
verification, and emit domain.created for a row with a tenant_id. An overlapping run that lost the
update discards its unused ID. The same hook completes the rows that pmail domains add --local-token
inserts (CLI and setup §18.1). Setup polls the
platform row’s monitor_do_id for up to 2 minutes and reports monitor: started, or a warning naming
the cron.
From then on the platform domain is monitored like any other domain. Without PM_CF_API_TOKEN in the
Worker, GET …/records for it returns the stored records_json (read from the API by setup) checked
against DNS.
SPF lookup count
core::dns::count_spf_lookups(record, resolve) -> SpfCount (RFC 7208 §4.6.4): each include, a, mx,
ptr, exists and the redirect modifier counts one lookup, recursively through include and
redirect targets fetched by DoH (depth ≤ 10, each target fetched once); all, ip4, ip6 and exp
count none. A lookup that returns no records is a void lookup. More than 10 lookups or more than 2 void
lookups is a permerror.
Domain health
DomainMonitor (one Durable Object per domain) runs verification and health checks as an alarm-driven
state machine (FR-DOM-4, FR-DOM-5, ADR 0005). core::domain_fsm holds
the pure transition function; the object does I/O and persistence.
Schedule
- A full check every 15 minutes (
alarm:check), and immediately afterInit, averifyrequest (rate-limited to one a minute), areprove, an address change on the domain, or a sending errorsender_domain_unavailablefrom a transport. - Ownership (NS and RDAP) weekly (
alarm:ownership), and onreprove(H4). - After a check whose agreed outcome would change the state, the confirming check runs after 5 minutes instead of 15. After a resolver error or disagreement, the next check runs after 2 minutes.
What each check verifies
Each check queries every expected record on both DoH resolvers (PM_DOH_RESOLVERS). Expected values
come from records_json (re-read from the provider APIs once a day and on GET …/records). For each
record the result is ok, missing, mismatch or unexpected, and its issue code has a level:
| Record | Applies to | ok when | Issue codes (level) |
|---|---|---|---|
| MX at the domain | receiving with inbound = routing (zone, delegated, platform) | The set of MX hosts equals the routing API’s | mx_missing (fail); mx_unexpected: an extra, non-Cloudflare MX host (degraded) |
| SPF at the domain | receiving with inbound = routing | Exactly one v=spf1 TXT, containing the routing API’s include, ≤ 10 lookups | spf_missing (degraded); spf_multiple (degraded); spf_too_many_lookups (H2, degraded) |
Routing DKIM (cf2024-1._domainkey.{domain}, per the routing API) | receiving with inbound = routing | TXT equals the API’s | routing_dkim_missing (degraded) |
Return path (cf-bounce.{domain} MX and SPF TXT, per the sending API) | sending with transport = cloudflare | Records equal the API’s | return_path_missing (fail) |
Sending DKIM ({dkim_selector}._domainkey.{domain}, from the sending API) | sending with transport = cloudflare | TXT p= equals the API’s | dkim_missing, dkim_mismatch (fail) |
| SES DKIM (three CNAMEs) | transport = ses or inbound = ses; informational on a Cloudflare-transport domain with ses_identity (below) | Each CNAME points to {token}.{SigningHostedZone}, and SES GetEmailIdentity reports DkimAttributes.Status = SUCCESS (once a day) | dkim_missing (fail); ses_dkim_failed (FAILED, fail) |
DMARC (_dmarc.{domain}, else the organisational domain) | sending | Exactly one valid v=DMARC1 record with p=quarantine or p=reject, and alignment possible (H3) | dmarc_missing (degraded); dmarc_policy_none (degraded); dmarc_multiple (degraded); dmarc_alignment_impossible (fail) |
Ownership TXT (_pylota-mail.{domain}) | all except platform | Contains pm-verify={ownership_token} | ownership_record_missing (ownership) |
| NS (weekly) | zone and platform | The NS set equals expected_ns_json | nameservers_changed (ownership) |
| RDAP (weekly) | zone, delegated and external | Fingerprint equals rdap_fingerprint | registration_changed (ownership) |
The failover identity while transport = cloudflare. On a zone or delegated domain that has
ses_identity (the optional step of Kind zone) and still sends through Cloudflare, the
three SES DKIM CNAMEs and the daily GetEmailIdentity check run, but only for information: each record’s
result is shown in GET …/records and GET …/health, and they add no issue to the outcome, so they never
change the domain’s state. After a PATCH to ses, the rows for transport = ses replace those for
transport = cloudflare (return path and sending DKIM): the SES DKIM CNAMEs and the SES identity check
count with their levels; the receiving, DMARC, ownership, NS and RDAP rows are unchanged; and the MAIL FROM
row of Domains on any DNS host § 6 does not apply,
because the failover identity has no custom MAIL FROM (mail_from_domain stays null), so failing over
never makes the domain degraded.
The checks that depend on the connection method (SES inbound MX, SES identity, MAIL FROM, the SES account, the alignment probe, SMTP login, parent delegation and doubled names) are in Domains on any DNS host › Health checks per method. They use the same levels, the same two-resolver agreement and the same state machine.
Alignment (H3). core::dns::check_alignment(dmarc, transport_dkim_domain, return_path_domain):
DKIM aligns when the transport’s DKIM d= equals the domain (adkim=s) or shares its organisational
domain (adkim=r); SPF aligns by the same rule over the return-path domain and aspf. Cloudflare signs
with d= the sending domain and uses cf-bounce.{domain} as return path, so aspf=s alone never
aligns but DKIM does; SES Easy DKIM signs with d= the domain. “Alignment impossible” means neither can
align under the record’s tags while p is quarantine or reject. For transport = smtp the relay signs, so
alignment is proved by the alignment probe instead
(§5.3).
RDAP. The registry is found from the IANA bootstrap file https://data.iana.org/rdap/dns.json
(longest label match, right to left; cached for 24 hours) and queried at {base}domain/{registrable domain}.
rdap_fingerprint = hex(sha256(registrar entity handle ‖ registrant entity handle or "" ‖ registration eventDate))
(RFC 9083 roles registrar and registrant, event registration). Requests follow Security § 9.3 (HTTPS, at most one
redirect to another bootstrap-listed host, 256 KB, 10 s). An RDAP error or timeout is ignored for that
week. A change must be seen on two RDAP queries an hour apart: the first stores meta.rdap_pending
({fingerprint, seen_at}) and sets alarm:ownership to an hour later; the second confirms it (the same
fingerprint) or clears it.
Outcome per resolver and agreement
- Per resolver: if any query failed (timeout, HTTP error,
SERVFAIL,REFUSED) →error. Otherwise: any ownership-level issue →ownership_changed; else any fail-level issue →fail; else any degraded-level issue →degraded; elsepass. - Each resolver’s result is stored in
checks(resolver,results_json,outcome); rows beyond the last 500 are deleted. - Agreement (H7). If either resolver’s outcome is
error, or the two outcomes differ, the cycle has no agreed outcome: nothing changes, and the next check runs in 2 minutes. One resolver’s failure or lie can therefore never change the state. - Two consecutive agreeing cycles.
meta.candidateandmeta.candidate_counttrack the agreed outcome that would change the state. The same outcome again increments the count; a different one resets it; an outcome that matches the current state clears it. The transition happens when the count reaches 2.
State machine
States: pending, verifying, healthy, degraded, failing, suspended, removing, removed.
“Recovered” is not a state: it is the event domain.recovered on a return to healthy.
| State | Event | Guard | Action | Next |
|---|---|---|---|---|
pending | Init | Zone active (or kind external) | First check now | verifying |
pending | Check | Method nameservers or delegated_subdomain, and the zone still pending | Reminders | pending |
pending | Check | The pending zone no longer exists (Cloudflare deletes it after 28 days, N23) | state_reason = zone_expired; domain.removed (reason: zone_expired) | removed |
pending | Check | Zone became active | Run onboarding steps 2–8 | verifying |
verifying | Agreed pass ×2 | – | ownership_verified_at = now; domain.verified; activate the domain’s pending addresses that are routed | healthy |
verifying | Agreed degraded ×2 | – | ownership_verified_at = now; domain.degraded; activate addresses | degraded |
verifying | Agreed fail or ownership_changed | – | Record issues; reminders | verifying |
healthy | Agreed degraded ×2 | – | domain.degraded with issues | degraded |
healthy, degraded | Agreed fail ×2 | – | failing_since = now; domain.failing with issues and fallback_active | failing |
degraded | Agreed pass ×2 | – | domain.recovered (from_state: degraded) | healthy |
failing | Agreed pass ×2 | – | failing_since = NULL; domain.recovered (from_state: failing) | healthy |
failing | Agreed degraded ×2 | – | failing_since = NULL; domain.degraded; sending from the domain resumes | degraded |
failing | now − failing_since ≥ 14 days | – | New ownership_token; domain.suspended (reason: failing_14_days) | suspended |
healthy, degraded, failing | Agreed ownership_changed ×2 | – | New ownership_token; domain.suspended with reason = nameservers_changed, ownership_record_missing or registration_changed | suspended |
suspended | POST …/reprove | – | New ownership_token; records_json updated; check now | suspended |
suspended | Check | The new ownership TXT is seen on both resolvers in two consecutive cycles | expected_ns_json and rdap_fingerprint re-recorded, ownership_verified_at = now | verifying |
any except removing, removed | DELETE …/domains/{id} | No active or retiring address on the domain (else 409 domain_in_use) | Delete its pending addresses; create a domain_remove job | removing |
removing | Job completed | – | domain.removed (reason: requested) | removed |
Every transition updates domains.state, state_reason (the first issue code) and state_changed_at in
D1 and appends the event to the monitor’s outbox in the object’s transaction (the D1 update runs after
commit and is retried by the alarm until it succeeds). Address activation runs on each transition into
healthy or degraded.
Reminders. While a domain is pending, verifying, degraded, failing or suspended,
domain.reminder is emitted at 24 hours, 72 hours and 7 days in that state (hours_in_state), tracked in
meta.reminders_sent_json and reset on every state change. A pending zone created by nameservers
also gets a final reminder at day 21 (Creating a zone, step 5).
Health response. GET /v1/domains/{id}/health returns the state, the reason, since, issues
(code, record, fix; the fix quotes the exact name and value from records_json), the last checks
per resolver, and fallback_active.
Fallback behaviour
fallback_active = (state ∈ {failing, suspended} OR (transport = ses AND SES sending is paused for the account)) AND policy.domain_fallback. The service never sends as a domain whose authentication records are broken (failing) or whose ownership signals changed (suspended) (FR-DOM-5). SES sending is paused when the platform check reportsses_sending_paused(Health checks per method).- Fallback works the same for every connection method. An
smtp_relaydomain whose alignment probe fails twice becomesfailinglike any other (§5.3). The platform address always sends through Email Sending on the platform domain. - While active, sends from the domain go out from the identity’s platform address, with the same
display name, a
Reply-Tocarrying the thread token on the platform address, the flagsent_via_fallback, andfallback_pinned = 1on the thread (Outbound, FR-DOM-6, H1). Withdomain_fallback = falsethey fail withdomain_failing_no_fallback. - Inbound mail to the domain is still accepted while it is
failingorsuspended. - Recovery. On
domain.recovered, new threads send from the domain again. Threads that used fallback stay pinned to the platform address until they have been quiet for 72 hours (Threading), so a conversation does not change From address mid-way. - A
retiringaddress’s domain may fail too; sends fall back the same way (G7). - The platform domain has no fallback.
Changing the transport (J5)
PATCH /v1/domains/{domain_id} with { "transport": "ses" | "cloudflare" } is the Email Sending
failover (J5). Only platform keys may call it (403 scope_denied otherwise). It
needs ses_identity set, SES configured and the SES DKIM records in records_json for ses
(422 transport_unavailable otherwise). A cloudflare_zone, nameservers or delegated_subdomain
domain gets all three during onboarding when SES is configured (the optional step of
Kind zone), so PATCH to ses works for any such domain that has ses_identity. Only the methods that put the domain on Cloudflare
(cloudflare_zone, nameservers, delegated_subdomain) can switch; any other method, and the platform
domain, gets 422 transport_unavailable with details.reason = "method_not_supported". An smtp_relay
domain changes its relay with PATCH and smtp instead (tenant, partner or platform key with domains:write);
the new values are kept pending until a probe passes
(§5.3). A transport change updates domains.transport,
writes an audit_log row (domain.transport), and asks the monitor for a check at once, because DKIM
alignment differs per transport. The outbound consumer reads the transport at transport time, so queued
mail moves with it.
Domain removal
The domain_remove job (JobRunner, steps journaled in steps) undoes onboarding, each step
idempotent:
delete_rules: delete every literal rule of the domain’s addresses (DELETE /zones/{zone_id}/email/routing/rules/{rule_id}), including retired ones.disable_catch_all(apex):PUT …/rules/catch_allwith"enabled": false.disable_routing:DELETE /zones/{zone_id}/email/routing/dnsfor the domain’s name.disable_sending:DELETE /zones/{zone_id}/email/sending/subdomains/{tag}(this also removes its DNS records; routing still active elsewhere is unaffected).delete_subscription: delete the event subscription byevent_subscription_id.delete_ses_identity(whenses_identityis set):DELETE /v2/email/identities/{domain}; on azoneordelegateddomain, also delete the three DKIM CNAMEs that onboarding published through the DNS records API.prune_retired_rules(inbound = ses): remove the domain’s retired addresses from theirpm-retired-{n}rules (read, merge, write, as in Domains on any DNS host § 4.6) and clearaddresses.ses_bounce_rule. With the SES identity gone, SES no longer accepts mail for the domain, so the rule entries only use capacity.delete_ownership_record(zone, delegated): delete the_pylota-mailTXT.delete_zone(nameservers,delegated_subdomain):DELETE /zones/{zone_id}, because this deployment created the zone for a domain used only for mail, then delete itszone_claimsrow. A zone found throughcloudflare_zonebelongs to the account owner and is never deleted.finish:UPDATE domains SET state = 'removed', smtp_sealed = NULL, smtp_pending_sealed = NULL, updated_at = ?and emitdomain.removed(reason: requested).
A provider 404 on a delete counts as done only when it carries the provider’s own “not found” error
code; any other 404 is retried. Failed steps retry with backoff (1, 5, 15, 60 minutes, then hourly).
Retired address rows stay, so their addresses are never reassigned.
Cloudflare API token permissions
Two Cloudflare tokens exist, and their permissions are listed once, in one table: Deploy to Cloudflare › Create a Cloudflare API token. That table names each permission as the dashboard shows it (for example Zone · Edit, which the API tab of Cloudflare’s permissions reference calls Zone Write) and marks which token needs it.
- The operator’s own token,
CLOUDFLARE_API_TOKEN, is used only by the CLI (CLI reference › Commands that use your Cloudflare token):pmail setuponboards the platform domain with it (The platform domain), andpmail domains add --local-tokenan apexcloudflare_zonedomain (CLI and setup §18.1). - The Worker’s token, the secret
PM_CF_API_TOKEN, is what this design calls during a domain’s life. It is needed forcloudflare_zone,nameserversanddelegated_subdomaindomains (Adding a domain); a deployment whose tenants use onlydns_records,send_onlyorsmtp_relaycan leave it unset.
What the Worker does with each permission marked for it in that table:
| Permission (dashboard name) | The Worker uses it for |
|---|---|
| Zone · Read | Finding zones (Kind zone step 1) and reading a new zone’s status (Creating a zone) |
| Zone · Edit | Creating a zone for nameservers and delegated_subdomain, and deleting it on removal (delete_zone). Whether a zone-scoped grant can create new zones is not stated; verify at build time |
| Zone Settings · Edit | Enabling routing, setting sub-addressing and reading the routing DNS records (steps 5 and 8); disable_routing on removal |
| Email Routing Rules · Edit | The catch-all rule on an apex and the literal rules per address on a subdomain (Routing an address) |
| DNS · Edit | The ownership TXT, MX removal for replace_mx (steps 2 and 4), and the SES DKIM CNAMEs of the failover identity (step 6) |
| Email Sending · Edit | Sending onboarding and its DNS records (step 6) and the suppression list (G4). It is named in the Email Service docs but not on the permissions page; its scope is verified at build time |
| Queues · Edit | Listing queues and creating a domain’s event subscription to pm-delivery-events (step 7) |
| Vectorize · Edit, Workers AI · Read and Edit | Only the REST fallbacks, if spike S6 fails |
Neither token needs an Email Routing Addresses permission, because nothing registers a destination address (role mail, under Username validation, is sent as new messages, never forwarded).
Tests
| Test | Covers |
|---|---|
core::address::a1_case_and_dots | Case-insensitive match, dots significant (A1) |
core::address::a3_smtputf8_refused | Non-ASCII local parts → address_unsupported; Unicode display names allowed (A3, FR-ADR-7) |
core::address::a4_reserved_and_confusable | Every reserved name and pattern; rnailer-daemon, p0stmaster, Cyrillic а in аbuse, mixed scripts → address_reserved (A4, FR-ADR-6) |
core::address::a12_local_part_budget | Username plus suffix over 40 → local_part_too_long (A12) |
core::address::skeleton_vectors | Fold vectors from UTS #39 test data for ASCII-prototype entries |
it::identities::client_id_idempotent | Replay returns 200 and the same identity; a changed body returns client_id_conflict (FR-IDN-1) |
it::identities::a5_tombstone_blocks_reuse | A deleted identity’s address cannot be created by any tenant (A5) |
it::identities::a13_delete_then_reply_rejected | After delete, mail to its addresses gets 550 5.1.1 (A13, FR-IDN-4) |
it::addresses::a11_newer_pending_replaces | A second pending address on a domain replaces the first (A11) |
it::addresses::promote_retire_rollback | Promote, retire after grace, rollback by promoting the retiring address, events emitted (FR-ADR-2–4) |
it::addresses::a14_platform_address_kept | Promoting away keeps the platform address active; retiring or deleting it is refused with 409 address_in_use; promoting it again rolls back: it is primary and the custom address is an active alias with retire_at = NULL (A14, FR-ADR-2, FR-DOM-6) |
core::address::a4_role_names_by_domain | support, sales, info, marketing are refused on the platform domain and allowed on a tenant domain; postmaster and abuse are refused on both (A4, FR-ADR-6) |
it::domains::transport_patch | With SES configured, a cloudflare_zone domain is onboarded with ses_identity, ses_region and its three DKIM CNAMEs (created through the Cloudflare API fake, required: false); while it sends through Cloudflare, a missing CNAME or a FAILED SES DKIM status changes no state; a platform key switches it to ses and back; a tenant key gets 403 scope_denied; a domain without an SES identity gives 422 transport_unavailable (J5) |
it::domains::remove_deletes_ses_identity | Removing a cloudflare_zone domain that onboarding gave a failover SES identity runs delete_ses_identity: the SES fake no longer has the identity and the Cloudflare DNS fake no longer has its three DKIM CNAMEs; a provider 404 with its own not-found code counts as done, any other 404 is retried; tenant erasure’s remove_domains does the same for every such domain of the tenant (Domain removal, Privacy §6.6) |
it::addresses::retirement_cron | retire_at reached → retired, identity.address_retired, inbound 550 5.1.6 (FR-ADR-3) |
it::addresses::c3_reply_from_retiring | Replies from the retiring address the counterparty used (C3) |
it::domains::h1_failing_fallback | DNS fake removes DKIM; after two agreeing checks failing; sends fall back with thread continuity; restore → recovered; pinned threads stay (H1, FR-DOM-5, FR-DOM-6) |
core::dns::h2_spf_lookup_count | Lookup and void-lookup counting; preflight refusal (H2) |
core::dns::h3_strict_alignment | adkim=s/aspf=s against Cloudflare and SES signing domains (H3) |
it::domains::h4_ownership_change | NS move, ownership TXT removed, RDAP change → suspended; reprove → verifying (H4) |
it::domains::h5_existing_mx | Apex with existing MX refused without replace_mx (H5) |
it::domains::h8_zone_permission | With the Cloudflare fake holding a zone claimed by tenant B, a zone listed in tenant A’s domains.cloudflare_zones, an unlisted zone, and the zone of PM_PLATFORM_DOMAIN: a tenant key and a partner key of tenant A get 403 scope_denied (zone_not_allowed, the same body for an existing and a missing zone) for cloudflare_zone on B’s zone (also when A’s policy lists it, and for a name under a listed parent zone that resolves to B’s zone), on the unlisted zone, and with replace_mx on any of them, and for nameservers or delegated_subdomain under the platform zone or B’s zone; nothing is written and no MX record is deleted; the listed zone and a zone created for A by nameservers are accepted; the zone_claims row is written with the domain and deleted by delete_zone and by zone_expired; a platform key may use every zone; a partner key cannot set domains.cloudflare_zones (403 scope_denied); with domains.cloudflare_zones: ["pylota.io"] listed, a tenant key adds notify.pylota.io, but adding the pylota.io apex, or replace_mx there, gets 403 scope_denied (zone_not_allowed) (H8) |
it::domains::h6_rule_failure | Literal rule creation fails → address stays pending with routing_rule_failed, retried, activated only with its rule (H6) |
core::domain_fsm::h7_resolver_disagreement | One resolver erroring or disagreeing never changes state; two consecutive agreeing cycles do (H7, FR-DOM-4) |
core::domain_fsm::transition_table | Every row of the state machine table, including 14 days in failing and reminders at 24 h, 72 h and 7 days |
it::send::g7_domain_states | Retiring, pending and failing domain behaviour at send time (G7) |
it::domains::onboarding_idempotent | Adding a cloudflare_zone domain (apex and subdomain) succeeds, and re-running a failed add against the recorded Cloudflare API fake creates nothing twice (FR-DOM-2, FR-DOM-3, FR-OPS-1; build plan M13) |
it::domains::onboarding_idempotent_created_zones | The same for nameservers and delegated_subdomain: a failed add, and onboarding steps 2–8 run by the monitor once the zone is active, re-run without creating anything twice (build plan M23) |
it::domains::records_from_api | Records in responses equal the fake provider’s API answers, never templates (FR-DOM-3) |
it::domains::s9_manual_delivery_events | With the Cloudflare fake answering 403 to the subscription create, a cloudflare_zone domain is created with event_subscription_id = NULL, delivery_events: "manual" and details.action = "run pmail domains subscribe {domain}"; a 503 answer fails the create with 502 upstream_error; once the subscription ID is recorded, delivery_events is active and a delivery event updates the recipient (spike S9 fallback, FR-DOM-3; build plan M13) |
it::domains::s9_manual_delivery_events_nameservers | The same fallback for a nameservers domain, whose step 7 runs in the monitor once the zone is active: the 403 leaves delivery_events: "manual" with the same details.action (build plan M23) |
it::domains::cron_mints_missing_monitor | A row with monitor_do_id = '' (platform, or inserted by pmail domains add --local-token) gets one DomainMonitor, Init, and domain.created when it has a tenant_id; two overlapping cron runs mint one ID |
it::domains::cf_token_required_by_method | Without PM_CF_API_TOKEN: adding a cloudflare_zone domain, and creating an address that needs a literal rule, → 422 cf_token_required (build plan M13) |
it::domains::cf_token_required_other_methods | Without PM_CF_API_TOKEN: nameservers and delegated_subdomain → 422 cf_token_required; dns_records, send_only and smtp_relay are unaffected (build plan M23) |