A healthy HTTP process does not prove the whole application is ready. Begin with GET /api/health for liveness, then test a signed-in read and a synthetic edit. Inspect the corresponding saved request, workflow run, or email state before assuming a background process succeeded. Keep private content and credentials out of shared logs.

SymptomCheck first
Sign-in fails or loopsMatching Clerk keys, exact authorized origins, callback registration, and production/development application alignment
AI request stays queuedATLAS_WORKER_MODE, supervised worker status, role selection, embedded startup flags, and worker database target
AI is unavailableProvider configuration and saved failure state; choose the built-in helper explicitly only for supported tasks
Search misses an emailMailbox ownership, completed sync, folder coverage, filters, and GET /api/search/status
A record edit conflictsWorkspace revision changed; refresh and review again instead of forcing the stale proposal
Attachments disappear after restartPersistent data mount, ownership, and whether the correct private directory was mounted
Email connection succeeds but sending failsDrafts/Sent capabilities, delegated Send As permission where relevant, and the saved provider error
Phone cannot reach the appDevice-reachable HTTPS API origin; physical-device localhost points to the device

Recover without creating a second problem#

For an uncertain send, reconcile the original message before attempting an authorized resend. For a failed AI job, use the saved retry flow so context and staged work can be preserved. For a database migration failure, inspect the failing migration and checksum; create a corrective migration rather than editing applied history.

A pending semantic index should not block normal editing or keyword search. The SQLite search sidecar is rebuildable: with the server stopped, move aside search-v1.sqlite and its WAL/SHM companions, then restart to requeue indexing. Never remove atlas.sqlite, which contains canonical application data. PostgreSQL indexing has its own durable job and source-version checks.

Before considering recovery complete, repeat the original user flow, restart the relevant service, and verify persistence. If the issue involved files or storage, include a restore-verification check rather than relying only on a successful page load.