Skip to content
NexusMortgageOS

Reliability

Platform status, monitoring & backups

A live health probe plus the monitoring, alerting and recovery commitments behind it — the page a lender's IT reviewer asks for before signing.

Checking systems…

Live probe refreshes every 60 seconds

Uptime target

99.9%

monthly

Data loss ceiling

5 min

RPO

Restore target

4 hrs

RTO, full platform

Probe frequency

60 sec

multi-region

Monitoring & alerting

What we watch, and what wakes someone up

Uptime probe

GET /api/public/health every 60s from multiple regions

2 consecutive failures → page on-call

Database health

Connection, latency and error-rate sampling

p95 query latency > 800ms for 5 min → alert

Vendor credential harness

Nightly golden-path re-verification of every sealed credential

Any failed verification → tenant admin notified

Authentication

Sign-in success rate and session errors

Success rate < 95% over 10 min → page

Billing webhooks

Payment event delivery and processing failures

Any unprocessed webhook > 15 min → alert

Error tracking

Client and server exception capture with release tagging

New error affecting > 5 sessions → alert

Backups

Encrypted, tenant-isolated, restore-tested

Continuous WAL archiving

Point-in-time recovery to any moment in the retention window

RPO ≤ 5 minutes
Daily full snapshot

Encrypted at rest, retained 7 days on the managed platform

Nightly 02:00 PT
Weekly long-term copy

Retained 90 days for regulatory restore requests

Sundays
Document storage

Object storage with versioning and tenant-scoped path isolation

Continuous

Disaster recovery

Objectives, testing and tenant-level rollback

Recovery Point Objective (RPO)

5 minutes — the maximum data loss in a total-failure scenario, backed by continuous write-ahead-log archiving.

Recovery Time Objective (RTO)

4 hours for full platform restoration; 1 hour for a single-tenant restore from snapshot.

Restore testing

A restore drill is performed quarterly against a scratch environment, and the result is recorded in the audit ledger.

Tenant-level restore

A single workspace can be rolled back to a point in time without affecting other tenants, on request through Sev-2 support.

Regional failure

Compute runs on a global edge network with automatic failover; the database platform supports cross-zone recovery from snapshot.

For your monitoring team

Point any external uptime service here

GET https://nexusmortgageos.com/api/public/health → 200 { "status": "ok" } · 503 when degraded

The endpoint is unauthenticated, returns no customer data, and is safe to poll every 30–60 seconds. Incidents are posted here and emailed to workspace administrators per the Support & SLA policy.