bRRAIn Certified Maintenance Specialist
Keep a bRRAIn pod healthy, upgraded and recoverable, and prove the organization's memory survives every change.
- Level
- Practitioner
- Learning time
- 20 hours
- Price
- $499
- Credential
- Valid 3 years
What changed in this edition.
- Teaches the real deployment model: one dedicated brain pod per organization, the eight-zone architecture (Z1 Vault to Z8 Code Sandbox) and the Handler as bRRAIn's own model running on the pod.
- Upgrades now follow the actual release channel: release-notes preflight, brrain upgrade with -check, -dry-run, -pin and -rollback, the console's in-place Upgrade, extension Update, and supervised auto-rollback after repeated crashes.
- Backup and recovery use the real tools: brrain dr snapshot, restore, failover, migrate and status, brrain export and brrain recover verify, with restore tests and DR drills.
- Capacity planning is rebuilt around the pod's GPU and vault storage using the console dashboard's compute and storage charts, and covers pod recreates that change the URL while the vault volume survives.
- New integrity routine: brrain audit verify, record-stream checks, POPE tag checks and silent-loss detection, with forward repair under the Correction & Supersession standard.
- Every module now opens with a diagnostic pretest and closes with a retrieval check; ten AI role-play labs including the capstone; a new exam bank assembled per candidate.
What you will be able to do.
- You will be able to build a monitoring routine for a bRRAIn pod from its real health signals, with baselined thresholds and disciplined noise tuning.
- You will be able to profile performance complaints and decide, with evidence, between efficiency fixes and capacity changes.
- You will be able to plan and execute in-place brain and extension upgrades from release notes, verify version agreement, and handle manual and automatic rollback.
- You will be able to prove, with before-and-after evidence, that an operation preserved files, POPE tags, records, canonical context, access grants, audit evidence, extensions and client reach.
- You will be able to operate encrypted snapshots against RPO and RTO, prove restorability and recovery-key unlock, and run DR drills and failovers.
- You will be able to run routine integrity verification, detect silent data loss, repair forward without rewriting history, and report findings others can act on.
- You will be able to plan GPU and storage capacity and execute capacity changes and hosting migrations without losing memory or reach.
- You will be able to coordinate maintenance with the Operations Controller, Security Controller, Access Controller and Care Analyst, including governed outside access.
Who it's for
- SREs and platform engineers responsible for an organization's bRRAIn pod
- IT operations leads who run upgrades, backups and capacity for bRRAIn
- Partner-firm Care-practice engineers maintaining customer deployments
- Administrators moving from Installation Specialist to steady-state operations
Not covered here
- Analysis of memory growth and usage (see Care Analyst: they analyze, you maintain)
- Permission architecture and role design (see Access Controller)
- Audit execution and security incident ownership (see Security Controller)
- Initial deployment (see Installation Specialist)
9 modules, 71 lessons.
About 20 hours of learning. Open a module to see every lesson.
-
The Maintainer's Map of a bRRAIn Deployment
The eight zones at maintenance depth, the one-pod-per-organization model, the install channel and on-disk layout, the real health signals, and the failure modes unique to AI memory.
- Module 1 pretest
- The eight zones at maintenance depth
- Where the health signals are
- Failure modes unique to an AI memory deployment
- The install channel and what lives where on disk
- Lab 1: First-week health triage on an inherited pod
- Module 1 retrieval check
-
Health Monitoring and Alerting
Build a monitoring routine from real signals, set thresholds from your own baseline, tune noise without losing signal, and recognize the five monitoring anti-patterns.
- Module 2 pretest
- Building the monitoring routine
- Setting thresholds from your own baseline
- Noise tuning without losing signal
- Monitoring anti-patterns
- Lab 2: Build and tune Calder Freight's monitoring routine
- Module 2 retrieval check
-
Performance Diagnosis
Profile slow surfaces, understand the search index, caches and model path behind them, choose remediation patterns, and separate capacity from efficiency.
- Module 3 pretest
- Profiling: timing each surface
- Search and the pod's caches
- The model path: Handler, LLM Registry and Ask latency
- Performance remediation patterns
- Capacity or efficiency?
- Lab 3: Profile and remediate a performance regression
- Module 3 retrieval check
-
Upgrades and Release Management
Plan upgrades from release notes, execute verified in-place upgrades with brrain upgrade or the console Upgrade, handle manual and supervised rollback, meet an internal security patch SLA, and write runbooks with version agreement checks.
- Module 4 pretest
- Planning an upgrade: the release-notes preflight
- Executing the upgrade: brrain upgrade and the in-place Upgrade button
- Security fixes and your patch SLA
- Rollback: manual and supervised
- The upgrade runbook and version agreement
- Lab 4: Run a brain upgrade in a maintenance window
- Module 4 retrieval check
-
Preserving Memory Through Change
Signature module. Catalog what must survive every operation (files, POPE tags, sidecars, records, canonical context, access, audit evidence, extensions and reach) and prove preservation with a pre/post evidence pack.
- Module 5 pretest
- What must survive: the memory artifact catalog
- POPE tags and frontmatter through change
- Session, decision and stage records
- Audit log continuity
- Canonical context and standards files
- Access grants and auditor sessions through change
- The pre/post preservation check
- Lab 5: A bulk import that damaged the memory
- Module 5 retrieval check
-
Backup, Restore and Disaster Recovery
Set snapshot cadence against RPO and RTO with brrain dr snapshot, prove restorability and recovery-key unlock, choose and execute restores and failover, and run DR drills.
- Module 6 pretest
- Snapshot cadence against RPO and RTO
- Proving a backup: restore tests and recovery-key checks
- Restore and failover in production
- Disaster-recovery drills
- Lab 6: Quarterly DR restore drill
- Module 6 retrieval check
-
Capacity Planning and Scaling
Model vault and snapshot storage growth, assess GPU and compute headroom on the org's pod, execute capacity changes including pod recreates, and run capacity reviews with the Care Analyst.
- Module 7 pretest
- Vault and storage growth
- GPU and compute sizing
- Executing a capacity change
- Coordinating capacity with the Care Analyst and the business
- Lab 7: Capacity review and a pod recreate plan
- Module 7 retrieval check
-
Integrity Verification
Run audit, tag and record verification on a cadence, detect silent data loss, repair forward without rewriting history, and write integrity reports others can act on.
- Module 8 pretest
- Routine audit verification
- Tag and graph integrity checks
- Verifying records: who decided what, and when
- Detecting silent data loss
- Remediating damage without rewriting history
- The integrity verification report
- Lab 8: Audit continuity after a restore
- Lab 9: Monthly integrity verification and report
- Module 8 retrieval check
-
Production Issues, Migrations and Coordination
Govern outside access during maintenance, diagnose recurring production issues from a catalog, migrate between hosting options with explicit cutover, coordinate with the other roles, and complete the capstone.
- Module 9 pretest
- Governed outside access during maintenance
- The production issue catalog
- Migrations between hosting options and versions
- Working with the Security Controller and the other roles
- Module 9 retrieval check
- Capstone: A maintenance month at Calder Freight
Practice against someone who pushes back.
Labs run in your browser as AI role-plays. An AI plays the person on the other side of the scenario — with their own goals and objections — and your work is scored against the published rubric. There is nothing to install.
-
Lab 1 · The Maintainer's Map of a bRRAIn Deployment
AI role-play: first-week health triage on an inherited pod with an IT operations lead
-
Lab 2 · Health Monitoring and Alerting
AI role-play: building and tuning a monitoring routine with a tired on-call engineer
-
Lab 3 · Performance Diagnosis
AI role-play: profiling a regression with a practice manager demanding a bigger GPU
-
Lab 4 · Upgrades and Release Management
AI role-play: planning and running an in-place brain upgrade in a maintenance window
-
Lab 5 · Preserving Memory Through Change
AI role-play: scoping memory damage from a bulk import with a clinical operations director
-
Lab 6 · Backup, Restore and Disaster Recovery
AI role-play: a quarterly DR restore drill with a managing partner
-
Lab 7 · Capacity Planning and Scaling
AI role-play: a capacity review and pod recreate plan with an Operations Controller
-
Lab 8 · Integrity Verification
AI role-play: verifying audit continuity after a partial restore with a Security Controller
-
Lab 9 · Integrity Verification
AI role-play: a monthly integrity verification and report for an Operations Controller
-
Lab 10 · Production Issues, Migrations and Coordination
AI role-play: a compressed month of maintenance at a growing organization, ending in a monthly report
A maintenance month at Calder Freight
AI role-play scored against the published rubric
Pass mark: 72%
Scored on
- Upgrade and rollback handling25%
- Backup and recovery assurance20%
- Capacity and performance reasoning15%
- Memory preservation and integrity verification25%
- Communication, coordination and reporting15%
One exam. A credential anyone can verify.
The exam
- Items per form
- 59
- Time allowed
- 120 min
- Pass mark
- 72%
- Performance tasks
- 4
- Attempts included
- 2
- Wait between attempts
- 7 days
- Online and timed, taken on learn.brrain.io.
- Your form is assembled for you from the course's item bank, so no two candidates sit the same paper.
- Performance tasks are conducted by an AI examiner: you work through a realistic scenario and are scored against a published rubric.
The credential
- A verifiable digital badge in your name.
- A public verification page at learn.brrain.io/verify, so an employer or client can confirm it.
- Valid for 3 years.
- Renewal: At 3 years, by completing current CE modules and passing the then-current exam
Where this course sits.
Stacks well with
Frequently asked.
Do I need to install anything for the labs?
No. Labs and the capstone run in your browser on learn.brrain.io as AI role-plays: an AI plays the person on the other side of the scenario, and your work is scored against the rubric published with the course.
How is the exam delivered?
Online and timed: 59 items in 120 minutes, on a form assembled for you from the course's item bank. 4 of the items are performance tasks conducted by an AI examiner: you do the work rather than pick an answer. The pass mark is 72%.
What if I don't pass first time?
You have 2 attempts, with a 7-day wait after an unsuccessful attempt. Further exam attempts can be bought for $299 each.
How long is the credential valid?
3 years. You receive a verifiable digital badge with a public verification page at learn.brrain.io/verify, so anyone can confirm it is genuine.
I hold the v1 credential. Is it still valid?
Yes. Credentials earned on v1 remain valid and verifiable at learn.brrain.io/verify. When you renew, you sit the then-current version of the exam.
Can my company enroll a team?
Yes. Firms can buy a certification bundle for $2,999 per firm per year — see the pricing page — or contact us to arrange enrollment for a larger group.
bRRAIn Certified Maintenance Specialist
Keep a bRRAIn pod healthy, upgraded and recoverable, and prove the organization's memory survives every change.