Before any real user reached either BHS-powered pathway, the system had to demonstrate it was safe to place in the hands of a vulnerable person caring for another vulnerable person. Not whether it could respond. Whether it could be trusted not to cause harm. This page documents the evaluation frameworks, the methodology, and the results that answer that question.
Both BHS pathways serve caregivers who are themselves under significant psychological stress, often in the absence of immediate human support. That context sets an evaluation standard that is not about quality of response in any subjective sense. It is about whether the system can be safely placed in that situation without making things worse. Synthetic scenario testing allows that question to be answered rigorously, repeatedly, and without exposing vulnerable individuals to an unverified system. Every scenario is grounded in what real caregivers actually experience, not constructed hypotheticals.
No real user interacts with a BHS-powered pathway until the system has demonstrated it will not cause harm across the full risk spectrum. Synthetic testing allows that bar to be met without placing vulnerable individuals in contact with an unverified system.
Scenario libraries are built from what real caregivers actually say, ask, and express, sourced from primary experience and direct community observation. The evaluation covers the real range of what this population goes through, including the moments most clinical instruments do not capture.
Misreading the emotional state of someone in distress is not a quality issue. It is a clinical failure. The evaluation framework treats emotional accuracy as a hard requirement, not a soft scoring dimension.
Every test suite includes scenarios spanning the full risk spectrum, including acute crisis and thoughts of harm. The hard gate requirement is on safe crisis handling: the system must respond appropriately and must never under-escalate in a way that puts a user at risk. Exact severity classification is scored separately. Conservative escalation, classifying a situation as more serious than the baseline expectation, is treated as a pass because it is clinically safer than under-detection.
The Bloomb evaluation process was built from the ground up, starting not with a test framework but with primary source material: the actual experiences, questions, and expressions of postpartum mothers. The methodology follows five distinct phases. The same principle, grounding scenario libraries in direct experience rather than constructed hypotheticals, applies to every new BHS pathway evaluation.
The scenario library began with first-hand documentation of postpartum experience: the questions that arise, the moments that feel unsupported, the things that are difficult to say to anyone available during office hours. This formed the epistemological foundation of the scenario library, not a constructed approximation of what the population might experience.
This was expanded through direct observation in postpartum communities on Reddit and TikTok, real spaces where mothers speak candidly about their experiences. Questions, confessions, fears, and moments of joy were collected to build a broader picture of what mothers actually express and ask.
Across the collected data, patterns emerged: recurring emotional themes, types of questions, and categories of experience. These were grouped into scenario families covering emotional check-ins, relationship strain, physical recovery, identity shift, breastfeeding challenges, crisis moments, and more.
Rather than recruiting real mothers for initial testing, which carries ethical and emotional risk, an automated system was built to synthesize caregivers asking questions from the ground truth and receive responses. This allows rigorous, repeatable testing without exposing vulnerable individuals to an unverified system.
Every response is evaluated across multiple dimensions: overall quality, emotional accuracy, risk classification, technique appropriateness, and crisis handling. Results are documented per scenario and aggregated into a versioned report available for IRB submission, insurer review, and enterprise due diligence.
The 45 Bloomb scenarios were designed to span the full range of what a postpartum mother might experience, from moments of joy and small wins to the most serious crisis situations. No sanitized version of the postpartum experience. The system had to demonstrate it could handle what real mothers go through before any real mother used it.
Everyday feelings, mood shifts, and the range of emotional states across the postpartum arc.
Pain, healing, body changes, and the physical reality of the postpartum period.
Feeling invisible, changed, or disconnected from who you were before becoming a mother.
Partner dynamics, family pressure, isolation, and the relational weight of early parenthood.
Difficulty, guilt, decisions about stopping, and the emotional complexity of infant feeding.
When things are going well. The system must not pathologize calm or misread denied feelings as present ones.
Yellow and orange risk levels. Sustained distress, emerging crisis signals, and the trajectory toward acute risk.
Red risk level, including thoughts of self-harm. These are the scenarios the system must never get wrong.
Clinical boundary enforcement and prompt injection resistance. The system must stay in its lane and cannot be manipulated out of it.
The SAFE Standard is a BHS-authored clinical evaluation framework built around four non-negotiable criteria: Secure, Accurate, Focused, and Explainable. Hard gate dimensions require a 1.0 score on safe crisis handling, privacy preservation, and medical guardrail adherence. Crisis severity classification is scored separately: exact match passes, and conservative escalation also passes because classifying a situation as more serious than expected is clinically safer than under-detection. What cannot pass is under-escalation that leaves a user without appropriate support. The framework was designed specifically for behavioral health AI serving vulnerable populations, grounded in the principle that a wrong answer in this context is not a product issue. It is a patient safety event.
The evaluation was completed in March 2026 before any real user interacted with the Bloomb system. Version 1.0. Pilot Phase.
The two partial results in CrisisEval and AlertEval reflect scenarios where the system classified crisis severity at an adjacent level rather than the expected exact level. In both cases the response was appropriate, the user would have received the correct level of support, and no harm resulted. These are classification precision gaps, not safety failures. They are tracked as improvement areas and will be re-evaluated before any model changes reach live users.
The SMART 40 framework was defined by the grant reviewer who independently assessed the Clover program, not by BHS. Passing an externally defined evaluation framework is a different kind of credibility from passing one you wrote yourself. The evaluation ran 40 cycles across stress, boundary, and standard scenario categories covering the autism caregiver population.
The 77% technique selection score reflects 4 scenarios where the system returned fallback techniques rather than population-specific ones. Root cause: missing context edges at the boundaries of the Clover ontology. These gaps have been identified, are being addressed in the knowledge graph, and will be re-evaluated before the next release. All safety, crisis, and guardrail dimensions held at 100%.
The SMART 40 was defined by the grant reviewer who independently assessed the Clover program. BHS did not author the evaluation criteria. The reviewer defined the scenario categories, the pass thresholds, and the dimensions to be assessed. Clover was run against those criteria as defined.
That distinction matters. Self-evaluation, however rigorous and however transparent about its methodology, does not carry the same weight as evaluation against criteria an external party set. The SAFE Standard is BHS-authored and explicit about that. The SMART 40 result is independent and explicit about that too.
VERA-MH is an open-source mental health AI safety benchmark published by Spring Health. BHS is implementing VERA-MH evaluation across both the Bloomb and Clover pathways. Running against an industry-published benchmark allows BHS results to be compared against a common standard rather than only against BHS-authored criteria.
VERA-MH implementation is currently in development. Results will be published on this page when available.
Dimensions Under Implementation
Evaluation of whether the system produces responses that are appropriate, safe, and clinically coherent from the perspective of the end user receiving them.
Assessment of whether the system maintains its clinical boundaries under adversarial conditions, including prompt injection attempts and attempts to redirect the system outside its intended scope.
Evaluation across a range of user personas representing different demographics, communication styles, and presentations of distress, to surface variance in system performance across population subgroups.
A compliance officer, IRB reviewer, or enterprise partner asking hard questions about a behavioral health AI product needs more than a pass rate. They need a documented methodology, named evaluation dimensions, honest disclosure of what did not score at threshold and why, and a clear re-evaluation commitment when the system changes. That documentation exists for both BHS pathways and is available on request.
Full scenario-level evaluation results, methodology documentation, and aggregate findings are available in formats suitable for IRB submission. Per-scenario detail, pass thresholds, and dimension-level scoring all included.
When your legal, compliance, or security team asks what the system did in a crisis scenario, BHS has a documented answer. Named frameworks, versioned results, honest disclosure of partial scores, and a re-evaluation commitment on every significant update.
You define the scenarios that matter to your population or your research question. BHS runs a full SAFE Standard evaluation against your inputs and delivers a custom report before any real users are involved. Suitable for platforms with IRB requirements or specific clinical focus areas.
The complete evaluation documentation, including per-scenario detail and methodology notes, is available to research partners, health systems, enterprise partners, and IRB reviewers under NDA on request. A 30-minute briefing with the founder is also available for investors and pilot partners.