Engine-agnostic · Multimodal Rev 02 / Live pilot in progress

Signal is not meaning

Emotion engines tell you a face moved. Emotrix tells you what it meant, how confident that reading is, and when the engine is contradicting itself.

Engine output stress0.75    surprise0.78
Ground truth the participant is laughing
Ten sections, about six minutes
01 The gap Market state 2026

Every system now
watches a human.
Almost none
understand one.

Cameras and microphones are default hardware. Voice agents, cabins, kiosks, clinical tools, research platforms: they all sit opposite a person and none of them can say what that person is going through with any defensible confidence.

Not because detection failed. Detection got good. What never got built is the layer that turns detection into something you would put in front of a stakeholder, a regulator, or a decision. The signal arrives raw and everyone is expected to interpret it themselves.

The independent middle has been bought out. What is left is closed full-stack ecosystems above, and raw score APIs below.
Consolidation, 2021 to 2026
AffectivaAcquired · Smart Eye
iMotionsAcquired · Smart Eye
emotion3DAcquired · indie Semiconductor
audEERINGAcquired · Agile Robots
FACET facial codingDiscontinued
Hume Expression MeasurementDiscontinued mid-2026
The row nobody is filling

Independent, engine-agnostic middleware

Vacant

Compiled from public announcements. Verify before citing.

Two of those did not sell. They were switched off. One of them was switched off underneath a live build of ours. That is the whole argument for a middle layer, and the market made it for us.

02 Signal to meaning Four stages, one pipeline

Meaning is
a resolution
problem

A raw affect signal is a diffuse cloud: dozens of muscle coefficients, prosodic features and confidence values arriving at ten times a second, most of which mean nothing on their own. Meaning is what is left after you calibrate it to the individual, resolve its internal contradictions, hold it against time, and throw away everything you cannot stand behind. Emotrix is the four steps between those two states.

Raw signal01
Calibrated02
Resolved03
Defensible04

Engine output. Fifty-two facial coefficients, head pose, voice activity, prosody. Noisy, individual, uncalibrated.

Referenced against this person's own resting state, not a population average. Squinting under bright light stops reading as strain.

Contradictions arbitrated. Competing constructs suppress one another rather than firing together and cancelling out trust.

What survives, with confidence attached and gaps left visibly empty. Fewer claims, every one of them supportable.

03 The middle layer Architecture

Engines are
interchangeable.
The layer is not.

Emotrix sits between the engines that produce signal and everything that consumes meaning. Below it, adapters normalise whatever vendor, model or sensor you happen to be using this quarter. Above it, one contract: research interfaces, product runtimes and agents all read the same session object. Change the engine and nothing above the adapter notices.

Layer 03

Consumers

Research UISession reviewCohort compareProduct runtimeAgent contextAPI / SDK

Read one session object. Timeline, moments, transcript, confidence, absence. Nothing here knows or cares which engine produced the signal.

One contract
Layer 02

The reducer

Baseline calibrationAbsolute scaleCompetitive suppressionTime constantsConfidenceAbsence

The product. Turns normalised signal into calibrated constructs with arbitration, uncertainty and honest gaps. Domain lenses configure how the same signal is read in a cabin, a clinic or a usability lab.

Normalised signal
Layer 01

Engine adapters

Face landmarksAction unitsVoice activityProsodySpeech to textGaze & head poseVehicle telemetry

Commercial APIs, open source models, on-device inference, or your own. Swappable by design, because in this market they will be swapped whether you designed for it or not.

The reducer is the product. Everything else is plumbing that someone will commoditise.

This is the part people get wrong. The instinct is to compete on detection accuracy, which is a race against companies with nine figures of funding and a decade of labelled data. The gap worth owning is one level up: the reasoning that turns coefficients into a claim, and the discipline to say when there is no claim to make.

04 From the field Live pilot, under NDA

Run against
real research
footage

We are working with real UX research footage. Not a demo reel, not a benchmark dataset: awkward real-world material, recorded on rigs never intended for affect analysis, in scenarios where speech is hard to resolve and facial capture is carrying the reading.

That is exactly why it was worth doing. Clean data proves nothing. The findings below are ours, about our method. No source material, participant detail or partner is identified anywhere on this site.

Session timeline, schematic reconstruction
Emotion (signed) Cognitive load Signal unavailable
Illustrative data. Not real footage, not real results. Hatched regions are where the face was not resolvable. Never interpolated.
Finding 01

Two confident readings, both wrong

In one segment the engine reported high stress and high surprise in the same frame. The participant was laughing. A laugh recruits the same facial actions that both mappings read as evidence, so both fired, and both were correct at the muscle level. No improvement in detection accuracy fixes this, because detection was not wrong. It is only resolvable one level up, where constructs are allowed to compete and suppress one another.

stress 0.75 · surprise 0.78 · state: positive
Finding 02

Everyone's neutral face is different

Population-normalised scores tell you how this face compares with a training distribution, which is rarely the question. Calibrating every construct against the participant's own resting state answers the question people actually ask: what changed, for this person, and when. It also removes an entire class of false positive, starting with squinting under bright light reading as strain.

baseline: participant-local · scale: absolute
Finding 03

The rig is the limit, not the model

The face was resolvable for roughly two thirds of the segment and usable speech was scarce enough to count by hand. Rather than smooth over it, the interface hatches every gap and refuses to interpolate. Losing tracking is never reported as looking away. A camera at eye level and a lapel microphone would move coverage past ninety percent, and that finding is worth more to a research team than a confident line drawn through nothing.

coverage ≈ 2/3 · gaps hatched, not filled
05 Proof by accident Engine swap, mid-pilot

The engine
died
mid-build

Halfway through the pilot, the vendor whose API the entire pipeline ran on discontinued it. Not deprecated. Switched off, with the authentication endpoint still answering so that a health check looked fine.

We swapped the adapter underneath for an open-source stack running locally on CPU. Nothing above the adapter changed: not the reducer, not the session object, not a single line of the interface.

We had been arguing for engine-agnosticism as a design position. The market turned it into a demonstration.

06 What the layer actually does Six operations

Six operations
between the score
and the claim

None of these are model improvements. Every one of them is a decision about what a number is allowed to mean, made once, in one place, and applied identically to every session so that two participants can be compared without an asterisk.

01Deviation from self

Every construct is expressed as movement away from this participant's own resting state, established from their own footage. Individual difference stops being noise and becomes the measurement.

02Absolute scale

Normalising within a session makes people incomparable and inflates rarely-used muscles into apparent drama. Constructs sit on fixed full-scale ranges so two sessions can be honestly placed side by side.

03Competitive suppression

Constructs contest the same evidence. Strong positive affect suppresses stress and surprise, because a laugh and a grimace share facial machinery. One state wins the frame instead of three firing at once.

04Four time constants

The same session read at instant, five second, twenty second and whole-arc resolution. A flicker and a mood are different objects, and most tools collapse them into one jittery line.

05Confidence travels with the claim

Every reading carries its own confidence, derived from tracking quality, signal availability and cross-modal agreement. When face and voice diverge, that divergence is the finding, not an error to be averaged away.

06Honest gaps

Where the signal is not there, nothing is drawn. Lanes are zeroed and hatched. No smoothing across a gap, no inference to fill a hole, no confident line through absent data. This is the single most requested behaviour from every researcher who has seen it.

07 Affective context Where this is going

Models are
fluent and
emotionally blind

A voice agent hears the words and infers the mood from the words. It cannot tell the difference between someone who is fine and someone who is saying they are fine. The signal to close that gap exists, in the microphone and camera already pointed at the user. What is missing is a representation an agent can actually reason over.

Emotrix session statet = 412.8s
{
  "t": 412.8,
  "affect": { "valence": -0.42, "arousal": 0.61 },
  "constructs": {
    "cognitive_load": 0.71,
    "stress":         0.58,
    "positive":       0.06,
    "surprise":       0.11,
    "attention_away": 0.00
  },
  "confidence": 0.78,
  "baseline": "participant_local",
  "unobservable": ["voice"],        // mic gated, not silent
  "suppressed": [
    { "construct": "surprise", "by": "positive", "w": 0.70 }
  ],
  "since": 398.2,                    // state held 14.6s
  "engine": "adapter/face-landmark-v2"
}

Why not just hand the model the raw scores?

Because it will write you a story. Give a language model fifty-two coefficients and it produces confident narrative prose about a person's inner life, with no way for anyone downstream to tell which parts were measured. A small calibrated object with its uncertainty and its blind spots declared is the difference between grounding and confabulation.

What does an agent do with it?

Change what it does next. Slow down and stop stacking questions when load is high and climbing. Stop offering help nobody asked for when the state is settled and positive. Hand over to a human when stress holds above threshold for long enough to matter. None of that requires the model to guess how someone feels.

Why does the blind spot field matter?

Because unobservable and zero are opposite facts, and almost every system in this category conflates them. An agent that knows it cannot currently see the user behaves differently, and more safely, than one that believes it is looking at calm.

Where does regulation land on this?

Emotion inference is a restricted and in places prohibited category under the EU AI Act, particularly in workplace and education settings. A layer that records provenance, calibration basis, confidence and absence per reading is not a compliance afterthought. It is the only version of this that survives an audit.

08 Where it applies One layer, configured lenses

The same signal
means different
things in different
rooms

Surprise in a usability test is a design defect. Surprise in a vehicle cabin is a safety event. Surprise in a consultation is something a clinician needs to hear about. The infrastructure does not change between them. The lens does, and the lens is configuration rather than a rebuild.

Primary · Live

Experience research

Emotional response mapped onto the test timeline, aligned to tasks and script. Where the participant struggled versus where they said they struggled.

Primary · Next

Cabin intelligence

Driver and passenger state beyond drowsiness. Load, strain, comfort and disengagement, read against what the vehicle was doing at the time.

Exploratory

Clinical & triage

Voice-led distress indicators to support, never replace, clinical judgement. Built with a healthcare partner on live voice infrastructure.

Proven adjacent

Coaching & rehearsal

Real-time voice affect driving live feedback. Already running in an alpha interview coaching product with real users.

Domain lens

Service & support

Emotional trajectory across an interaction. Escalation predicted from slope, not from a keyword list.

Domain lens

Creative testing

Second-by-second genuine response to work, before the budget commits behind it.

Domain lens

Learning

Confusion and cognitive overload as they happen, subject to the EU AI Act constraints that apply in education.

Domain lens

Agent runtimes

Affective state as structured context for voice and multimodal agents. The section above, in production.

09 Where this actually is Stated plainly

No claims
we cannot
stand behind

The whole argument of this project is that confidence should be attached to claims. It would be a strange product to build and then oversell. So here is the position, without the usual staging.

Reducer and session modelBuilt, on real footage
Engine adapters, face and voiceSwapped once, under load
Research interfaceWorking demonstrator
Field pilot on real study dataIn progress · under NDA
Voice affect in a live productAlpha · external users
Public API and SDKNot yet
Commercial entity and pricingNot yet
RevenueNone

Last reviewed August 2026.

10 Contact Design partners, a few

Who owns the
interpretation?

Somebody on your side is already doing this job. Deciding which of two contradictory readings to believe. Normalising by eye because the scores are not comparable between participants. Explaining to someone about to make a decision why the number says stress and the video shows a person laughing.

It is being done by hand, in a spreadsheet, and it is written down nowhere. That is the layer. The only question is whether anybody owns it on purpose.

gideon@emotrix.cloud

Say which of the two below applies, and one sentence on the problem underneath it. That is enough to start.

If you already have the data

Bring us your hardest footage. You supply existing session recordings and, ideally, the study protocol and any subjective measures. We return a calibrated timeline, the findings, and an honest account of what your rig will and will not support. Difficult material is more useful to us than clean material: compromised rig, wrong lighting, participants who do not talk, that is the interesting case.

If you are building something

Bring the problem instead. You have a product that has to know how a person is doing, and you have reached the point where raw scores are not enough to act on. We work through the contract before anyone writes it: what the state object carries, what it must refuse to claim, how confidence and absence travel with a reading, and where the boundary sits between your product and the layer underneath it.

Mutual NDA before any material moves either way. Processing runs locally on CPU with no third-party API calls required. Nothing identifiable is published, here or anywhere.