Built by operators who ran ABA clinics.

IOA computed from the data, not the memory.

A second observer collects alongside the technician, and Wilma® works out the interobserver agreement from what both of them recorded — across every measurement type, scored only against what the observer actually watched.

measurement types scored
Seven
from the data you collect
Computed
to what the observer saw
Windowed
when a score is too thin
Flagged

What is ABA IOA software?

Interobserver agreement, evidenced on the session it describes.

ABA IOA software calculates interobserver agreement — the reliability check that asks whether two trained people watching the same session recorded the same thing. When they did not, the data describes the observers as much as the client, and every clinical decision resting on it inherits that doubt.

The usual version of this is a percentage worked out in a spreadsheet after the fact and typed into a note, which means the figure cannot be re-derived when someone asks how it was reached. In Wilma the observer collects alongside the technician on the same targets, and agreement is computed from the two sets of data that already exist — scored per measurement type, limited to the window the observer was present for, and withheld when it would rest on too few points to mean anything.

IOA is half of the data-quality question. It establishes that the measurement was reliable; treatment fidelity establishes that the plan was delivered as written. Both sit on the same record as the supervision trail, because a funder asking how you know your data is trustworthy is asking all three questions at once.

Two people collect. Wilma® does the comparing.

The technician runs the session and collects as normal. The observer collects alongside them on the same targets. Nobody transcribes anything into a spreadsheet afterwards, because the agreement is calculated from the two sets of data that already exist.

Collect side by side

The observer records on the same targets, in the same session.

No second entry

IOA comes from the session data, not a re-typed copy of it.

In the room or over video

A remote observer watching over a video link is still an observer.

Scored per goal and per session

Agreement rolls up, and you can still see which target dragged it down.

Agreement scored the way each target was measured.

IOA on a trial-by-trial target is a different calculation from IOA on a duration target, and a tool that only scores trials leaves most of a real programme unmeasured. Every capture type Wilma supports is scored on its own terms.

Trial-by-trial

Including agreement on the prompt level, not just correct/incorrect.

Frequency and rate

Counts compared across the observed window.

Duration and latency

Scored against how long, and how long until.

Interval, task analysis, ABC

Interval records, chained steps, and antecedent-behaviour-consequence entries.

Only the part the observer actually watched.

A supervisor who joins forty minutes into a two-hour session did not disagree about the first eighty minutes — they were not there. Scoring an observer against data collected before they arrived produces a low number that means nothing, and the obvious fix of scoring the whole session produces a high one that means just as little.

Scored to the window

Only the period the observer was present counts toward the score.

Grace at both ends

A point captured a moment either side of the window still pairs.

Pairing you choose

Match points by trial order, or by when each was recorded.

Partial observation is normal

A twenty-minute probe is a valid IOA sample, not a broken one.

A number that refuses to flatter you.

Two people who recorded three trials between them can agree perfectly by luck. Most of the ways IOA goes wrong are ways it goes wrong upward, so the scoring is built to withhold a score rather than hand over a comfortable one.

Sufficiency flagged

A score resting on too few points is marked as such, not published quietly.

No phantom agreement

A target neither person recorded on is excluded, never scored 100%.

Unmatched points count

A trial one person recorded and the other did not counts against the score.

Thresholds you set

Minimum points and targets per session are yours to set, not ours.

The answer to "how do you know this data is reliable?"

That question arrives from funders, from auditors, and from your own clinical director. IOA is the part of the answer that can be shown rather than asserted — and it is only showable if it was captured against the session it describes, at the time, on the same record.

On the session record

Agreement sits with the data and the supervision it belongs to.

Exportable results

Take the results out as a document when someone asks for them.

Trended over time

Agreement per technician and per target, not a single snapshot.

Pairs with fidelity

IOA covers the measurement; treatment fidelity covers the delivery.

Buyer's Checklist

Evaluating IOA software? Demand every line.

FAQ

Questions BCBAs ask about interobserver agreement.

What is IOA software?

What is interobserver agreement, in plain terms?

Which measurement types can Wilma score?

What if the observer only watched part of the session?

Can a supervisor run IOA remotely?

How do you stop a meaningless score looking like a good one?

Is IOA the same as treatment fidelity?

Does this replace our supervision records?

How much does it cost?

Is Wilma HIPAA compliant?

Customer Story · Clinical

Why a clinical team left "click-after-click" software behind

How fewer clicks for notes, data, and supervision gave a BCBA their day back.

See IOA scored on a real session.

Thirty minutes. Bring a programme with a mix of measurement types — we'll run two observers over it and show you what comes back, including where it declines to give you a number.

Choosing an IOA tool

How to evaluate ABA IOA software

Almost every platform will tell you it "supports IOA". The differences that matter are not in whether a number is produced, but in which numbers it refuses to produce — because the failure mode of a reliability metric is that it reassures you.

  • 1Is agreement computed, or typed in?A field where a BCBA records the IOA percentage they worked out elsewhere is a text box, not a reliability system. It inherits every arithmetic error and every optimistic rounding in the spreadsheet it came from, and it cannot be re-derived later when an auditor asks how the figure was reached.Red flag: An "IOA %" field on the session note, or an export you are expected to reconcile in Excel.
  • 2Does it score more than trials?Trial-by-trial agreement is the easy case and often the only one implemented. A real programme also carries duration targets, interval records, task analyses and frequency counts, and if those cannot be scored then the reliability evidence covers whichever part of the caseload happens to be discrete-trial.Red flag: Documentation that only ever gives trial examples, or a percentage that appears for some goals and silently not for others.
  • 3What happens when the observer arrives late?Partial observation is the normal case, not the exception. A tool that scores the observer against the whole session punishes them for the part they did not see; one that quietly scores only the goals they touched can flatter them instead. Ask specifically how the observation window is decided and what happens to points falling on its edges.Red flag: No concept of an observation window at all, which usually means the whole session is compared.
  • 4Can the score come back "not enough data"?This is the question that separates a reliability tool from a reassurance tool. Agreement computed over a handful of points is dominated by chance, and a system that always returns a number will hand you 100% on two trials. A system that can decline to score is one you can trust when it does score.Red flag: A percentage for every session regardless of sample size, and no configurable minimum.
  • 5Are empty targets counted as agreement?If neither person recorded anything against a goal, there was nothing to agree about. Counting that as perfect agreement is the single easiest way to inflate an average, and it is a common implementation shortcut because it makes the arithmetic simpler.Red flag: Session averages that rise when an observer opens more goals than they collect on.
  • 6Does the evidence live with the session?Reliability evidence is only useful at the moment someone challenges a clinical decision, which is usually months later. If IOA lives in a separate report rather than against the session, the reconstruction work lands on whoever is least able to refuse it.Red flag: IOA that exists only as a standalone report, with no link back to the session and supervision record.

Keep reading

Related guides on running an ABA practice — the same operation, from a different angle.