How to Calculate IOA in ABA: Every Method, and When Each One Is Honest
Interobserver agreement is the only evidence you have that your data describes the learner rather than the observer.
Interobserver agreement is the check that keeps behavioral data honest. Two people observe the same behavior at the same time, independently, and you compare what they recorded. High agreement does not prove your data is accurate — both observers can be wrong the same way — but low agreement proves it is not trustworthy.
The complication is that IOA is not one calculation. There are several, they produce meaningfully different numbers from identical data, and choosing the flattering one is the most common quiet dishonesty in behavioral measurement.
Total count IOA
The simplest and the weakest.
Formula: smaller count ÷ larger count × 100.
Observer A counts 18 instances, Observer B counts 20. Agreement is 18 ÷ 20 × 100 = 90%.
The problem is that this compares only the totals. If A recorded 18 instances in the first half of the session and B recorded 20 in the second half, agreement still computes to 90% while the two observers did not agree on a single actual occurrence. Total count IOA is appropriate for a quick check on frequency data and nothing more.
Exact count-per-interval IOA
The strict fix for the problem above. Divide the session into intervals, and score an interval as agreement only if both observers recorded exactly the same count in it.
Formula: intervals with exact agreement ÷ total intervals × 100.
This is demanding and produces lower numbers than total count, which is the point. It is the most rigorous IOA for frequency data and the one to use when the stakes are high.
Mean count-per-interval IOA
A middle position. Calculate agreement within each interval using the total count formula, then average across intervals.
More sensitive than total count because it requires agreement across the session rather than just in aggregate, more forgiving than exact count because near-misses still earn partial credit.
Trial-by-trial IOA
For discrete trial data, where each trial has a defined outcome — correct, incorrect, prompted.
Formula: trials with agreement ÷ total trials × 100.
Clean, appropriate, and the default for discrete trial teaching. The one caution is that agreement inflates when a learner is at floor or ceiling: if a learner gets every trial correct, two observers agree perfectly without demonstrating any discrimination.
Interval-by-interval IOA
The method for interval recording data. Compare scores interval by interval, count agreements, divide by total intervals.
Also called point-by-point agreement. This is where the inflation problem is most severe. If a behavior occurs in 5% of intervals, two observers who both score almost everything as non-occurrence will agree on roughly 95% of intervals while agreeing on almost none of the actual behavior.
Which is why interval data needs one of the following instead.
Scored-interval IOA
Count only intervals where at least one observer scored an occurrence. Ignore intervals both scored blank.
This is the honest calculation for low-rate behavior. It strips out the free agreement on empty intervals and reports agreement on the behavior that actually happened.
Unscored-interval IOA
The mirror image: count only intervals where at least one observer scored non-occurrence.
This is the honest calculation for high-rate behavior, where the free agreement comes from both observers scoring nearly every interval as occurrence.
Duration and latency IOA
Total duration IOA: shorter total duration ÷ longer total duration × 100. Same aggregation weakness as total count.
Mean duration-per-occurrence IOA: calculate agreement for each occurrence, then average. More rigorous and the better default.
Latency IOA follows the same pattern, comparing latencies occurrence by occurrence.
What number is acceptable
The conventional floor in applied behavior analysis is 80%, with 90% or above preferred, particularly for data supporting treatment decisions or publication. Those are conventions rather than statutes — funders, accreditors and journals set their own expectations, and some are stricter.
More important than the threshold is what you do when you miss it. Agreement below the floor is not a number to report and move past. It means retraining the observers, tightening the operational definition, or both — and it means the data collected during that period should not drive a phase change.
How often to run it
Common practice is IOA on a minimum of 20% to 33% of sessions, distributed across phases, conditions, observers and times of day rather than clustered wherever it was convenient.
The distribution matters as much as the percentage. Twenty percent of sessions all collected in baseline, by the same two people, on Tuesday mornings, tells you very little about the reliability of your data in treatment.
The three failures that make IOA meaningless
Observers who can see each other's data. Independence is the entire premise. Two people scoring on one screen, or one calling out scores, produces a number that measures compliance rather than agreement.
An operational definition that gets negotiated afterwards. If observers discuss discrepancies and revise their sheets before the calculation, you are measuring their ability to reach consensus, not their agreement.
Reporting the flattering method. Running interval-by-interval agreement on a behavior that occurs in 4% of intervals and reporting 94% is technically a correct calculation and substantively misleading. The right method for that data is scored-interval agreement, and it will produce a much lower number — which is the information you needed.
Making IOA routine rather than an event
IOA collapses in practice for logistical reasons: getting two people observing simultaneously and independently is a scheduling problem, and the calculation is fiddly enough that it gets deferred and then skipped.
The practices that sustain it treat IOA as a scheduled property of the program rather than a task someone remembers. That means the second observer is assigned when the session is scheduled, both collect on independent devices, and the agreement calculation runs automatically against the method the program specified rather than by hand in a spreadsheet weeks later.
Worth being straightforward here: interobserver agreement scoring is on Wilma's roadmap and is not something the platform calculates for you today. If IOA automation is a requirement for you this quarter, weigh that honestly — what Wilma does support now is independent data capture per observer on the same session, which is the input any IOA calculation needs.
The summary
Match the IOA method to the measurement system, and match it to how often the behavior actually occurs. Total count for a rough check on frequency. Exact count-per-interval when it matters. Trial-by-trial for discrete trials. Scored-interval for low-rate behavior and unscored-interval for high-rate behavior, never plain interval-by-interval at the extremes. Aim for 80% at minimum and 90% where decisions ride on it, collect on at least 20% of sessions spread across conditions, and when agreement is poor, fix the definition rather than the report.
Frequently asked questions
What is an acceptable IOA percentage in ABA?
The conventional minimum is 80%, with 90% or higher preferred for data supporting treatment decisions. These are professional conventions rather than fixed rules, and funders, accreditors and journals may set stricter expectations.
How often should IOA be collected?
Common practice is a minimum of 20% to 33% of sessions. The distribution matters as much as the proportion — spread collection across phases, conditions, observers and times of day rather than clustering it where it was convenient.
Why is interval-by-interval IOA misleading?
It gives free credit for agreement on intervals where neither observer scored the behavior. With a behavior occurring in 5% of intervals, two observers can reach 95% agreement without agreeing on any actual occurrence. Use scored-interval IOA for low-rate behavior and unscored-interval IOA for high-rate behavior instead.
What should I do when IOA falls below threshold?
Treat it as a signal that the data cannot support decisions, not as a number to report and move past. Retrain the observers, tighten the operational definition, or both — and do not make a phase change based on data collected during that period.