Applied Behavior Analysis on Intelligence
No, behavioral therapy cannot raise IQ by ten points
Introduction
Y’know, autism is a very interesting thing. It’s also a surprisingly controversial thing. There are several interesting questions to be asked about it, such as about theory of mind, empathy, whether or not it’s actually a spectrum, etc. However, this article is not about autism in general. Instead, this article is about a treatment for it called applied behavior analysis. Commonly abbreviated as ABA (and from here on), it is a method in which the psychological principles of reinforcement and punishment are used to extinguish behaviors or shape new ones.
Now obviously, there are ethical concerns with this, but that’s not the point of this article either. This article is about a claim which I feel is way more ridiculous than commonly acknowledged. That is; the claim that ABA can raise IQ over ten points. For a quick datapoint on how widespread this claim is, if you search the question “How much does ABA raise IQ in autistics?” then the AI will answer with the following:
Between 9 and 18 IQ points being behaviorally malleable seems extremely prima facie unlikely to me. Sometimes, however, surprising things turn out to be correct. Therefore, in this article, I’ll be assessing this utterly extraordinary claim.
The Research
It’s probably most reasonable to start this section with very early research on the matter, which interestingly isn’t even on autistics. The two earliest studies I could find are Edlund (1972) as well as Clingman & Fowler (1974), both of which have opposite results.
Firstly, Edlund compared children between the ages of 5yo and 7yo (n = 22) on Stanford-Binet scores. The kids were given the test twice, one week apart. The experimental group, meaning those were given candy when getting answers correct, had a much higher score on the retest whereas the control group had approximately the same score.
Clingman & Fowler, however, reported this same behavioral intervention to have a much smaller effect on the IQ of first- and second- grade children (n = 36). The change in IQ was by merely 3 to 4 IQ points.
Both of these studies are obviously flawed because of threats on external validity. For examples,
The control group mean in Edlund is 83 IQ and so there could be a floor effect. This means that the results of this behavioral invention are not necessarily generalizable to people with dull, average, bright, very bright, or gifted intelligence.
The df in Edlund is 10, which may distort the distributions from a normal pattern. This is particularly problematic if the change is not on g, which is not tested herein.
The text of Clingman & Fowler have no information on (1) which IQ test was utilized, (2) the time between take and retake, (3) any way that ceiling effects may or may not have been addressed.
So the early research is very questionable, and the results are opposite to one another. However, since ABA is most commonly done on autistics, it makes more sense to look at that research instead. Virués-Ortega (2010) meta-analyzed Randomized Controlled Trials that been conducted between 1985 and 2009 (k = 20). It was found that the pooled effects of ABA on IQ scores are statistically insignificant.
Some studies reported evidence of harm whereas other studies reported evidence of benefit. There was also evidence of publication bias, although the author questions whether that’s an artifact (p. 398). More recently, Dixon et al. (2021) conducted another RCT on “…twenty-eight participants… (24 male, 4 female; 17 clinic participants, 11 waitlist participants). Twenty-five participants had a diagnosis of ASD, and three had no diagnosis but presented with language delays”. Notably, the control group and the group which received traditional ABA didn’t significantly differ in IQ scores.
The impression of “p=0.001” is given from ANOVA, which is basically a multi-directional t-test. This means that, even though the effects of traditional ABA aren’t significantly different from the control condition, the significant difference between the control condition and comprehensive ABA can still give that impression.
If our test of the null hypothesis is rejected, we conclude that not all the means are equal: that is, at least one mean is different from the other means. The ANOVA test itself provides only statistical evidence of a difference, but not any statistical evidence as to which mean or means are statistically different.
Actually, Beaujean & Farmer (2021). have criticized Dixon et al. on the basis that there are violated assumptions in the randomization process that cast doubt on the similarity of pre-intervention group characteristics (i.e., see the section titled Waitlist Comparison for elaboration). Additionally, they added other aspects to the ANOVA, such as within-group differences over time. The reanalysis, including the flawed waitlist comparison group, resulted in the following.
“The results for the one-way ANOVA and split-plot ANOVA are provided in Tables 2 and 3, respectively. They indicate that the group differences are not statistically significant. The same was true for the Kruskal–Wallis (p=0.34) and Wilcoxon (p=0.37) tests. The effect size (Hedges’ g) for the between-group effect is 0.69, but the confidence interval (CI) ranges from −0.41 to 1.80. While the point estimate of Hedges’ g is small-to-moderate, the 95% CI contains 0.00 which suggests that the estimate is too unstable to be conclusive.”
Another meta-analysis by Rodgers et al. (2021) reported that, among autistic preschool children, that ABA increases IQ by a mean difference of 10.12 (CIs=[5.81, 14.44]) after one year and 11.97 (CIs=[6.74, 17.20]) after two years. There are, however, a few problems with this meta-analysis that cast a lot of doubt on these results.
The change in increase from the first year to the second is evidence of non-linearity. This means that ABA is only beneficial to IQ in the first year but these benefits don’t occur afterwards. Actually, as shown in Figure 6, the beneficial effects of ABA on IQ disappear after three- to four- years.
The Confidence Intervals are very wide, with the lower-bound of the pooled effects being nearly halved.
Assessment of Risk of Bias reveals (1) serious risk of confounding in all studies, (2) moderately biased participant selection processes, (3) moderate risk of deviating from standard ABA procedure, and (4) moderate risks of validity threats to outcome measurement. The possibility of publication bias is not tested nor is it even mentioned anywhere in the text.
So hopefully we can all agree at this point that the extraordinary claims about intelligence are unjustified by the evidence. However, there’s also an obvious theoretical problem worth noting. It’s obviously well-known that the narrow-sense heritability of intelligence is higher amongst older cohorts (e.g., Haworth et al., 2009; Briley & Tucker-Drob, 2015). This is because, as people get older and older, they have an increasing amount of control over their environment. For example, a child cannot choose to buy an electric guitar if their parents don’t want that. An adult, however, can do this. The same logic applies for hanging out with gangs, going hiking in the mountains, reading at libraries, partying and clubbing with friends, etc. Adolescents and adults have an increasing amount of freedom from their families. This is simply an obvious fact. Because of this, you would expect changes in IQ from ABA to disappear after it ends, since the wholly controlled environment is gone. Indeed, Rodgers et al. reported evidence of just that:
In Summary
I think that the conclusion of ABA raising IQ to such a comical extent is very questionable. The research is actually incredibly mixed and, for all the aforementioned reasons, I think the research evidencing these large increases are fatally flawed. Such research is also theoretically problematic because changes in heritability are not considered. These changes in heritability explain the decline of these supposed IQ gains over time, because ABA is a wholly controlled environment, and such extreme controls obviously disappear over time.









If I recall correctly, this study is one of many over decades that has attempted to raise intelligence in the mentally defective, and as I also recall, all have failed. So in that grand tradition, why not try ABA (assuming it’s not been done already). The was a summary of such studies I once read, but cannot put a finger on it right now (old publication, not in wide circulation).
The one problem I tend to have with these “studies” and reviews thereof is the confusion between IQ and intelligence, which was first termed as “g”—a hypothetical concept—by Spearman in 1904. IQ is a test score based on what is considered (hoped?) an approximation of the intelligence construct and such is still debated to this day. I admit to using such measures as IQ freely myself, however.