Friday, October 16, 2009

J. Williams (2007, Texas) IQ MR Atkins death penalty decision posted: Many issues raised

Another interesting Atkins court decision has been posted to the court decisions section of this blog.  Jeffrey Williams (Williams v Quarterman, Texas, 2007).  Augmenting this court decision is a copy of a psychological report included in the decision.  An initial read of the decision and the psychological report raises a number of interesting questions and issues, such as:



  • The apparent role of a judge becoming the psychometric expert when she decides which part score (verbal, nonverbal, full scale) to use to establish Williams level of intellectual functioning.  Kevin Foley has provided a guest post about this issue, which will be the next post at this blog.
  • The continued issue of part vs full scale IQ scores in Atkins cases. 
  • The role of  school records and school special education decisions made during elementary and secondary schooling in Atkins decisions.
  • The role of measured academic achievement in Atkins decisions, particularly the issue of whether a person who may be mentally retarded can achieve above their measured IQ (note - I've written about this previously and will make comments with appropriate links in a separate post..the answer is "yes").
  • The conceptually and technically messy issue of determining pre-incarceration levels of adaptive behavior retrospectively.  I'm hoping that a few experts in Atkins AB assessment will review the documents and weigh in on the information provided, expert opinions, and final decision.  This area is a definite methodological quagmire in Atkins cases. 
  • The whole issue of malingering and how to detect it
I will be making my own follow-up comments regarding some of these issues in the near future.

Technorati Tags: , , , , , , , , , , , , , , , , , , , ,

Thursday, October 15, 2009

The psychology of expert testimony: Select research

Below are three empirical studies addressing the evaluation of expert testimony and variables impacting the acceptance of expert testimony.  They have been sitting in my "to read" in box since I started this blog.  I had intended to digest all three and post some general comments.  I now realize I'll likely never get to these in the near future, so rather than having them sit and be of no value, I instead have decided to post the references, abstracts, and provide links to the articles to those who want to read more in depth.  These are only a handful of studies.  It is my guess that there is a large body of empirical research (and books and/or book chapters) on expert testimony.  This post is not intended to be comprehensive.

The decision to post these three empirical studies was spurred by my last nights reading of a very interesting article in the current issue of the American Psychologist (2009, Vol 64, No. 6).  The article (Conditions for intuitive expertise: A failure to disagree) is by Daniel Kahneman and Gary Klein.  Below is the abstract following by some key conclusions.
This article reports on an effort to explore the differences between two approaches to intuition and expertise that are often viewed as conflicting: heuristics and biases (HB) and naturalistic decision making (NDM). Starting from the obvious fact that professional intuition is sometimes marvelous and sometimes flawed, the authors attempt to map the boundary conditions that separate true intuitive skill from overconfident and biased impressions. They conclude that evaluating the likely quality of an intuitive judgment requires an assessment of the predictability of the environment in which the judgment is made and of the individual’s opportunity to learn the regularities of that environment. Subjective experience is not a reliable indicator of judgment accuracy.

I zeroed in on the articles material that dealt with issues surrounding the development of expertise and conditions impacting expert judgments.  Although the research is not specific to expert testimony in court proceedings (as are the other three articles listed in this post), I found some specific conclusions worthy of mention and reflection, particularly after reading and thinking about a number of Atkins rulings discussed at this blog (all that included dueling psychological experts).  Below are a few conclusions and statements worthy of consideration.  Unless otherwise noted via underlining or bracket comments, these are all direct quotes from the article.
  • Kahneman coined the term illusion of validity for the unjustified sense of confidence that often comes with clinical judgment.
  • The intuitive judgments of some professionals are impressively skilled, while the judgments of other
    professionals are remarkably flawed.
  • Skilled judges are often unaware of the cues that guide them, and individuals whose intuitions are not skilled are even less likely to know where their judgments come from.
  • True experts, it is said, know when they don’t know. [emphasis added by blogmaster...I think this is a critical finding] However, nonexperts (whether or not they think they are) certainly do not know when they don’t know. Subjective confidence is therefore an unreliable indication of the validity of intuitive judgments and decisions.
  • The situation that we have labeled fractionation of skill is another source of overconfidence. Professionals who have expertise in some tasks are sometimes called upon to make judgments in areas in which they have no real skill.... It is difficult both for the professionals and for those who observe them to determine the boundaries of their true expertise. [emphasis added by blogmaster - this is so true and dangerous.  People often assume, since I develop intelligence tests and conduct research on intelligence theories, that I must know everything about these two topics.  Often when asked questions that are at my boundaries of expertise, it is tempting to provide an answer based on partial knowledge....but, consistent with the prior point emphasized above, the more professional response is to recognize one's limits of expertise and simply say "I don't know."  I wonder how often psychological experts in Atkins cases are faced with this fractionalization of skill expertise conflict?]

Kruass, D. A. & Sales, B. D. (2001). The effects of clinical and scientific expert testimony on juror decision making in capital sentencing. Psychology, Public Policy, and Law, 7 (2), 267-310. (click to view)
The Supreme Court and many state courts have assumed that jurors are capable of differentiating less accurate clinical opinion expert testimony from expert testimony based on more sound scientific footing and of appropriately weighing these two types of testimony in their decisions. Persuasion and jury decision-making research, however, both suggest that this assumption is dubious. The authors investigated whether mock jurors are more influenced by clinical opinion expert testimony or actuarial expert testimony. Results suggested that jurors are more influenced by clinical opinion expert testimony than by actuarial expert testimony and that this preference for clinical opinion expert testimony remains even after the presentation of adversary procedures. Limited empirical evidence was found for the notion that various types of adversary procedures will have a differential impact on the influence of expert testimony on juror decisions.

Levett, L. M. & Kovera, M. B. (2009).  Psychological mediators of the effects of opposing expert testimony on juror decisions.  Psychology, Public Policy, and Law, 15 (2), 124–148. (click to view)
This study examined the effectiveness of the opposing expert safeguard against unreliable expert testimony and whether beliefs about experts as hired guns and general acceptance mediate the effect of opposing expert testimony on juror decisions. We found strong evidence that the presence, but not the content, of opposing expert testimony affected jurors’ trial judgments and that these effects were mediated by mock jurors’ beliefs about general acceptance. The presence of an opposing expert affected jurors’ ratings of the general acceptance of research investigating sexual harassment in the workplace. Jurors’ beliefs about general acceptance then affected jurors’ ratings of plaintiff expert competence and research, which affected juror ratings of the probability that the plaintiff experienced a hostile work environment.

Schweitzer, N. J. & Saks, M. J. (2009).  The gatekeep effect: The Impact of Judges’ Admissibility Decisions on the Persuasiveness of Expert Testimony. Psychology, Public Policy, and Law, 15 (1), 1-18) (click to view)
In a pair of mock-trial studies of a possible “gatekeeper” effect, our participants were presented with a summary of a trial that included a piece of expert scientific evidence. The judge’s decision was manipulated to admit the scientific evidence, as well as the quality of the evidence and the credibility of the expert. Participants were found to be less critical of and more persuaded by expert evidence when it was presented within a trial, compared with the same evidence presented outside of a courtroom context. These findings suggest that, when judges allow expert testimony to reach the jury although the evidence is of low quality, they imbue it with undeserved credibility. Furthermore, no changes in participants’ perceptions of the evidence were found if the mock jurors were explicitly informed that the judge had evaluated the evidence, suggesting that the participants assumed that judges normally review evidence before allowing it to reach the jury. In addition, implications for basic research are discussed, as the moderating effects of a gatekeeper have not previously been considered by established models of persuasion.


Technorati Tags: , , , , , , , , , , , , ,


CHC theory and IQ testing in Atkins decisions: Recognition by intelligence scholars


By now regular readers of this blog know I have a significant concerns re: the first prong of mental retardation determination (intelligence) in Atkins cases being determined, over and over, primarily on the basis of scores from the WAIS-R/III/IV.  I repeatedly mention the need for Atkins intelligence experts to base their intelligence testing and testimony on state-of-the-art psychometric theories of intelligence.  In particular, I frequently reference the need for the Cattell-Horn-Carroll (CHC) theory of cognitive abilities to be used as the organizational framework for intellectual test data and the determination of intellectual functioning.  I typically provide links to two sources (one a pre-pub version of a book chapter that was eventually published; the other an invited 2009 editorial in the journal Intelligence). 

If readers take time to read these sources, they will learn that CHC theory is the combination of Cattell-Horn Gf-Gc theory and Carroll's three-stratum Gf-Gc theory [Carroll, J. B. (1993). Human cognitive abilities:  A survey of factor analytic studies. New York: Cambridge University Press].  I cannot stress enough the importance of the development of CHC theory for evidence-based intelligence theories and test development and interpretation.

To add external credibility to my professional opinion, I suggest skeptical readers read the words of major intelligence scholars as they rendered judgment on the Carroll portion of the CHC model.  Below are a few select quotes.  The conclusion should be obvious. Top notch intelligence scholars recognize the seminal work of Carroll, which is a major cornerstone of CHC theory.  I'll let the words of these giants speak for themselves.

Richard Snow (1993; back cover jacket of Carroll's, 1993 book):
 “John Carroll has done a magnificent thing. He has reviewed and reanalyzed the world’s literature on individual differences in cognitive abilities…no one else could have done it… it defines the taxonomy of cognitive differential psychology for many years to come.”

Burns, R. B. (1994). Surveying the cognitive terrain. Educational Researcher, 35-37.
Carroll’s book “is simply the finest work of research and scholarship I have read and is destined to be the classic study and reference work on human abilities for decades to come” (p. 35).

Horn, J. (1998). A basis for research on age differences in cognitive abilities. In J.J. McArdle, & R.W. Woodcock (Eds.), Human Cognitive Abilities in Theory and Practice (pp. 57-92). Mahwah, NJ: Lawrence Erlbaum.
A “tour de force summary and integration” that is the “definitive foundation for current theory” (p. 58).  Horn compared Carroll’s summary to “Mendelyev’s first presentation of a periodic table of elements in chemistry” (p. 58). 
Jensen, A. R. (2004). Obituary - John Bissell Carroll. Intelligence, 32(1), 1-5.
…on my first reading this tome, in 1993, I was reminded of the conductor Hans von Bülow’s exclamation on first reading the full orchestral score of Wagner’s Die Meistersinger, ‘‘It’s impossible, but there it is!’’

“Carroll’s magnum opus thus distills and synthesizes the results of a century of factor analyses of mental tests. It is virtually the grand finale of the era of psychometric description and taxonomy of human cognitive abilities. It is unlikely that his monumental feat will ever be attempted again by anyone, or that it could be much improved on. It will long be the key reference point and a solid foundation for the explanatory era of differential psychology that we now see burgeoning in genetics and the brain sciences” (p. 5).


Technorati Tags: , , , , , , , , , , , , , , , , , , , , , , ,


FYI: 2/3 of US population support death penalty: New Gallup poll results



FYI post.  No comment.  More information available at Gallup.

Technorati Tags: , , , , , ,

Thanks to In the News for the mention



Thanks to In the News for the mention re: one of my recent blog posts.  In the News is a great source of information regarding forensic psychology, criminology, and psychology-law.

Technorati Tags: , , , , , , , , , , , ,

AP101 Brief #1b: g or not to g in Atkins MR death penalty cases (part b in series)



Applied Psychometrics (AP) 101 Brief #1b:  g or not to g in Atkins MR death penalty cases (second in a series)

If you have not read the first  post in this series, you should read the first post now.  Then return and resume reading.

As described in the first post, the g-loadings (of the tests or composite scores in an IQ battery) on the first principal-component in principal component analysis (PCA) is a traditional index of the g-ness (aka., saturation of general intellectual ability) of a measure.  Furthermore, g-loadings calculated within a specific intelligence battery only tell you the relative g-ness of measures as defined by that specific collection of measures within that particular IQ battery.  It is when one moves to joint-analysis of intelligence batteries (and the more batteries in the analysis the better) that a more accurate picture of a measures g-ness can be determined.

Having access to the mixed LD/normal university young adult sample reported in the WJ III technical manual (conflict of interest note - I'm a coauthor of the WJ III), I just ran a joint PCA on on this sample of 200.  A description of the sample and instruments administered, abstracted from the WJ III technical manual (click here for a brief WJ III technical manual bulletin summary) can be found by clicking here.   I selected this data set as the adult subjects had all been administered the WAIS-III.  In addition, they had also been administered the CHC-based WJ III Tests of Cognitive Ability (WJ III) and the Gf-Gc based Kaufman Adolescent and Adult Intelligence test (KAIT).  I analyzed the composite scores from the three batteries via the PCA procedures conceptually described in the first post in this series.  Below is a summary of the results.


Interestingly, the top four g-measures are from the KAIT (Fluid and Crystallized composites) and the WJ III (Gc=Comprehension-Knowledge; Gf=Fluid Reasoning).  Even more interesting was the finding that when a more broad array of cognitive ability composites are included in a g-analysis, the WAIS-III VC (Verbal Comprehension Index; also classified as a strong measure of Gc as per CHC theory) is only the 9th strongest g-measure, and even falls behind the WAIS-III Perceptual Organization (PO; primarily Gv and some Gf as per CHC theory) and Working Memory Indexes (WM--Gsm as per CHC theory). 

A hypothesis that has been advanced to explain the differences in how Gc/verbal abilities are measured by the Wechsler verbal scales and Gc/verbal abilities on other cognitive batteries is grounded in Jim Cummins distinction between two types of language proficiency---BICS (Basic Interpersonal Communication Skills) and CALPS (Cognitive Academic Language Proficiency; click here for additional on-line information).  Briefly, BICS is language proficiency in more contextualized everyday language contexts while CALP is the more context-reduced conceptual-linguistic knowledge that occures in a context of semantics, and abstractions.  CALP is more cognitively demanding.  Anyone familiar with the Wechsler verbal subtests of Vocabulary, Comprehension and Similarities knows that they allow for subjects to provide lengthy verbal responses in their own everyday language.  In contrast, verbal items on other IQ tests (e.g., WJ III), require one-word responses and tend to focus more on cognitive processing involving language (e.g., antonyms, synonyms, verbal analogies).  It has been hypothesized that the Wechsler verbal tests and scales are more BICS-influenced while other IQ tests tend to use verbal/Gc test formats that require more CALP.  This might explain the findings reported above--that the WAIS III Verbal Comprehension Index is less cognitively demanding that the Gc/Verbal scales from the WJ III and KAIT.

Why is this finding important?  Because in a number of Atkins decisions, where concerns were raised about the person's WAIS-III Full Scale score being a good estimate of the persons g-ness (general intelligence), considerable stock was placed in the Verbal IQ (of which the Verbal Comprehension Index is now a purer factor measure) as being the best estimate of the person's general intelligence (see posts re: Maldonado and Vidal decisions)

Lets also examine these same data through the lens of multidimensional scaling analysis (specifically, Guttman's Radex Model).  The radex model statistically classifies measures as per two dimensions--cognitive complexity and stimulus content.  As noted in the prior post in this series, cognitively complexity is considered an index of general intelligence (g). For interested readers, a classic article on the use of the MDS radex model in the analysis of intelligence measure was published in Intelligence in 1983 (Marshalek, Lohman & Snow, 1983). Below is a visual-spatial representation of the MDS results in the current university sample, using the same composite measures as reported above (in the PCA g-loading analysis).  The interpretations below are mine.



The two broad cognitive processing continuim interpretaions (X- and Y-axis) are not the focal point of the current discussion.  The most critical finding in the current context, as per the radex model, is how close to the center of the figure a composite measure score is placed.  Measures that are the closest to the center are considered the most cognitively complex.  As measures move further away from the center, they are judged to be less cognitively complex

The results, although using a different method than PCA, produce the same conclusions.  The most cognitively complex measure in this sample (which could be thus interpreted as the best index of cognitive complexity or g-ness) is the KAIT Fluid Intelligence Scale.  The next closest to the center of the figure are the WJ III Gc (Comprehension-Knowledge) and Gf (Fluid Reasoning) clusters.  Again of interest is the location of the WAIS VC composite...it is much less cognitively complex than these other measures, and interestingly, much less cognitively complex than the two similar measures of Gc abilities (WJ III Gc composite; KAIT Crystallized Intelligence composite).  I've also provided my stimulus content hypothesis interpretations of groupings of composties (designated by ovals) from across the batteries (e.g., Processing Speed- WAIS Processing Speed and WJ III Gs or Processing Speed).

The findings in this one sample (which therefore warrants caution in generalization), suggests that the WAIS-III Verbal Comprehension composite, which is the most valid measure of Gc or verbal abilities on the WAIS-III, may NOT be all it is thought to be--when it comes to tapping cognitively complex cognitive processing.  Other Gc or verbal measures from intelligence batteries with adult norms (KAIT; WJ III) were found, in a relative sense, to be much better indicators of a person's g-ness (general intelligence).  The data suggest that, in this sample, even the WAIS-III Working Memory Index score may be a better relative proxy for g-ness than the WAIS verbal composite.

These analyses raise interesting questions about Atkins decisions that have relied either exclusively on the WAIS-R/WAIS-III scores, particularly when part scores (Verbal IQ, Verbal Comprehension; Peformance IQ; Perceptual Organization; etc.) are used instead of the Full-Scale IQ to determine mental retardation, or when the respective WAIS verbal composite is considered better than other potential test global IQ scores (or similar Gc/verbalcomposite scores from other batteries) in making a determination of level of general intellectual functioning.

How can this be?  How can a major scale from the "gold standard" of IQ tests (as it is commonly called in Atkins decisions) be a poorer estimate of general intelligence (g-ness) than most psychologists think?  More importantly, what are the implications for Atkins decisions, when the WAIS-R/III Full Scale score has been questioned as an accurate g-estimate in the face of considerable profile variability, and then the Verbal IQ/Verbal Comprehension Index is used to estimate g-ness (general intelligence)?

In both the Maldonado and Vidal decisions considerable stock was placed on the respective Wechsler verbal composite scores as being the best indicator of general intelligence (for making a determination of mental retardation).  In Maldonado, the reliance of the Wechsler verbal composite trumped a more comprehensive CHC-based IQ battery (BAT-R) administered in Maldonado's native language (Spanish).  Of concern in the Vidal decision, is that he had been administered four different versions of the Wechsler IQ batteries (over many decades), and they consistently revealed a large verbal/nonverbal (performance) IQ split.  Thus, arguments hinged extensively on the Verbal IQ vs the Full Scale IQ.  I'm perplexed why the experts in intelligence and intelligence testing, even without knowing the result of the above g-analysis, did not say "we've got consistent Wechsler V-P split information, I think it would be important to administer more contemporary intelligence tests, or parts of some of these batteries, to find out more information about important g-related cognitive abilities (for the defendant) not measured by the WAIS battery."  The Wechsler batteries had consistently captured Vidals abilities as measured per that battery--wouldn't time have been better spent, and a decision made on a higher quality array of cognitive information, by requesting administration of other IQ tests (or parts of other IQ tests) instead of arguing over old and consistent limited cognitive data? In fact, I, together with Flanagan and Ortiz, published a book in 2000 (The Wechsler Intelligence Scales and Gf-Gc theory:  A contemporary approach to interpretation) that presents procedures for augmenting the various Wechsler batteries to provide for a more comprehensive CHC/Gf-Gc based assessment of a persons intellectual functioning.  This information was also available as early as 1998 (see ITDR by McGrew and Flanagan).  [conflict of interest note - I coauthored these two books which made little in the way of ching-$ for the authors.  They are now both not being printed and none of the authors are receiving any royalties from their sales].

I continue to be baffled/troubled by the over-reliance, and almost god-like stature of the various versions of the WAIS (R/III/IV), in Atkins rulings.  It has been well known (and written about in articles and books; click here) since the early 1990's, that contemporary CHC (aka, Gf-Gc) theory had emerged as the consensus model of intelligence and, more importantly, instruments had been designed (with adult norms) to measure many of the unmeasured or poorly measured CHC abilities not taped by the WAIS-R/III batteries.  If I was an attorney arguing an Atkins case, on either side of the fence, I would seek intellectual testing beyond the so-called "gold standard."  I would want the best possible estimate of g-ness (since this seems to be the crux of the first prong of MR determination in most Atkins cases).

IMHO, the major problem is that of the "inertia of tradition" in intelligence testing, particularly in psychology disciplines that deal with adult populations.  Many practicing psychologists, esp. those working in adult settings whose professional associations and journals have paid less attention to contemporary intelligence theory and test development (less than school and educational psychologists), simply have not kept abreast of these developments. 

How long will Atkins expert intelligence testimony, expert debates, and decisions be made in the face of the Atkins MR IQ Theory-Test gap?  Isn't this simply wrong?  Professionally and ethically shouldn't psychologists who offer judgements in life-or-death decisions hinging on IQ test results be "up to speed" regarding contemporary intelligence theory and instruments?  Should the courts continue to handicapped by the presentation of intelligence test results that are not based on the best evidence from intelligence theory, research, and test development?  The courts are at the mercy of experts who testify, experts who I believe need to be familiar with the cutting edge empirical and theoretical information on the structure of human intelligence and various IQ batteries that are available, beyond the Wechslers.

Given the data presented above, it is possible that the decisions in at least two cases (and I'm sure there are more), may have had a different outcome, or at least an outcome based on a more comprehensive set of intelligence information.  Justice could have been better served via more contemporary intellectual testing practice and interpretation.

I continue to be troubled by this issue.....I need to stop writing and reflect...and will post more on it in the future.

Stay tuned...this series may continue as I analyze other data sets.

Technorati Tags: , , , , , , , , , , , , , , , , ,



Wednesday, October 14, 2009

Calling all Atkins MR death penalty Amicus Briefs


I just added another Amicus Brief, this one filed by AAIDD in the case of Briseno v Quarterman, to the Amicus Brief section of this blog.  I've also added specificity to the labeling of the other briefs by designating  who filed the brief.

If readers are aware of other Amicus Briefs that have been filed in Atkins MR death penalty cases, please forward to me.  In addition to building a library of Atkins MR death penalty decisions, I'd similarly like to build an accessible list of Amicus Briefs filed in such cases.

Thank you.

Technorati Tags: , , , , , , , , ,


Tuesday, October 13, 2009

APA Div. 41, 33 and Law and Human Behavior journal



This weekend I joined Division 41 (American Psychology-Law Society) of the American Psychological Association.   I will be monitoring publications and activities related to the Intellectual Competence and Death Penalty blog.  As a result I've add the divisions journal, Law and Human Behavior, to the list of professional journals monitored by this blog (see listings on right-side of blog).

I've also mailed my application and dues to Division 33 (Intellectual and Developmental Disabilities) and will make a similar post once I receive my membership notification.

Technorati Tags: , , , , , , , , , , , ,

Monday, October 12, 2009

More voodoo psychometrics in Atkins MR death penalty cases? This time adaptive behavior

Voodoo psychometrics strikes again!

A few days ago I made a post (and a subsequent brief FYI follow-up post)  regarding a number of theoretical and psychometric issues that surfaced in the Maldonado (2009) Atkins MR death penalty court decision (click here and here).  It was my opinion that a number of questionable arguments and decisions had been made re: the entirety of the psychometric intelligence test data available in the Maldonado (2009) decision. I won't repeat them here.

Probably my biggest criticism was the use of a non-empirical, unvalidated n=1 psychologist clinical procedure to upwardly adjust IQ scores based on educational and cultural background variables.  Today I learned  that similar "social-cultural" upward adjustment of adaptive behavior scale scores, which are the foundation of the second prong of the determination of mental retardation (or not) in Atkins cases, has also occurred in a number of Atkins cases.  This time I was able to locate an article by Denkowski and Denkowski (2008) that outlined the logic and reasoning for the recommended "systematic" procedure.  I read it in psychometric disbelief!

There is no reason for me to outline my psychometric criticisms, as a number of authors replied with most of the arguments I would have made.  These response, in the same journal, are by Widaman & Siperstein (2009) and Olley (2009) (click here for post re: AB-related chapter by Olley &Cox, 2008). Denkowski and Denkowski (2009) then reply.

As an applied psychometrician, I concur with most of the reactions and arguements of Widaman, Siperstein and Olley.  There is simply insufficent scientifc and psychometric grounds for the upward adjustment of AB scores as outlined.  Yes, psychologists are trained to use clinical adjustment when interpreting test scores, and I so invoked such judgement when conducting countless intellectual assessments during my years as a practicing school psychologist, but clinical judgement is not the same as the development of special score adjustments of nationally standardized psychological instruments based primarily on logic and reason.

Technorati Tags: , , , , , , , , , , , ,


Thursday, October 8, 2009

AP101 Brief #1a: g or not to g in Atkins MR death penalty cases



Applied Psychometrics (AP) 101 Brief #1a:  g or not to g in Atkins MR death penalty cases (first in a series)

Despite whether one believes that general intelligence (g) exists, or not (e.g., John Horn), and ignoring the search for the essence of g (via elementary cognitive tasks measuring reaction time, temporal processing, etc.) at the level of brain mechanisms (e.g., Jensen's neural efficiency hypothesis), it is clear from a reading of most Atkins IQ MR death penalty cases that psychological experts testifying in these cases [primarily because of the emphasis on a "deficit in general intellectual functioning" as the first prong in MR diagnosis in the courts, as per recognized professional association definitions of mental retardation; APA, AAIDD] often argue for different IQ scores as being more accurate estimates of the persons g-ness (IQ) than others.

For example, both in Davis (2009), and especially in Vidal (2007), major arguments focused on whether the Full Scale IQ score from theWAIS-III/IV was the best index of g-ness (and thus mental retardation or mental capacity), or whether one of the part scores (e.g., Verbal IQ, Performance IQ) should be used as the best estimate of the persons g-ness (due to extreme variability in the part scores). My "g-estimate is better than your g-estimate" appears a fundamental point of contention at the core of many Atkins cases,  given the assumption that mental retardation is a global deficit in intelligence (see guest post by Watson for some alternative thoughts and excellent insights on the global vs modular nature of intelligence),

Then, along comes Maldonado (2009) where the g-ness argument, at one juncture, is based on the belief that the Spanish WAIS-III Verbal IQ, which is best interpreted as a CHC measure of crystallized intelligence (Gc), should take precedence over the BAT-R total composite score that is comprised of Gc and six other broad CHC abilities.

"My g-estimate....your g-estimate......this special "nonverbal" g-estimate is more accurate for this individual....that is not a good g-estimate....etc......" back-and-forth arguments beg for empirical scrutiny.  So....buckle up and lets examine some real data.......in search of g-ness.  This is the introduction to a small series of posts that will eventually examine, with empirical data, the relative g-ness of the "gold standard" (WAIS-III/IV) composite scores that are most often debated in these matters.

But first a definition and some methodological background information.  According to the APA Dictionary of Psychology,  general intelligence (the general factor) is:
  • a hypothetical source of individual differences in GENERAL ABILITY (emphasis in original) , which represents individuals' abilities to perceive relationships and to derive conclusions from them.  The general factor is said to be a basic ability that underlies the performance of different varieties of intellectual tasks, in contrast to SPECIFIC ABILITIES (emphasis in original), which are alleged each to be unique to a single task (p. 403).
[Note - some of the the text below comes from Flanagan, McGrew & Oritz (2000).  The Wechsler Intelligence Scales and Gf-Gc theory.  Boston:  Allyn & Bacon.

Intelligence tests have been interpreted often as reflecting a general mental ability referred to as g (Anastasi & Urbina, 1997; Bracken & Fagan, 1990; Carroll, 1993a; French & Hale, 1990; Horn, 1988; Jensen, 1984, 1998; Kaufman, 1979, 1994; Keith, 1997; Sattler, 1992; Sattler & Ryan, 1999; Thorndike & Lohman, 1990).  The g concept was associated originally with Spearman (1904, 1927) and is considered to represent an underlying general intellectual ability (viz., the apprehension of experience and the eduction of relations) that is the basis for most intelligent behavior. The g concept has been one of the more controversial topics in psychology for decades (French & Hale, 1990; Jensen, 1992, 1998; Kamphaus, 1993; McDermott, Fantuzzo, & Glutting, 1990; McGrew, Flanagan, Keith, & Vanderwood, 1997; Roid & Gyurke, 1991; Zachary, 1990).

According to Arend et al., (2003),  Jensen (1998a, 1998b) proposed that cognitive complexity  might represent a fundamental aspect of g an could be quantified based on inspection of the test measures loadings on the first unrotated factor, because complex tasks show higher factor loadings than simple tasks on that factor.  In many respects when psychologists are discussing mental retardation and general intelligence, there is an implicit assumption that low general intelligence (e.g., mental retardation) is reflected most clearly on performance on the most cognitively complex measures (i.e., high g measures). 

As with the controversy surrounding the nature and meaning of g, disagreements exist about how best to calculate and report psychometric g estimates.  Most all methods are based on some variant of principal component, principal factor, hierarchical factor, or confirmatory factor analysis (Jensen, 1998; Jensen & Weng, 1994).  Although a hierarchical analysis is generally preferred (see Jensen, 1998, p. 86), as long as the number of tests factored is relatively large, the tests have good reliability, a broad range of abilities is represented by the tests, and the sample is heterogeneous, (preferably a large random sample of the general population), the psychometric g's produced by the different methods are typically very similar (Jensen, 1998; Jensen & Weng, 1994).  For the interested reader, Jensen’s (1998) treatise on g (The g Factor) is suggested, as it represents the most comprehensive and contemporary integration of the g related theoretical and research literature.

Operationally the determination of high, moderate or low g-ness of tests or composites has typically been based on each measures correlation (aka., factor or principal component loading) with a single common factor, component, or dimension extracted from the correlations among the set of measures in question.  Measures that "load" high on the g-factor are considered to be the better estimates of general intelligence.

Consider the following simple analogy (which is not original...I borrowed the conceptual idea from Cohen et al., 2006).  You have a special pole that posses a special form of  magnetism (general intelligence). You throw a bunch of  metal marbles (which are the test measures), which have different degrees of the same magnetic force, into a box with the pole at the center.  You gently shake the box.  When you open the box, there is one "king" marble at the top of the poll (it has the highest degree of shared magnetism with the strongest part of the pole), followed next by the next strongest....and so on until the metal marble with the least amount of shared magnetic force is at the bottom.  The pole represents g (general intelligence) and the ordering of the metal marbles (the test measures) represents the ordering of the g-ness (degree of shared magnetic force) of the measures.  The "king" test/marble is assigned the highest numerical index, with each succeeding (and lower) test/marble assigned a slightly lower numerical index of g-ness (shared magnetism).

This is what principal component analysis conceptually accomplishes with a collection of IQ test measures.  It statistically orders the various psychometric measures from strong g-loading to low-g-loading.  This is the typical and traditional statistical currency used by psychometericians and psychologists when discussing the degree of g-ness or g-saturation of different measures--those measures most important for establishing an estimate of a person's general intelligence.


The problem with within-battery factor analysis is that it can affect the g-estimates.  For example, a test’s loading [note- g-loadings are most often computed for the individual subests in a test battery, and not the composite scores such as Verbal IQ, processing speed, etc.-- it is the later, the g-ness of composite scores, which appears to be a critical issue in many Atkins cases.  Thus, when reading the this text I will refer to the measures g...which could mean test or composite] on the general intelligence (g) factor will depend on the specific mixture of measures used in the analysis (Gustafsson & Undheim, 1996; Jensen, 1998; Jensen & Weng, 1994; McGrew, Untiedt, & Flanagan, 1996; Woodcock, 1990).  If a single vocabulary measure is combined with nine visual processing measure, the vocabulary measure will most likely display a relatively low g loading because the general factor will be defined primarily by the visual processing measures.  In contrast, if the vocabulary measure is included in a battery of measure that is an even mixture of verbal and visual processing measures, the loading of the vocabulary measure on the general factor will probably be higher.  It is important to understand that measures g loadings, as typically reported, only reflect each measures relation to the general factor within a specific intelligence battery.  Although in many situations a measure g loading will not change dramatically when computed in the context of a different collection of diverse cognitive tests (Jensen, 1998; Jensen & Weng, 1994), this will not always be the case.

Within (internal-validity) vs across (joint; external validity) estimation of test measures g-ness

When measures from different batteries are combined in the joint-battery approach, the battery-bound g  estimates for some measures may be altered significantly.   Flanagan et al. (2000) demonstrated these when they calculated within- and joint-battery g estimates for the WISC-III.  These estimates were derived from a sample of 150 subjects who were administered the WISC-III and WJ III cognitive measures as part of the Phelps validity study reported for the WJ III cognitive technical manual.  Within-battery g estimates were calculated with the WISC-III data based on the first unrotated principal component.  Next the joint-battery factor analysis allowed for an examination of the WISC-III g estimates when calculated together with another intelligencet test battery (WJ III), one that included a broader array of CHC abilitiy measures.

Flanagan et al. (2000) reported that the within- and joint-battery WISC-III g loadings were similar for many of the individual measures.  For example, the within- and joint-battery test g loadings are generally similar (i.e., do not differ by more than .05) for the Similarities (.76 vs .71), Vocabulary (.78 vs .74), Digit Span (.48 vs .49), Block Design (.60 vs .61), Object Assembly (.50 vs .45), and Symbol Search (.57 vs .54) measures.  These six WISC-III measures appear to have similar g characteristics when examined from the perspective of either the WISC-III or CHC (WJ III battery) frameworks.  However, the joint-battery g loadings were noticeably lower than the within-battery g loadings (i.e., lower by .06 or more) for Information  (.77 vs .68), Arithmetic (.70 vs .64), Comprehension (.59 vs .51), Picture Completion (.50 vs .40), Picture Arrangement (.37 vs .31), and Coding (.46 vs .37).   The results suggested that the latter WISC-III measures were relatively weaker g indicators than is suggested by within-battery WISC-III g analysis.

This example demonstrates the potential chameleon nature of test measures g estimates that are calculated within the confines of individual intelligence batteries when compared to those calculated within a comprehensive set of ability measures. 

And, yet to be mentioned is another, older, and for some reasons under-utilized statistical method for examing the g-ness (congitive complexity) of IQ test measures...multidmensional scaling (MDS).  We will save that for the next post in this seires.

To be continued........................

.

Maldonado (2009) miscarriage of psychometric justice PS

No sooner had I made my Maldonado (2009) Atkins IQ MR death penalty miscarriage of psychometric justice post and someone sends me a link to a post at the StandDown Texas Project regarding a case where the clinical adjusting of IQ scores, as described and challenged in my prior post, was questioned.....and even referred to as "junk science."

Technorati Tags: , , , , , , , , ,


Maldonado (2009) IQ MR Atkins death penalty decision: A psychometric miscarriage of justice?

I'm no longer shocked by what passes as credible psychological/psychometric evidence or testimony in some Atkins IQ MR death penalty court decisions.  Another such decision (Maldonado, 2009) has come to my attention.  A link to a PDF copy of the decision is now available in the Court Decisions section of this blog. 

Readers can gather all the relevant background information re: the case by reading the entire decision.  I intend to focus primarily on the psychological interpretation and psychometric issues in the decision that are troubling. 

But first, before delving into these issues, I  present a few quotes (from the Maldonado record) that capture the confusion and uncertainty surrounding many Atkins decisions, a situation that is resulting in considerable variability in the quality of psychological assessment data reported/interpreted and the common occurrence of "dueling expert witnesses".

Quotes from the record:
Because the Supreme Court did not establish a bright-line test to identify mental retardation, the Atkins inquiry has become a fact-intensive question that heavily relies on the opinions provided by mental-health experts. This case, like most involving Atkins claims, requires consideration of testimony from competing experts who disagree about the nature of mental retardation, the means by which it may be identified, the manner in which it manifests in a criminal defendant’s life, and the psychological profession’s role in making the legal decision of whether mental capacity precludes execution. (p.32)

A “welter of uncertainty” followed the Atkins decision because “[t]he Supreme Court neither conclusively defined mental retardation nor provided guidance on how its ruling should be applied to prisoners already convicted of capital murder.” Bell v. Cockrell, 310 F.3d 330, 332 (5th Cir. 2002). Accordingly, federal courts have approached the implementation of Atkins with some trepidation. (p.36)
After reading a number of of Atkins rulings (see Court Decisions section of this blog), I could not agree more with these statements.  The courts appear ill-equipped to handle the complex psychological measurement issues presented, issues that are, at times, confounded by the inclusion of data from dubious procedures, interpretations of test scores that are not grounded in any solid empirical research, and the deference to a single intelligence battery (the WAIS series) as the "gold standard" when a more appropriate instrument (or combination of WAIS-III/IV and other measures) might have been administered, but the results of the more appropriate measure are summarily dismissed based on personal opinion (and not sound theory or empirical research).


Below are some of the troubling psychological/psychometric issues I see in the Maldonado (2009) decision

Lack of English language proficiency:
  "A theme developed by both parties is that Maldonado’s lack of English proficiency has madehis exact intellectual capacity difficult to gauge." (p.38).
  • As a result, the prosecution psychologist administered the WAIS-III "through a translator" (p.42).  The translator administered WAIS-III resulted in Verbal, Performance, and Full Scale scores of 74, 74, and 72 respectively.  A defense psychologist captures the essence of my reaction to the use of a translator-administered WAIS-III.  “The accepted practice in the evaluation on Spanish-speakers is to communicate with the client in Spanish without the use of translators. In addition, tests should be scientifically translated and validated and the most appropriate norms available should be applied" (p.49).  I agree.  I am unaware of any professionally established and endorsed procedure for the translated administration of the English-normed WAIS-III.  Such a procedure violates a fundamental backbone of the science of individual intelligence testing--standardized test administration.
Use of non-empirical clinical judgement procedures to upwardly adjust IQ scores.  On page 53 of the record, it is indicated that the prosecutions psychological expert believed that the translated English-normed WAIS-III scores needed to be upwardly adjusted due to Maldonado's educational and cultural background.  Additional support came from the finding of a Verbal IQ score of 83 on the WAIS Espanol (p.56).  Although psychologists are appropriately trained to recognize the potential impact of such environmental variables when interpreting scores, the psychologist  upwardly adjusted the scores to a specific IQ score estimate ("It’s around the 80s, I guess, if you had to pin me down. Around the 80s; somewhere in there"- p.48) and this expert "conceded that only 'clinical judgment,' not any statistical formula or established methodology, informed how much to alter an IQ score because of cultural and educational factors" (p. 53). 

My concern with this procedure mirrors the testimony of the defense experts in the case.  Adjusting obtained IQ scores, either up or down, based on an n=1 professionals clinical judgement, in the absence of any scientifically established procedure for adjusting IQ scores, is troubling and is not consistent with accepted psychological assessment practices or standards.  In fact, this IQ adjustmend procedure sounds similar to a notable empirical effort (in the late 1970s and early 1980s) to produce IQ scores that better reflected a persons social-cultural backgroundJane Mercer's SOMPA (System of Multicultural Pluralistic Assessment) was a valiant effort to adjust Wechsler IQ scores for African-American individuals based on their social-cultural knowledge and history.  The result was a new score called Estimated Learning Potential (ELP).  From the start, SOMPA was controversial and eventually was found to be flawed for many reasons (see Hellfinger, 1987; also Jirsa, 1983).  If a reasonably conceived theoretical and empirical IQ adjustment procedure (i.e, SOMPA), which was intended to account for a person's social-cultural background, was found to be flawed, how can a specific n=1 psychologist be endowed with unique insights that allow for the invoking of an unspecified personal algrorithm to make IQ score adjustments?  This is indeed troubling.  Also, to the best of my knowledge, SOMPA is no longer around and is not, or is seldom, used.  Race-based or adjusted norms have not been recognized as an acceptable professional psychological assessment practice for over twenty years!

Dismissal of the BAT-R intelligence results.  Maldonado had also been administered the Woodcock-Muñoz Bateria-R (“Bateria-R”), the Spanish-language counerpart of the Woodcock-Johnson Test of Cognitive Abilities--Revised [conflict of interest notice:  I am not a co-author of the BAT-R  or WJ-R, but was a paid measurement consultant on the WJ-R project and have since become a coauthor of the subsequent edition, the WJ III].  Maldonado obtained a Broad Cognitive Ability (BCA) score, which is analagous to the full-scale score from other intelligence batteries, of 61.  Of all the cognitive tests administered, this is the only comprehensive intelligence battery that was administered in Maldonado's natural language and where his performance is compared against appropriate US-equated Spanish norms [note-- the WAIS Español administered was also administered in his natural language and makes comparisons against Spanish norms, but only a portion, the Verbal section, was administered].  Also, as previously noted at this blog, the WJ-R/BAT-R and WJ III/BAT III provide for the most comprehensive assessment of intellectual functioning as per the consensus model of human intelligence (CHC theory) among serious intelligence scholars.  This is what I've termed the "Atkins MR death penalty IQ test-theory gap.

Why were the BAT-R scores dismissed? 

The prosecution expert "testified that the AAMR has not cited the Bateria-R as a predicate test to the evaluation of mental retardation and “[i]t’s not well suited for that purpose, although you can use it for that" (p. 70).  He also stated that the "Bateria-R test score was especially suspect because it was inconsistent with Dr. _____'s administration of the WAIS Español in which Maldonado scored well above the range for mental retardation." (p.70). Furthermore, this expert "opined that the usefulness of the test was impaired because it 'is used by school psychologists to diagnose learning disabilities and it measures a lot of things like visual and auditory processing. It really measures very little in terms of general intelligence.' " (p.70).  As a result, "the state habeas court dismissed the Bateria-R score because it “is not one of the tests the AAMR cites for mental retardation evaluation,” but instead “is generally used by school psychologists to diagnose learning disabilities” and, in fact, “is not very relevant for establishing general intellectual functioning, so it is not well-suited for determination of the first prong . . . to determine mental retardation" (p.70-71).

There are many problems with the reasons given for dismissing the most culturally appropriate (for Maldonado) and comprehensive measure of intelligence (BAT-R). 

First, dismissing an instrument because it is used primarily by a particular set of psychologists (school psychologists) is non-sensical, and frankly, condescending.  School psychologists typically give many more intelligence tests than psychologists working in adults settings and use these instruments to diagnose mental retardation. School psychologists, in many respects, have more intimate familiarity with intelligence testing than most other professional psychologists.  

Second, the previously mentioned national expert panel that examined the Dx of MR at the same cut-point as Atkin's cases (for SSA benefits) indicated that most intelligence tests would be moving towards measuring the Cattell-Horn and Carroll Gf-Gc models of intelligence (now collectively referred to as CHC theory; also see McGrew, 2009), and instruments based on this model are very relevant to the Dx of MR.  The English version of the WJ-R WJ III (a CHC-based revision of the WJ-R) was listed as one of the approved instruments by the expert panel...which should also implicitly argue for use of the Spanish-language versions.  The prosecution witness, who appears  stuck in the land where the WAIS is the "gold standard," appeared unaware of the recent advancements in understanding the psychometric nature of human intelligence, which has converged on the CHC model of intelligence.  This is particularly ironic given that both John Horn (of Cattell-Horn) and Jack Carroll served as theoretical consultants on the WJ-R...which was the foundation of the BAT-R.

Third, stating that the BAT-R (and WJ-R by implication) is not a respected measure of general intellectual functioning reflects a complete lack of awareness of the CHC-foundation of the instruments, as well as published research (including the WJ-R and BAT-R technical manuals and bulletins).  If CHC theory is the consensus model of psychometric intelligence, then the only battery administered to Maldonado that measured most of the model, which in most conceptualizations has general intelligence (g) at the apex, should have been given serious weight.  I, and others, heard Dr. Arthur Jensen, the most prominent pscyhometric expert on g, at an ISIR conference in Nashville, TN, state, during a discussion of a presentation in front of the entire audience, that (at the time) he considerd the WJ-R (and, thus, the BAT-R by implicit endorsement) the best available intelligence battery for measuring g.  I will return to this point in future posts as I've been analyzing WAIS-III data together with the WJ batteries (as well as other accepted intelligence batteries, e.g., KAIT, K-ABC; SB-IV; SB-IV) to evalute the g-ness of each battery when jointly analyzed.

Fourth, prosecution psychologist used the WAIS Español Verbal score as evidence that the BAT-R was not accurate.  The problem with this logic is that the WAIS Verbal scale is known to be an excellent measure of crystallized intelligence/comprhension-knowledge (Gc), only one of the major 7-8 domains in the CHC model of intelligence.  Conversely, the BAT-R includes indicators from seven of the major CHC broad ability domains, only one of which is Gc.  No attempt was made to compare the WAIS Gc score with the corresponding Gc measure(s) on the BAT-R.  They might have been similar...or not.  More importantly, you cannot compare a measure of Gc to one that includes Gc and fluid reasoning (Gf), visual-spatial processing (Gv), short-term memory (Gsm), long-term retrieval (Glr), auditory processing (Ga), and processing speed (Gs).

I could go on and comment on other issues, such as the use of only the Verbal scale of the Spanish WAIS-III, as well as the whole area of adaptive behavior, but these are issues for another post, possibly by a guest blogger.

Technorati Tags: , , , , , , , , , , , , , , , , , , ,

Wednesday, October 7, 2009

Guest bloggers make a difference: Atkins MR IQ death penalty blog guest posts



The above graph represents new people who have visited the Intellectual Competence and Death Penalty blog over the past month.  Notice the various "spikes."  Most all of are directly related to guest blog posts by psychologists or legal professionals involved in Atkins IQ MR death penalty cases.  Today's spike (10-7-09) is directly related to the guest blog post by Dr. Dale Watson.  Thank you Dr. Watson.

If you are interested in issues related to this blog, and want an occasional  platform from which to express thoughtful, scholarly, and logical comments, and/or research synthesis, etc. regarding Atkins IQ MR death penalty cases, a guest post at this blog would be welcome...and, in turn, would provide you some good exposure.  Psychologists and legal professionals are both welcome.

If you are interested, contact me privately at:  iap@earthlink.net.

The blogmaster.




Davis (2009) Atkins decision: IQ part scores and modular nature of intelligence: Guest post by Dr. Dale Watson

A number of recent Atkins-related IQ/MR death penalty court decisions have raised important issues re: the interpretation of variability in part scores when compared to the total (full scale-general intelligence; g) composite index from an intelligence test.  In particular, Davis (2009) and Vidal (2007) are excellent examples of the complex issues...and how they relate to a diagnosis (or not) of mental retardation.

Dr. Dale Watson has studied the Davis (2009) decision and has provided the following thoughtful guest blog post (click here for other guest blog posts].  Thank you Dr. Watson for the excellent analysis and commentary.  I (the blog dictator) have reproduced his post "as is" with a few exceptions.  I've added a few URL links to other sources.  I've also added emphasis to certain statements via underlining, followed by a note that I added the emphasis.  Finally, I've yet to figure out if it is possible to add real footnote superscripts to blog text via the blog editor I use.  So, I've adopted the format of putting footnote numbers in brackets [ ] and placing the footnotes at the end of the blog post.

I encourage other psychologists and measurement specialists to review some of the Atkins court decisions posted at this blog and send me their analysis and thoughts for potential blog posts. 


By Dr. Dale Watson

In a recent Maryland case, U.S. v. Davis, (2009 WL 1117401 (D.Md.)) the trial court found the defendant to be mentally retarded and thus ineligible for the death penalty in line with Atkins v. Virginia, 536 U.S. 304, 122 S.Ct. 2242, 153 L.Ed.2d 335 (2002).  An analysis of the IQ results in this case highlights issues regarding the impact of part-score variability in the IQ profile and the diagnosis of mental retardation as well as the modular nature of intelligence. [emphasis added by blogmaster.  Blogmaster note--Also, see prior post regarding the recommended use of total vs part scores from a national panel report ].

Earl Davis, the defendant, had been charged with federal crimes, including murder, which made him potentially eligible for the death penalty unless he were found to be intellectually and developmentally disabled.  In a pre-trial proceeding the court heard evidence regarding whether Mr. Davis was in fact mentally retarded.  The defense presented the testimony of five imminently qualified experts supporting their position.  The prosecution, relying on the opinions of two board certified neuropsychologists, alternatively argued that Mr. Davis did not meet the criteria for being mentally retarded and instead demonstrated a learning disability.  This argument was founded, in part, on the finding of a significant discrepancy between the Verbal Comprehension (VCI) and Perceptual Reasoning Indices (PRI) of the WAIS-IV.[1]   The prosecution contended that the 15-point discrepancy between these indices made the Full Scale IQ (FSIQ) unreliable as an overall measure of his intellectual functioning.[2]   A discrepancy of this magnitude (VCI = 66; PRI = 81) could be expected to occur with a base rate of 12.7 percent within his ability range and could thus be considered “abnormal.” [3]   This argument was advanced despite the fact that Mr. Davis’ WAIS-IV FSIQ of 70 nominally met the first prong of the definition of an intellectual disability (mental retardation).
 
The prosecution argument was not without precedent.  For example, it was noted in testimony that the Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition, Text Revision (DSM-IV-TR) reported:
When there is a significant scatter in the subtest scores, the profile of strengths and weaknesses, rather than the mathematically derived full-scale IQ, will more accurately reflect the person’s learning abilities.  When there is a marked discrepancy across verbal and performance scores, averaging to obtain a full-scale IQ can be misleading (U.S. v. Davis citing DSM-IV-TR, p. 42).
Fiorello et al. (2002) summarized the literature addressing this viewpoint as applied to the Wechsler scales:
Sattler (2001) cautions that the FSIQ may misrepresent a child’s cognitive functioning level if the Verbal IQ (VIQ) and Performance IQ (PIQ) are significantly different; however he indicates that no empirical evidence exists to indicate when the FSIQ should not be reported or used in eligibility decisions.  Prifitera, Weiss, and Saklofske (1998) recommend that FSIQ should not be interpreted when differences between VIQ/PIQ or Verbal Comprehension Index (VCI) and Perceptual Organization Index (POI) scores are extreme, which they define as differences found in less than 10% of the population (p. 117). [4]
The reluctance to make the diagnosis of mental retardation in the face of significant discrepancies between measures of verbal and non-verbal abilities is thus perhaps not surprising or new.  There has been a long-held view that mental retardation is marked by a “flat” cognitive profile.  When confronted with a profile that instead shows a pattern of ipsative strengths and weaknesses perhaps many psychologists would be hesitant to make the diagnosis of mental retardation.  However, recent empirical evidence does not support that position. [emphasis added by blogmaster]

It is the case, as noted in the WAIS-IV Technical and Interpretive Manual, that “The prevalence of large and unusual discrepancies between verbal and nonverbal composite scores [on the Wechsler scales] has been shown to decrease with decreasing levels of ability (internal citations omitted)” p. 102.  However, Bergeron and Floyd (2006), using the Woodcock-Johnson III Tests of Cognitive Abilities (WJ-III Cog.), demonstrated that individuals with mental retardation “will not likely display a flat cognitive profile on comprehensive assessments of CHC broad cognitive abilities—especially when measures vary widely in g loadings—regardless of the information presented in test manuals for mental retardation groups (emphasis added).” [5]  These researchers found that, in children with mental retardation, there was significantly greater intra-individual score ranges (scatter) than in average-achieving individuals.  Specifically, nearly 37% of these children obtained at least one CHC factor score within the average range.  Bergeron and Floyd concluded,
…it is likely that the increasing number of specific cognitive abilities measured by intelligence test batteries and the variation of these scores in individual profiles inadvertently muddies the waters of mental retardation diagnosis.  As a result, when faced with IQs in the range described in the diagnostic criteria for the disorder and part scores that are much higher, practitioners may believe that an individual cannot be diagnosed with mental retardation because of evidence of “intact” or “unimpaired abilities” p. 428.
…denying special education eligibility or failing to make a diagnosis of mental retardation based on significant part score variability may do these children disservice when other ecologically valid evidence of mental retardation (e.g., adaptive behavior skill deficits) indicates genuine need. P. 429.
These findings have implications for the role of g as well as domain specific disabilities in the genesis of mental retardation.  Bergeron and Floyd posited that the centrality of the impaired CHC factors, as measured by the factor’s g-loading, would determine its sensitivity to mental retardation.  In fact, the Comprehension-Knowledge and Fluid Reasoning clusters, with the highest g-loadings, were most commonly associated with the diagnosis of mental retardation, though low average or average scores on one of these clusters did not preclude the diagnosis.

In another line of research, Anderson (1998) has suggested the need to distinguish “between mental retardation as a general deficit of thinking and mental retardation that might result from the global effects of a specific deficit in a cognitive module.” [6]  In a similar vein, Frith and Happé (1998) have argued “that general impairments (e.g. low IQ) in developmental disorders need not be the result of primary damage to domain-general mechanisms.  Rather, they may be the developmental consequence of damage to very specific, even modular, mechanisms which act as gatekeepers in development” p. 270. [7]

Neuroimaging data further supports the view of the modularity of cognitive functions.  In such a view, though g may represent a relatively unitary phenomenon in the normal brain it can fractionate in the face of modular neuropathology.  Gläscher et al. (2009) used CT and MR lesion maps to localize the impairments found on the WAIS-III Index scores in focal brain-damaged patients.   These investigators “found (1) impairments in VCI (Verbal Comprehension Index) were associated with damage in left hemisphere, in particular in the left inferior frontal cortex, (2) impairments in POI (Perceptual Organization Index) were associated with damage in right parietal, occipito-parietal, and superior temporal cortex, (3) impairments in WMI (Working Memory Index) were associated with left hemispheric lesions particularly focused in superior parietal cortex, and (4) impairments in PSI (Processing Speed Index) correlated with a number of small regions distributed across both hemispheres” (p. 686). [8]  Each of these indices, with the exception of the PSI, significantly predicted a lesion in the associated region.  These results largely confirm clinical lore regarding the significance of specific impairments in the Wechsler indices.

Williams syndrome, a genetically determined developmental disorder commonly marked by mental retardation, has frequently been characterized by relatively intact language abilities and profoundly impaired visual-spatial functions.  Thompson et al. (2005), using MRI scans, identified specifically increased cortical thickness within the perisylvian and inferior temporal regions of the right hemisphere.  This pattern of regional neuropathology appears to correspond well with the differentiated pattern of cognitive deficits.

Mental retardation may thus result from both general impairments or as the “developmental consequence of damage to very specific, even modular, mechanisms.”  [emphasis added by blogmaster]. In the former case one could anticipate a relatively “flat” profile of intellectual abilities.  In the latter instance, a more extreme pattern of strengths and weaknesses might emerge and yet result in sufficient impairment of overall intellectual functioning as to meet the criteria for mental retardation.  Evaluators would do well to remember that the diagnostic criteria for mental retardation do not limit the diagnosis to a particular ipsative pattern of scores, that mental retardation can arise from diverse etiologies and that “strengths co-exist with weaknesses.” (emphasis added by blogmaster).  These cautions apply equally within a clinical context and within a capital punishment context when the diagnosis is “a matter of life or death.”
  • [1] The opinion reflected some confusion on this issue indicating that the discrepancy was between the VIQ and PIQ despite the fact that these scores are no longer included in the WAIS-IV.
  • [2] See People v. Superior Court of Tulare County (Vidal) for an example of even greater discrepancies between verbal and performance IQs in an Atkins case.
  • [3] The “abnormality” of such discrepancies is somewhat arbitrary but authors have variously set the cut-off for an unusual finding at a base-rate of either 10 or 15 percent level.
  • [4] Fiorello, C.A., Hale, J.B., McGrath, M., Ryan, K. & Quinn, S. (2002).  IQ interpretation for children with flat and variable test profiles.  Learning and Individual Differences, 13, 115-125.
  • [5] Bergeron, R. & Floyd, R.G. (2006). Broad cognitive abilities of children with mental retardation: An analysis of group and individual profiles.  American Journal on Mental Retardation, 111(6), 427.
  • [6] Anderson, M. & Miller, K.L. (1998).  Modularity, mental retardation and speed of processing.  Developmental Science, 1(2), 239-245.
  • [7] Frith, U. & Happé, F. (1998).  Why specific developmental disorders are not specific: On-line and developmental effects in autism and dyslexia.  Developmental Science, 1(2), 267-272.
  • [8] Gläscher, J., Tranel, D., Paul, L.K., Rudrauf, D., Rorden, C., Hornaday, A., Grabowski, T., Damasio, H., and Adolphs, R. (2009).  Lesion mapping of cognitive abilities linked to intelligence.  Neuron, 61, 681-691.

Technorati Tags: , , , , , , , , , , , , , , , , , , , , , , , , ,