Thursday, October 15, 2009

AP101 Brief #1b: g or not to g in Atkins MR death penalty cases (part b in series)



Applied Psychometrics (AP) 101 Brief #1b:  g or not to g in Atkins MR death penalty cases (second in a series)

If you have not read the first  post in this series, you should read the first post now.  Then return and resume reading.

As described in the first post, the g-loadings (of the tests or composite scores in an IQ battery) on the first principal-component in principal component analysis (PCA) is a traditional index of the g-ness (aka., saturation of general intellectual ability) of a measure.  Furthermore, g-loadings calculated within a specific intelligence battery only tell you the relative g-ness of measures as defined by that specific collection of measures within that particular IQ battery.  It is when one moves to joint-analysis of intelligence batteries (and the more batteries in the analysis the better) that a more accurate picture of a measures g-ness can be determined.

Having access to the mixed LD/normal university young adult sample reported in the WJ III technical manual (conflict of interest note - I'm a coauthor of the WJ III), I just ran a joint PCA on on this sample of 200.  A description of the sample and instruments administered, abstracted from the WJ III technical manual (click here for a brief WJ III technical manual bulletin summary) can be found by clicking here.   I selected this data set as the adult subjects had all been administered the WAIS-III.  In addition, they had also been administered the CHC-based WJ III Tests of Cognitive Ability (WJ III) and the Gf-Gc based Kaufman Adolescent and Adult Intelligence test (KAIT).  I analyzed the composite scores from the three batteries via the PCA procedures conceptually described in the first post in this series.  Below is a summary of the results.


Interestingly, the top four g-measures are from the KAIT (Fluid and Crystallized composites) and the WJ III (Gc=Comprehension-Knowledge; Gf=Fluid Reasoning).  Even more interesting was the finding that when a more broad array of cognitive ability composites are included in a g-analysis, the WAIS-III VC (Verbal Comprehension Index; also classified as a strong measure of Gc as per CHC theory) is only the 9th strongest g-measure, and even falls behind the WAIS-III Perceptual Organization (PO; primarily Gv and some Gf as per CHC theory) and Working Memory Indexes (WM--Gsm as per CHC theory). 

A hypothesis that has been advanced to explain the differences in how Gc/verbal abilities are measured by the Wechsler verbal scales and Gc/verbal abilities on other cognitive batteries is grounded in Jim Cummins distinction between two types of language proficiency---BICS (Basic Interpersonal Communication Skills) and CALPS (Cognitive Academic Language Proficiency; click here for additional on-line information).  Briefly, BICS is language proficiency in more contextualized everyday language contexts while CALP is the more context-reduced conceptual-linguistic knowledge that occures in a context of semantics, and abstractions.  CALP is more cognitively demanding.  Anyone familiar with the Wechsler verbal subtests of Vocabulary, Comprehension and Similarities knows that they allow for subjects to provide lengthy verbal responses in their own everyday language.  In contrast, verbal items on other IQ tests (e.g., WJ III), require one-word responses and tend to focus more on cognitive processing involving language (e.g., antonyms, synonyms, verbal analogies).  It has been hypothesized that the Wechsler verbal tests and scales are more BICS-influenced while other IQ tests tend to use verbal/Gc test formats that require more CALP.  This might explain the findings reported above--that the WAIS III Verbal Comprehension Index is less cognitively demanding that the Gc/Verbal scales from the WJ III and KAIT.

Why is this finding important?  Because in a number of Atkins decisions, where concerns were raised about the person's WAIS-III Full Scale score being a good estimate of the persons g-ness (general intelligence), considerable stock was placed in the Verbal IQ (of which the Verbal Comprehension Index is now a purer factor measure) as being the best estimate of the person's general intelligence (see posts re: Maldonado and Vidal decisions)

Lets also examine these same data through the lens of multidimensional scaling analysis (specifically, Guttman's Radex Model).  The radex model statistically classifies measures as per two dimensions--cognitive complexity and stimulus content.  As noted in the prior post in this series, cognitively complexity is considered an index of general intelligence (g). For interested readers, a classic article on the use of the MDS radex model in the analysis of intelligence measure was published in Intelligence in 1983 (Marshalek, Lohman & Snow, 1983). Below is a visual-spatial representation of the MDS results in the current university sample, using the same composite measures as reported above (in the PCA g-loading analysis).  The interpretations below are mine.



The two broad cognitive processing continuim interpretaions (X- and Y-axis) are not the focal point of the current discussion.  The most critical finding in the current context, as per the radex model, is how close to the center of the figure a composite measure score is placed.  Measures that are the closest to the center are considered the most cognitively complex.  As measures move further away from the center, they are judged to be less cognitively complex. 

The results, although using a different method than PCA, produce the same conclusions.  The most cognitively complex measure in this sample (which could be thus interpreted as the best index of cognitive complexity or g-ness) is the KAIT Fluid Intelligence Scale.  The next closest to the center of the figure are the WJ III Gc (Comprehension-Knowledge) and Gf (Fluid Reasoning) clusters.  Again of interest is the location of the WAIS VC composite...it is much less cognitively complex than these other measures, and interestingly, much less cognitively complex than the two similar measures of Gc abilities (WJ III Gc composite; KAIT Crystallized Intelligence composite).  I've also provided my stimulus content hypothesis interpretations of groupings of composties (designated by ovals) from across the batteries (e.g., Processing Speed- WAIS Processing Speed and WJ III Gs or Processing Speed).

The findings in this one sample (which therefore warrants caution in generalization), suggests that the WAIS-III Verbal Comprehension composite, which is the most valid measure of Gc or verbal abilities on the WAIS-III, may NOT be all it is thought to be--when it comes to tapping cognitively complex cognitive processing.  Other Gc or verbal measures from intelligence batteries with adult norms (KAIT; WJ III) were found, in a relative sense, to be much better indicators of a person's g-ness (general intelligence).  The data suggest that, in this sample, even the WAIS-III Working Memory Index score may be a better relative proxy for g-ness than the WAIS verbal composite.

These analyses raise interesting questions about Atkins decisions that have relied either exclusively on the WAIS-R/WAIS-III scores, particularly when part scores (Verbal IQ, Verbal Comprehension; Peformance IQ; Perceptual Organization; etc.) are used instead of the Full-Scale IQ to determine mental retardation, or when the respective WAIS verbal composite is considered better than other potential test global IQ scores (or similar Gc/verbalcomposite scores from other batteries) in making a determination of level of general intellectual functioning.

How can this be?  How can a major scale from the "gold standard" of IQ tests (as it is commonly called in Atkins decisions) be a poorer estimate of general intelligence (g-ness) than most psychologists think?  More importantly, what are the implications for Atkins decisions, when the WAIS-R/III Full Scale score has been questioned as an accurate g-estimate in the face of considerable profile variability, and then the Verbal IQ/Verbal Comprehension Index is used to estimate g-ness (general intelligence)?

In both the Maldonado and Vidal decisions considerable stock was placed on the respective Wechsler verbal composite scores as being the best indicator of general intelligence (for making a determination of mental retardation).  In Maldonado, the reliance of the Wechsler verbal composite trumped a more comprehensive CHC-based IQ battery (BAT-R) administered in Maldonado's native language (Spanish).  Of concern in the Vidal decision, is that he had been administered four different versions of the Wechsler IQ batteries (over many decades), and they consistently revealed a large verbal/nonverbal (performance) IQ split.  Thus, arguments hinged extensively on the Verbal IQ vs the Full Scale IQ.  I'm perplexed why the experts in intelligence and intelligence testing, even without knowing the result of the above g-analysis, did not say "we've got consistent Wechsler V-P split information, I think it would be important to administer more contemporary intelligence tests, or parts of some of these batteries, to find out more information about important g-related cognitive abilities (for the defendant) not measured by the WAIS battery."  The Wechsler batteries had consistently captured Vidals abilities as measured per that battery--wouldn't time have been better spent, and a decision made on a higher quality array of cognitive information, by requesting administration of other IQ tests (or parts of other IQ tests) instead of arguing over old and consistent limited cognitive data? In fact, I, together with Flanagan and Ortiz, published a book in 2000 (The Wechsler Intelligence Scales and Gf-Gc theory:  A contemporary approach to interpretation) that presents procedures for augmenting the various Wechsler batteries to provide for a more comprehensive CHC/Gf-Gc based assessment of a persons intellectual functioning.  This information was also available as early as 1998 (see ITDR by McGrew and Flanagan).  [conflict of interest note - I coauthored these two books which made little in the way of ching-$ for the authors.  They are now both not being printed and none of the authors are receiving any royalties from their sales].

I continue to be baffled/troubled by the over-reliance, and almost god-like stature of the various versions of the WAIS (R/III/IV), in Atkins rulings.  It has been well known (and written about in articles and books; click here) since the early 1990's, that contemporary CHC (aka, Gf-Gc) theory had emerged as the consensus model of intelligence and, more importantly, instruments had been designed (with adult norms) to measure many of the unmeasured or poorly measured CHC abilities not taped by the WAIS-R/III batteries.  If I was an attorney arguing an Atkins case, on either side of the fence, I would seek intellectual testing beyond the so-called "gold standard."  I would want the best possible estimate of g-ness (since this seems to be the crux of the first prong of MR determination in most Atkins cases).

IMHO, the major problem is that of the "inertia of tradition" in intelligence testing, particularly in psychology disciplines that deal with adult populations.  Many practicing psychologists, esp. those working in adult settings whose professional associations and journals have paid less attention to contemporary intelligence theory and test development (less than school and educational psychologists), simply have not kept abreast of these developments. 

How long will Atkins expert intelligence testimony, expert debates, and decisions be made in the face of the Atkins MR IQ Theory-Test gap?  Isn't this simply wrong?  Professionally and ethically shouldn't psychologists who offer judgements in life-or-death decisions hinging on IQ test results be "up to speed" regarding contemporary intelligence theory and instruments?  Should the courts continue to handicapped by the presentation of intelligence test results that are not based on the best evidence from intelligence theory, research, and test development?  The courts are at the mercy of experts who testify, experts who I believe need to be familiar with the cutting edge empirical and theoretical information on the structure of human intelligence and various IQ batteries that are available, beyond the Wechslers.

Given the data presented above, it is possible that the decisions in at least two cases (and I'm sure there are more), may have had a different outcome, or at least an outcome based on a more comprehensive set of intelligence information.  Justice could have been better served via more contemporary intellectual testing practice and interpretation.

I continue to be troubled by this issue.....I need to stop writing and reflect...and will post more on it in the future.

Stay tuned...this series may continue as I analyze other data sets.

Technorati Tags: , , , , , , , , , , , , , , , , ,



Wednesday, October 14, 2009

Calling all Atkins MR death penalty Amicus Briefs


I just added another Amicus Brief, this one filed by AAIDD in the case of Briseno v Quarterman, to the Amicus Brief section of this blog.  I've also added specificity to the labeling of the other briefs by designating  who filed the brief.

If readers are aware of other Amicus Briefs that have been filed in Atkins MR death penalty cases, please forward to me.  In addition to building a library of Atkins MR death penalty decisions, I'd similarly like to build an accessible list of Amicus Briefs filed in such cases.

Thank you.

Technorati Tags: , , , , , , , , ,


Tuesday, October 13, 2009

APA Div. 41, 33 and Law and Human Behavior journal



This weekend I joined Division 41 (American Psychology-Law Society) of the American Psychological Association.   I will be monitoring publications and activities related to the Intellectual Competence and Death Penalty blog.  As a result I've add the divisions journal, Law and Human Behavior, to the list of professional journals monitored by this blog (see listings on right-side of blog).

I've also mailed my application and dues to Division 33 (Intellectual and Developmental Disabilities) and will make a similar post once I receive my membership notification.

Technorati Tags: , , , , , , , , , , , ,

Monday, October 12, 2009

More voodoo psychometrics in Atkins MR death penalty cases? This time adaptive behavior

Voodoo psychometrics strikes again!

A few days ago I made a post (and a subsequent brief FYI follow-up post)  regarding a number of theoretical and psychometric issues that surfaced in the Maldonado (2009) Atkins MR death penalty court decision (click here and here).  It was my opinion that a number of questionable arguments and decisions had been made re: the entirety of the psychometric intelligence test data available in the Maldonado (2009) decision. I won't repeat them here.

Probably my biggest criticism was the use of a non-empirical, unvalidated n=1 psychologist clinical procedure to upwardly adjust IQ scores based on educational and cultural background variables.  Today I learned  that similar "social-cultural" upward adjustment of adaptive behavior scale scores, which are the foundation of the second prong of the determination of mental retardation (or not) in Atkins cases, has also occurred in a number of Atkins cases.  This time I was able to locate an article by Denkowski and Denkowski (2008) that outlined the logic and reasoning for the recommended "systematic" procedure.  I read it in psychometric disbelief!

There is no reason for me to outline my psychometric criticisms, as a number of authors replied with most of the arguments I would have made.  These response, in the same journal, are by Widaman & Siperstein (2009) and Olley (2009) (click here for post re: AB-related chapter by Olley &Cox, 2008). Denkowski and Denkowski (2009) then reply.

As an applied psychometrician, I concur with most of the reactions and arguements of Widaman, Siperstein and Olley.  There is simply insufficent scientifc and psychometric grounds for the upward adjustment of AB scores as outlined.  Yes, psychologists are trained to use clinical adjustment when interpreting test scores, and I so invoked such judgement when conducting countless intellectual assessments during my years as a practicing school psychologist, but clinical judgement is not the same as the development of special score adjustments of nationally standardized psychological instruments based primarily on logic and reason.

Technorati Tags: , , , , , , , , , , , ,


Thursday, October 8, 2009

AP101 Brief #1a: g or not to g in Atkins MR death penalty cases



Applied Psychometrics (AP) 101 Brief #1a:  g or not to g in Atkins MR death penalty cases (first in a series)

Despite whether one believes that general intelligence (g) exists, or not (e.g., John Horn), and ignoring the search for the essence of g (via elementary cognitive tasks measuring reaction time, temporal processing, etc.) at the level of brain mechanisms (e.g., Jensen's neural efficiency hypothesis), it is clear from a reading of most Atkins IQ MR death penalty cases that psychological experts testifying in these cases [primarily because of the emphasis on a "deficit in general intellectual functioning" as the first prong in MR diagnosis in the courts, as per recognized professional association definitions of mental retardation; APA, AAIDD] often argue for different IQ scores as being more accurate estimates of the persons g-ness (IQ) than others.

For example, both in Davis (2009), and especially in Vidal (2007), major arguments focused on whether the Full Scale IQ score from theWAIS-III/IV was the best index of g-ness (and thus mental retardation or mental capacity), or whether one of the part scores (e.g., Verbal IQ, Performance IQ) should be used as the best estimate of the persons g-ness (due to extreme variability in the part scores). My "g-estimate is better than your g-estimate" appears a fundamental point of contention at the core of many Atkins cases,  given the assumption that mental retardation is a global deficit in intelligence (see guest post by Watson for some alternative thoughts and excellent insights on the global vs modular nature of intelligence),

Then, along comes Maldonado (2009) where the g-ness argument, at one juncture, is based on the belief that the Spanish WAIS-III Verbal IQ, which is best interpreted as a CHC measure of crystallized intelligence (Gc), should take precedence over the BAT-R total composite score that is comprised of Gc and six other broad CHC abilities.

"My g-estimate....your g-estimate......this special "nonverbal" g-estimate is more accurate for this individual....that is not a good g-estimate....etc......" back-and-forth arguments beg for empirical scrutiny.  So....buckle up and lets examine some real data.......in search of g-ness.  This is the introduction to a small series of posts that will eventually examine, with empirical data, the relative g-ness of the "gold standard" (WAIS-III/IV) composite scores that are most often debated in these matters.

But first a definition and some methodological background information.  According to the APA Dictionary of Psychology,  general intelligence (the general factor) is:
  • a hypothetical source of individual differences in GENERAL ABILITY (emphasis in original) , which represents individuals' abilities to perceive relationships and to derive conclusions from them.  The general factor is said to be a basic ability that underlies the performance of different varieties of intellectual tasks, in contrast to SPECIFIC ABILITIES (emphasis in original), which are alleged each to be unique to a single task (p. 403).
[Note - some of the the text below comes from Flanagan, McGrew & Oritz (2000).  The Wechsler Intelligence Scales and Gf-Gc theory.  Boston:  Allyn & Bacon.

Intelligence tests have been interpreted often as reflecting a general mental ability referred to as g (Anastasi & Urbina, 1997; Bracken & Fagan, 1990; Carroll, 1993a; French & Hale, 1990; Horn, 1988; Jensen, 1984, 1998; Kaufman, 1979, 1994; Keith, 1997; Sattler, 1992; Sattler & Ryan, 1999; Thorndike & Lohman, 1990).  The g concept was associated originally with Spearman (1904, 1927) and is considered to represent an underlying general intellectual ability (viz., the apprehension of experience and the eduction of relations) that is the basis for most intelligent behavior. The g concept has been one of the more controversial topics in psychology for decades (French & Hale, 1990; Jensen, 1992, 1998; Kamphaus, 1993; McDermott, Fantuzzo, & Glutting, 1990; McGrew, Flanagan, Keith, & Vanderwood, 1997; Roid & Gyurke, 1991; Zachary, 1990).

According to Arend et al., (2003),  Jensen (1998a, 1998b) proposed that cognitive complexity  might represent a fundamental aspect of g an could be quantified based on inspection of the test measures loadings on the first unrotated factor, because complex tasks show higher factor loadings than simple tasks on that factor.  In many respects when psychologists are discussing mental retardation and general intelligence, there is an implicit assumption that low general intelligence (e.g., mental retardation) is reflected most clearly on performance on the most cognitively complex measures (i.e., high g measures). 

As with the controversy surrounding the nature and meaning of g, disagreements exist about how best to calculate and report psychometric g estimates.  Most all methods are based on some variant of principal component, principal factor, hierarchical factor, or confirmatory factor analysis (Jensen, 1998; Jensen & Weng, 1994).  Although a hierarchical analysis is generally preferred (see Jensen, 1998, p. 86), as long as the number of tests factored is relatively large, the tests have good reliability, a broad range of abilities is represented by the tests, and the sample is heterogeneous, (preferably a large random sample of the general population), the psychometric g's produced by the different methods are typically very similar (Jensen, 1998; Jensen & Weng, 1994).  For the interested reader, Jensen’s (1998) treatise on g (The g Factor) is suggested, as it represents the most comprehensive and contemporary integration of the g related theoretical and research literature.

Operationally the determination of high, moderate or low g-ness of tests or composites has typically been based on each measures correlation (aka., factor or principal component loading) with a single common factor, component, or dimension extracted from the correlations among the set of measures in question.  Measures that "load" high on the g-factor are considered to be the better estimates of general intelligence.

Consider the following simple analogy (which is not original...I borrowed the conceptual idea from Cohen et al., 2006).  You have a special pole that posses a special form of  magnetism (general intelligence). You throw a bunch of  metal marbles (which are the test measures), which have different degrees of the same magnetic force, into a box with the pole at the center.  You gently shake the box.  When you open the box, there is one "king" marble at the top of the poll (it has the highest degree of shared magnetism with the strongest part of the pole), followed next by the next strongest....and so on until the metal marble with the least amount of shared magnetic force is at the bottom.  The pole represents g (general intelligence) and the ordering of the metal marbles (the test measures) represents the ordering of the g-ness (degree of shared magnetic force) of the measures.  The "king" test/marble is assigned the highest numerical index, with each succeeding (and lower) test/marble assigned a slightly lower numerical index of g-ness (shared magnetism).

This is what principal component analysis conceptually accomplishes with a collection of IQ test measures.  It statistically orders the various psychometric measures from strong g-loading to low-g-loading.  This is the typical and traditional statistical currency used by psychometericians and psychologists when discussing the degree of g-ness or g-saturation of different measures--those measures most important for establishing an estimate of a person's general intelligence.


The problem with within-battery factor analysis is that it can affect the g-estimates.  For example, a test’s loading [note- g-loadings are most often computed for the individual subests in a test battery, and not the composite scores such as Verbal IQ, processing speed, etc.-- it is the later, the g-ness of composite scores, which appears to be a critical issue in many Atkins cases.  Thus, when reading the this text I will refer to the measures g...which could mean test or composite] on the general intelligence (g) factor will depend on the specific mixture of measures used in the analysis (Gustafsson & Undheim, 1996; Jensen, 1998; Jensen & Weng, 1994; McGrew, Untiedt, & Flanagan, 1996; Woodcock, 1990).  If a single vocabulary measure is combined with nine visual processing measure, the vocabulary measure will most likely display a relatively low g loading because the general factor will be defined primarily by the visual processing measures.  In contrast, if the vocabulary measure is included in a battery of measure that is an even mixture of verbal and visual processing measures, the loading of the vocabulary measure on the general factor will probably be higher.  It is important to understand that measures g loadings, as typically reported, only reflect each measures relation to the general factor within a specific intelligence battery.  Although in many situations a measure g loading will not change dramatically when computed in the context of a different collection of diverse cognitive tests (Jensen, 1998; Jensen & Weng, 1994), this will not always be the case.

Within (internal-validity) vs across (joint; external validity) estimation of test measures g-ness

When measures from different batteries are combined in the joint-battery approach, the battery-bound g  estimates for some measures may be altered significantly.   Flanagan et al. (2000) demonstrated these when they calculated within- and joint-battery g estimates for the WISC-III.  These estimates were derived from a sample of 150 subjects who were administered the WISC-III and WJ III cognitive measures as part of the Phelps validity study reported for the WJ III cognitive technical manual.  Within-battery g estimates were calculated with the WISC-III data based on the first unrotated principal component.  Next the joint-battery factor analysis allowed for an examination of the WISC-III g estimates when calculated together with another intelligencet test battery (WJ III), one that included a broader array of CHC abilitiy measures.

Flanagan et al. (2000) reported that the within- and joint-battery WISC-III g loadings were similar for many of the individual measures.  For example, the within- and joint-battery test g loadings are generally similar (i.e., do not differ by more than .05) for the Similarities (.76 vs .71), Vocabulary (.78 vs .74), Digit Span (.48 vs .49), Block Design (.60 vs .61), Object Assembly (.50 vs .45), and Symbol Search (.57 vs .54) measures.  These six WISC-III measures appear to have similar g characteristics when examined from the perspective of either the WISC-III or CHC (WJ III battery) frameworks.  However, the joint-battery g loadings were noticeably lower than the within-battery g loadings (i.e., lower by .06 or more) for Information  (.77 vs .68), Arithmetic (.70 vs .64), Comprehension (.59 vs .51), Picture Completion (.50 vs .40), Picture Arrangement (.37 vs .31), and Coding (.46 vs .37).   The results suggested that the latter WISC-III measures were relatively weaker g indicators than is suggested by within-battery WISC-III g analysis.

This example demonstrates the potential chameleon nature of test measures g estimates that are calculated within the confines of individual intelligence batteries when compared to those calculated within a comprehensive set of ability measures. 

And, yet to be mentioned is another, older, and for some reasons under-utilized statistical method for examing the g-ness (congitive complexity) of IQ test measures...multidmensional scaling (MDS).  We will save that for the next post in this seires.

To be continued........................

.

Maldonado (2009) miscarriage of psychometric justice PS

No sooner had I made my Maldonado (2009) Atkins IQ MR death penalty miscarriage of psychometric justice post and someone sends me a link to a post at the StandDown Texas Project regarding a case where the clinical adjusting of IQ scores, as described and challenged in my prior post, was questioned.....and even referred to as "junk science."

Technorati Tags: , , , , , , , , ,


Maldonado (2009) IQ MR Atkins death penalty decision: A psychometric miscarriage of justice?

I'm no longer shocked by what passes as credible psychological/psychometric evidence or testimony in some Atkins IQ MR death penalty court decisions.  Another such decision (Maldonado, 2009) has come to my attention.  A link to a PDF copy of the decision is now available in the Court Decisions section of this blog. 

Readers can gather all the relevant background information re: the case by reading the entire decision.  I intend to focus primarily on the psychological interpretation and psychometric issues in the decision that are troubling. 

But first, before delving into these issues, I  present a few quotes (from the Maldonado record) that capture the confusion and uncertainty surrounding many Atkins decisions, a situation that is resulting in considerable variability in the quality of psychological assessment data reported/interpreted and the common occurrence of "dueling expert witnesses".

Quotes from the record:
Because the Supreme Court did not establish a bright-line test to identify mental retardation, the Atkins inquiry has become a fact-intensive question that heavily relies on the opinions provided by mental-health experts. This case, like most involving Atkins claims, requires consideration of testimony from competing experts who disagree about the nature of mental retardation, the means by which it may be identified, the manner in which it manifests in a criminal defendant’s life, and the psychological profession’s role in making the legal decision of whether mental capacity precludes execution. (p.32)

A “welter of uncertainty” followed the Atkins decision because “[t]he Supreme Court neither conclusively defined mental retardation nor provided guidance on how its ruling should be applied to prisoners already convicted of capital murder.” Bell v. Cockrell, 310 F.3d 330, 332 (5th Cir. 2002). Accordingly, federal courts have approached the implementation of Atkins with some trepidation. (p.36)
After reading a number of of Atkins rulings (see Court Decisions section of this blog), I could not agree more with these statements.  The courts appear ill-equipped to handle the complex psychological measurement issues presented, issues that are, at times, confounded by the inclusion of data from dubious procedures, interpretations of test scores that are not grounded in any solid empirical research, and the deference to a single intelligence battery (the WAIS series) as the "gold standard" when a more appropriate instrument (or combination of WAIS-III/IV and other measures) might have been administered, but the results of the more appropriate measure are summarily dismissed based on personal opinion (and not sound theory or empirical research).


Below are some of the troubling psychological/psychometric issues I see in the Maldonado (2009) decision

Lack of English language proficiency:
  "A theme developed by both parties is that Maldonado’s lack of English proficiency has madehis exact intellectual capacity difficult to gauge." (p.38).
  • As a result, the prosecution psychologist administered the WAIS-III "through a translator" (p.42).  The translator administered WAIS-III resulted in Verbal, Performance, and Full Scale scores of 74, 74, and 72 respectively.  A defense psychologist captures the essence of my reaction to the use of a translator-administered WAIS-III.  “The accepted practice in the evaluation on Spanish-speakers is to communicate with the client in Spanish without the use of translators. In addition, tests should be scientifically translated and validated and the most appropriate norms available should be applied" (p.49).  I agree.  I am unaware of any professionally established and endorsed procedure for the translated administration of the English-normed WAIS-III.  Such a procedure violates a fundamental backbone of the science of individual intelligence testing--standardized test administration.
Use of non-empirical clinical judgement procedures to upwardly adjust IQ scores.  On page 53 of the record, it is indicated that the prosecutions psychological expert believed that the translated English-normed WAIS-III scores needed to be upwardly adjusted due to Maldonado's educational and cultural background.  Additional support came from the finding of a Verbal IQ score of 83 on the WAIS Espanol (p.56).  Although psychologists are appropriately trained to recognize the potential impact of such environmental variables when interpreting scores, the psychologist  upwardly adjusted the scores to a specific IQ score estimate ("It’s around the 80s, I guess, if you had to pin me down. Around the 80s; somewhere in there"- p.48) and this expert "conceded that only 'clinical judgment,' not any statistical formula or established methodology, informed how much to alter an IQ score because of cultural and educational factors" (p. 53). 

My concern with this procedure mirrors the testimony of the defense experts in the case.  Adjusting obtained IQ scores, either up or down, based on an n=1 professionals clinical judgement, in the absence of any scientifically established procedure for adjusting IQ scores, is troubling and is not consistent with accepted psychological assessment practices or standards.  In fact, this IQ adjustmend procedure sounds similar to a notable empirical effort (in the late 1970s and early 1980s) to produce IQ scores that better reflected a persons social-cultural background.  Jane Mercer's SOMPA (System of Multicultural Pluralistic Assessment) was a valiant effort to adjust Wechsler IQ scores for African-American individuals based on their social-cultural knowledge and history.  The result was a new score called Estimated Learning Potential (ELP).  From the start, SOMPA was controversial and eventually was found to be flawed for many reasons (see Hellfinger, 1987; also Jirsa, 1983).  If a reasonably conceived theoretical and empirical IQ adjustment procedure (i.e, SOMPA), which was intended to account for a person's social-cultural background, was found to be flawed, how can a specific n=1 psychologist be endowed with unique insights that allow for the invoking of an unspecified personal algrorithm to make IQ score adjustments?  This is indeed troubling.  Also, to the best of my knowledge, SOMPA is no longer around and is not, or is seldom, used.  Race-based or adjusted norms have not been recognized as an acceptable professional psychological assessment practice for over twenty years!

Dismissal of the BAT-R intelligence results.  Maldonado had also been administered the Woodcock-Muñoz Bateria-R (“Bateria-R”), the Spanish-language counerpart of the Woodcock-Johnson Test of Cognitive Abilities--Revised [conflict of interest notice:  I am not a co-author of the BAT-R  or WJ-R, but was a paid measurement consultant on the WJ-R project and have since become a coauthor of the subsequent edition, the WJ III].  Maldonado obtained a Broad Cognitive Ability (BCA) score, which is analagous to the full-scale score from other intelligence batteries, of 61.  Of all the cognitive tests administered, this is the only comprehensive intelligence battery that was administered in Maldonado's natural language and where his performance is compared against appropriate US-equated Spanish norms [note-- the WAIS Español administered was also administered in his natural language and makes comparisons against Spanish norms, but only a portion, the Verbal section, was administered].  Also, as previously noted at this blog, the WJ-R/BAT-R and WJ III/BAT III provide for the most comprehensive assessment of intellectual functioning as per the consensus model of human intelligence (CHC theory) among serious intelligence scholars.  This is what I've termed the "Atkins MR death penalty IQ test-theory gap." 

Why were the BAT-R scores dismissed? 

The prosecution expert "testified that the AAMR has not cited the Bateria-R as a predicate test to the evaluation of mental retardation and “[i]t’s not well suited for that purpose, although you can use it for that" (p. 70).  He also stated that the "Bateria-R test score was especially suspect because it was inconsistent with Dr. _____'s administration of the WAIS Español in which Maldonado scored well above the range for mental retardation." (p.70). Furthermore, this expert "opined that the usefulness of the test was impaired because it 'is used by school psychologists to diagnose learning disabilities and it measures a lot of things like visual and auditory processing. It really measures very little in terms of general intelligence.' " (p.70).  As a result, "the state habeas court dismissed the Bateria-R score because it “is not one of the tests the AAMR cites for mental retardation evaluation,” but instead “is generally used by school psychologists to diagnose learning disabilities” and, in fact, “is not very relevant for establishing general intellectual functioning, so it is not well-suited for determination of the first prong . . . to determine mental retardation" (p.70-71).

There are many problems with the reasons given for dismissing the most culturally appropriate (for Maldonado) and comprehensive measure of intelligence (BAT-R). 

First, dismissing an instrument because it is used primarily by a particular set of psychologists (school psychologists) is non-sensical, and frankly, condescending.  School psychologists typically give many more intelligence tests than psychologists working in adults settings and use these instruments to diagnose mental retardation. School psychologists, in many respects, have more intimate familiarity with intelligence testing than most other professional psychologists.  

Second, the previously mentioned national expert panel that examined the Dx of MR at the same cut-point as Atkin's cases (for SSA benefits) indicated that most intelligence tests would be moving towards measuring the Cattell-Horn and Carroll Gf-Gc models of intelligence (now collectively referred to as CHC theory; also see McGrew, 2009), and instruments based on this model are very relevant to the Dx of MR.  The English version of the WJ-R WJ III (a CHC-based revision of the WJ-R) was listed as one of the approved instruments by the expert panel...which should also implicitly argue for use of the Spanish-language versions.  The prosecution witness, who appears  stuck in the land where the WAIS is the "gold standard," appeared unaware of the recent advancements in understanding the psychometric nature of human intelligence, which has converged on the CHC model of intelligence.  This is particularly ironic given that both John Horn (of Cattell-Horn) and Jack Carroll served as theoretical consultants on the WJ-R...which was the foundation of the BAT-R.

Third, stating that the BAT-R (and WJ-R by implication) is not a respected measure of general intellectual functioning reflects a complete lack of awareness of the CHC-foundation of the instruments, as well as published research (including the WJ-R and BAT-R technical manuals and bulletins).  If CHC theory is the consensus model of psychometric intelligence, then the only battery administered to Maldonado that measured most of the model, which in most conceptualizations has general intelligence (g) at the apex, should have been given serious weight.  I, and others, heard Dr. Arthur Jensen, the most prominent pscyhometric expert on g, at an ISIR conference in Nashville, TN, state, during a discussion of a presentation in front of the entire audience, that (at the time) he considerd the WJ-R (and, thus, the BAT-R by implicit endorsement) the best available intelligence battery for measuring g.  I will return to this point in future posts as I've been analyzing WAIS-III data together with the WJ batteries (as well as other accepted intelligence batteries, e.g., KAIT, K-ABC; SB-IV; SB-IV) to evalute the g-ness of each battery when jointly analyzed.

Fourth, prosecution psychologist used the WAIS Español Verbal score as evidence that the BAT-R was not accurate.  The problem with this logic is that the WAIS Verbal scale is known to be an excellent measure of crystallized intelligence/comprhension-knowledge (Gc), only one of the major 7-8 domains in the CHC model of intelligence.  Conversely, the BAT-R includes indicators from seven of the major CHC broad ability domains, only one of which is Gc.  No attempt was made to compare the WAIS Gc score with the corresponding Gc measure(s) on the BAT-R.  They might have been similar...or not.  More importantly, you cannot compare a measure of Gc to one that includes Gc and fluid reasoning (Gf), visual-spatial processing (Gv), short-term memory (Gsm), long-term retrieval (Glr), auditory processing (Ga), and processing speed (Gs).

I could go on and comment on other issues, such as the use of only the Verbal scale of the Spanish WAIS-III, as well as the whole area of adaptive behavior, but these are issues for another post, possibly by a guest blogger.

Technorati Tags: , , , , , , , , , , , , , , , , , , ,

Wednesday, October 7, 2009

Guest bloggers make a difference: Atkins MR IQ death penalty blog guest posts



The above graph represents new people who have visited the Intellectual Competence and Death Penalty blog over the past month.  Notice the various "spikes."  Most all of are directly related to guest blog posts by psychologists or legal professionals involved in Atkins IQ MR death penalty cases.  Today's spike (10-7-09) is directly related to the guest blog post by Dr. Dale Watson.  Thank you Dr. Watson.

If you are interested in issues related to this blog, and want an occasional  platform from which to express thoughtful, scholarly, and logical comments, and/or research synthesis, etc. regarding Atkins IQ MR death penalty cases, a guest post at this blog would be welcome...and, in turn, would provide you some good exposure.  Psychologists and legal professionals are both welcome.

If you are interested, contact me privately at:  iap@earthlink.net.

The blogmaster.




Davis (2009) Atkins decision: IQ part scores and modular nature of intelligence: Guest post by Dr. Dale Watson

A number of recent Atkins-related IQ/MR death penalty court decisions have raised important issues re: the interpretation of variability in part scores when compared to the total (full scale-general intelligence; g) composite index from an intelligence test.  In particular, Davis (2009) and Vidal (2007) are excellent examples of the complex issues...and how they relate to a diagnosis (or not) of mental retardation.

Dr. Dale Watson has studied the Davis (2009) decision and has provided the following thoughtful guest blog post (click here for other guest blog posts].  Thank you Dr. Watson for the excellent analysis and commentary.  I (the blog dictator) have reproduced his post "as is" with a few exceptions.  I've added a few URL links to other sources.  I've also added emphasis to certain statements via underlining, followed by a note that I added the emphasis.  Finally, I've yet to figure out if it is possible to add real footnote superscripts to blog text via the blog editor I use.  So, I've adopted the format of putting footnote numbers in brackets [ ] and placing the footnotes at the end of the blog post.

I encourage other psychologists and measurement specialists to review some of the Atkins court decisions posted at this blog and send me their analysis and thoughts for potential blog posts. 


By Dr. Dale Watson

In a recent Maryland case, U.S. v. Davis, (2009 WL 1117401 (D.Md.)) the trial court found the defendant to be mentally retarded and thus ineligible for the death penalty in line with Atkins v. Virginia, 536 U.S. 304, 122 S.Ct. 2242, 153 L.Ed.2d 335 (2002).  An analysis of the IQ results in this case highlights issues regarding the impact of part-score variability in the IQ profile and the diagnosis of mental retardation as well as the modular nature of intelligence. [emphasis added by blogmaster.  Blogmaster note--Also, see prior post regarding the recommended use of total vs part scores from a national panel report ].

Earl Davis, the defendant, had been charged with federal crimes, including murder, which made him potentially eligible for the death penalty unless he were found to be intellectually and developmentally disabled.  In a pre-trial proceeding the court heard evidence regarding whether Mr. Davis was in fact mentally retarded.  The defense presented the testimony of five imminently qualified experts supporting their position.  The prosecution, relying on the opinions of two board certified neuropsychologists, alternatively argued that Mr. Davis did not meet the criteria for being mentally retarded and instead demonstrated a learning disability.  This argument was founded, in part, on the finding of a significant discrepancy between the Verbal Comprehension (VCI) and Perceptual Reasoning Indices (PRI) of the WAIS-IV.[1]   The prosecution contended that the 15-point discrepancy between these indices made the Full Scale IQ (FSIQ) unreliable as an overall measure of his intellectual functioning.[2]   A discrepancy of this magnitude (VCI = 66; PRI = 81) could be expected to occur with a base rate of 12.7 percent within his ability range and could thus be considered “abnormal.” [3]   This argument was advanced despite the fact that Mr. Davis’ WAIS-IV FSIQ of 70 nominally met the first prong of the definition of an intellectual disability (mental retardation).
 
The prosecution argument was not without precedent.  For example, it was noted in testimony that the Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition, Text Revision (DSM-IV-TR) reported:
When there is a significant scatter in the subtest scores, the profile of strengths and weaknesses, rather than the mathematically derived full-scale IQ, will more accurately reflect the person’s learning abilities.  When there is a marked discrepancy across verbal and performance scores, averaging to obtain a full-scale IQ can be misleading (U.S. v. Davis citing DSM-IV-TR, p. 42).
Fiorello et al. (2002) summarized the literature addressing this viewpoint as applied to the Wechsler scales:
Sattler (2001) cautions that the FSIQ may misrepresent a child’s cognitive functioning level if the Verbal IQ (VIQ) and Performance IQ (PIQ) are significantly different; however he indicates that no empirical evidence exists to indicate when the FSIQ should not be reported or used in eligibility decisions.  Prifitera, Weiss, and Saklofske (1998) recommend that FSIQ should not be interpreted when differences between VIQ/PIQ or Verbal Comprehension Index (VCI) and Perceptual Organization Index (POI) scores are extreme, which they define as differences found in less than 10% of the population (p. 117). [4]
The reluctance to make the diagnosis of mental retardation in the face of significant discrepancies between measures of verbal and non-verbal abilities is thus perhaps not surprising or new.  There has been a long-held view that mental retardation is marked by a “flat” cognitive profile.  When confronted with a profile that instead shows a pattern of ipsative strengths and weaknesses perhaps many psychologists would be hesitant to make the diagnosis of mental retardation.  However, recent empirical evidence does not support that position. [emphasis added by blogmaster]

It is the case, as noted in the WAIS-IV Technical and Interpretive Manual, that “The prevalence of large and unusual discrepancies between verbal and nonverbal composite scores [on the Wechsler scales] has been shown to decrease with decreasing levels of ability (internal citations omitted)” p. 102.  However, Bergeron and Floyd (2006), using the Woodcock-Johnson III Tests of Cognitive Abilities (WJ-III Cog.), demonstrated that individuals with mental retardation “will not likely display a flat cognitive profile on comprehensive assessments of CHC broad cognitive abilities—especially when measures vary widely in g loadings—regardless of the information presented in test manuals for mental retardation groups (emphasis added).” [5]  These researchers found that, in children with mental retardation, there was significantly greater intra-individual score ranges (scatter) than in average-achieving individuals.  Specifically, nearly 37% of these children obtained at least one CHC factor score within the average range.  Bergeron and Floyd concluded,
…it is likely that the increasing number of specific cognitive abilities measured by intelligence test batteries and the variation of these scores in individual profiles inadvertently muddies the waters of mental retardation diagnosis.  As a result, when faced with IQs in the range described in the diagnostic criteria for the disorder and part scores that are much higher, practitioners may believe that an individual cannot be diagnosed with mental retardation because of evidence of “intact” or “unimpaired abilities” p. 428.
…denying special education eligibility or failing to make a diagnosis of mental retardation based on significant part score variability may do these children disservice when other ecologically valid evidence of mental retardation (e.g., adaptive behavior skill deficits) indicates genuine need. P. 429.
These findings have implications for the role of g as well as domain specific disabilities in the genesis of mental retardation.  Bergeron and Floyd posited that the centrality of the impaired CHC factors, as measured by the factor’s g-loading, would determine its sensitivity to mental retardation.  In fact, the Comprehension-Knowledge and Fluid Reasoning clusters, with the highest g-loadings, were most commonly associated with the diagnosis of mental retardation, though low average or average scores on one of these clusters did not preclude the diagnosis.

In another line of research, Anderson (1998) has suggested the need to distinguish “between mental retardation as a general deficit of thinking and mental retardation that might result from the global effects of a specific deficit in a cognitive module.” [6]  In a similar vein, Frith and Happé (1998) have argued “that general impairments (e.g. low IQ) in developmental disorders need not be the result of primary damage to domain-general mechanisms.  Rather, they may be the developmental consequence of damage to very specific, even modular, mechanisms which act as gatekeepers in development” p. 270. [7]

Neuroimaging data further supports the view of the modularity of cognitive functions.  In such a view, though g may represent a relatively unitary phenomenon in the normal brain it can fractionate in the face of modular neuropathology.  Gläscher et al. (2009) used CT and MR lesion maps to localize the impairments found on the WAIS-III Index scores in focal brain-damaged patients.   These investigators “found (1) impairments in VCI (Verbal Comprehension Index) were associated with damage in left hemisphere, in particular in the left inferior frontal cortex, (2) impairments in POI (Perceptual Organization Index) were associated with damage in right parietal, occipito-parietal, and superior temporal cortex, (3) impairments in WMI (Working Memory Index) were associated with left hemispheric lesions particularly focused in superior parietal cortex, and (4) impairments in PSI (Processing Speed Index) correlated with a number of small regions distributed across both hemispheres” (p. 686). [8]  Each of these indices, with the exception of the PSI, significantly predicted a lesion in the associated region.  These results largely confirm clinical lore regarding the significance of specific impairments in the Wechsler indices.

Williams syndrome, a genetically determined developmental disorder commonly marked by mental retardation, has frequently been characterized by relatively intact language abilities and profoundly impaired visual-spatial functions.  Thompson et al. (2005), using MRI scans, identified specifically increased cortical thickness within the perisylvian and inferior temporal regions of the right hemisphere.  This pattern of regional neuropathology appears to correspond well with the differentiated pattern of cognitive deficits.

Mental retardation may thus result from both general impairments or as the “developmental consequence of damage to very specific, even modular, mechanisms.”  [emphasis added by blogmaster]. In the former case one could anticipate a relatively “flat” profile of intellectual abilities.  In the latter instance, a more extreme pattern of strengths and weaknesses might emerge and yet result in sufficient impairment of overall intellectual functioning as to meet the criteria for mental retardation.  Evaluators would do well to remember that the diagnostic criteria for mental retardation do not limit the diagnosis to a particular ipsative pattern of scores, that mental retardation can arise from diverse etiologies and that “strengths co-exist with weaknesses.” (emphasis added by blogmaster).  These cautions apply equally within a clinical context and within a capital punishment context when the diagnosis is “a matter of life or death.”
  • [1] The opinion reflected some confusion on this issue indicating that the discrepancy was between the VIQ and PIQ despite the fact that these scores are no longer included in the WAIS-IV.
  • [2] See People v. Superior Court of Tulare County (Vidal) for an example of even greater discrepancies between verbal and performance IQs in an Atkins case.
  • [3] The “abnormality” of such discrepancies is somewhat arbitrary but authors have variously set the cut-off for an unusual finding at a base-rate of either 10 or 15 percent level.
  • [4] Fiorello, C.A., Hale, J.B., McGrath, M., Ryan, K. & Quinn, S. (2002).  IQ interpretation for children with flat and variable test profiles.  Learning and Individual Differences, 13, 115-125.
  • [5] Bergeron, R. & Floyd, R.G. (2006). Broad cognitive abilities of children with mental retardation: An analysis of group and individual profiles.  American Journal on Mental Retardation, 111(6), 427.
  • [6] Anderson, M. & Miller, K.L. (1998).  Modularity, mental retardation and speed of processing.  Developmental Science, 1(2), 239-245.
  • [7] Frith, U. & Happé, F. (1998).  Why specific developmental disorders are not specific: On-line and developmental effects in autism and dyslexia.  Developmental Science, 1(2), 267-272.
  • [8] Gläscher, J., Tranel, D., Paul, L.K., Rudrauf, D., Rorden, C., Hornaday, A., Grabowski, T., Damasio, H., and Adolphs, R. (2009).  Lesion mapping of cognitive abilities linked to intelligence.  Neuron, 61, 681-691.

Technorati Tags: , , , , , , , , , , , , , , , , , , , , , , , , ,


Monday, October 5, 2009

Race Rears Its Ugly Head in Atkins case: Guest post by Kevin Foley

Race and IQ tests has been one of the most hotly debated controversies in the field of intelligence testing (click here for some recent post over at sister blog--IQs Corner).  To most psychometrician/measurement experts, the mere presence of mean IQ score differences does not meet the psychometric definition of test bias.  We measurement types typically define test bais as: (a) item content bias, which is now commonly removed via expert panels and differential item functioning [DIF] item analysis methods, (b) structural bias, which means that a test doesn't measure the same constructs across groups [and which is typically evaluated with confirmatory factor analysis structural invariance methods], and (c) predictive bias, which means that a test differentially predicts outcomes for different groups [and which is typically evaluated via the examination of potential differential prediction regression slopes].  Most contemporary intelligence tests address all forms of psychometric bias during test development.

The above mini-course not withstanding, Kevin Foley has written another excellent and though-provoking guest  post (click here to see his prior guest post) dealing with the introduction of racial bias claims, with an unusual "spin" on the interpretation of bias, in the context of experts testifying in an Atkins MR death penalty case in Tennessee.  K. Foley's guest post is reproduced "as is" with any URL links added by the blogmaster.  Thanks Kevin Foley for another well written and provocative post.

In an Atkins case out of Tennessee, two prosecution experts “testified that I.Q. tests have historically been biased against minorities in that they tend to underestimate the intelligence of minorities.” Memorandum Decision, Black v. Bell, ( U.S. Dist. Ct., M.D.Tenn., No. 3:00-0764, Apr. 24, 2008). In an Ohio Atkins case involving convicted murderer Kevin Yarbrough, the State “contend[ed] that the early tests were given at a time that IQ tests were culturally biased against minorities and could have lowered the test results”. Decision Order/Entry, State v. Yarbrough, Ohio Common Pleas Court, Shelby County, Case No. 96CR000023 (Feb. 28, 2007). In a similar vein, the trial court in Ex Parte Chester (unpub., Tex. Ct. Crim. App., No. AP-75,037 (2007)) refused to accord appropriate weight to childhood IQ scores obtained by Chester because the WAIS-R “would not adequately account for cultural, regional, or other types of factors that may have influenced [Chester’s] test results.” Since Chester is black, and his native “region” is reported to be Jefferson County, Texas, the import of this comment obviously refers to his racial background. Two other examples include the case of Eldridge v. Quarterman, 2008 U.S. Dist. LEXIS 19647 (S.D. Tex. 2008), where there was testimony from a psychologist who “acknowledged evidence that minorities score artificially low”, and Maldonado v. Thaler, 2009 U.S. Dist. LEXIS 88988 (S.D. Tex.) (prosecution expert testified that “cultural differences” probably artificially lowered immigrant defendant’s scores).

The author of a recent law review article about the consequences of Atkins asserted that, “There is evidence that some IQ tests feature an inherent cultural bias that leads some minority groups to score lower than other individuals.” M. Libell, Atkins’ Wake: How the States Have Shunned Responsibility for the Mentally Retarded, 31 Law & Psychol. Rev. 155, 162 (2007). Similarly, the author of another law review article claimed that, as a consequence of “flawed” IQ tests, “[e]conomically deprived people and ethnic minorities are . . . often erroneously found to be mentally retarded.” Lori M. Church, Mandating Dignity: The United States Supreme Court's Extreme Departure From Precedent Regarding the Eighth Amendment and the Death Penalty, 42 Washburn L.J. 305, 325 (2003).

Starting with Arthur Jensen’s tome, Bias in Mental Testing (1987), there is long list of scientific literature which would seem to indicate that IQ tests are not biased against minorities. Even though there is gap of approximately one standard deviation between the mean IQ scores of blacks and whites, and about half that amount between whites and Hispanics, that does not mean that IQ tests are biased.

Are the psychologists referred to above showing a bias of their own in attempting to give their party the testimony they want, and are the prosecutors, courts and legal commentators simply providing spin based on myth, not science? It would be nice to hear from some objective psychologists on this issue. Should comments like those above by made by experts in Atkins cases and in the discussions of commentators?
Technorati Tags: , , , , , , , , , , , , , , , , , , , ,


Sunday, October 4, 2009

Atkins MR death penalty IQ test-theory gap: Is this a problem?



[Double cllck on image to enlarge: Click here for a more comprehensive figure that includes IQ tests without adult norms]

Contemporary Cattell-Horn-Carroll (CHC) theory has emerged as the psychometric cognitive/intelligence theory with the largest body of supporting evidence (Kaufman, 2009).  Evidence for the emergence of CHC theory can be seen in a recently invited editorial on CHC theory in the prestigious journal Intelligence (McGrew, 2009).  More importantly, CHC theory “has formed the foundation for most contemporary IQ tests” (Kaufman, 2009, p. 91).

If CHC theory is now the consensus model of psychometric intelligence, and if the intellectual component of most Atkins cases hinges on one of the latest versions of the WAIS (WAIS-III, WAIS-IV), which is often described as the "IQ test standard" in Atkins related decisions, isn't this a serious problem?

Shouldn't life-or-death decisions hinging (to a major degree) on a person's tested level of intelligence be based on the assessment of intelligence as per the most validated model of intelligence?

Would the Vidal (2007) CA decision, which hinged largely on expert debates surrounding existing (prior) Wechsler Verbal, Performance and Full Scale IQ scores, benefited from less debate regarding the meaning of a consistently documented Verbal vs Performance IQ difference and more time spent asking for more comprehensive assessments of Vidal's complete CHC cognitive abilities, many that are either not measured, or are poorly represented, by the Wechsler batteries?

Should the Wechslers continue to be considered "the IQ standard" in these cases?  Shouldn't a standard be consistent with the consensus model of contemporary of intelligence?

It appears that many Atkins IQ MR-determination cases are decided in the presence of an IQ test-contemporary intelligence theory gap.  How long will this continue?   What are the implications?

Many questions to ponder.

Technorati Tags: , , , , , , , , , , , , , , , , , , ,


Saturday, October 3, 2009

Total vs part IQ composite scores in mental retardation determination


In a prior post I recommended a book that dealt with eligibility for social security benefits due to mental retardation.  The content is relevant to the current blog since it represents one of the few recommendations from a national panel of experts re: mental retardation eligibility issues (at the same IQ cut-point pivotal in Atkins cases), issues that are very similar to those often raised in Atkins mental retardation death penalty cases.

I've now extracted the main recommendations/statements from the book that are relevant to the total vs part IQ score issue raised in the CA Vidal (2007) decision.  Emphasis (underline/italics) in the material below was added by the blogmaster.  As noted in the prior post, there was a dissenting opinion expressed by Dr. Keith Widaman.  I plan to add his dissenting comments later. 

As someone who was a practicing school psychologist for 10+ years (where I saw many profiles with wild splits in scores due to many different reasons), and who now is primarily an applied psychometrician conducting research in the domain of intelligence theory and testing, I have a complex set of opinions re: the total vs part score issue.  I have yet to crystallize my opinion regarding the Vidal (2007) decision....in fact...I find it has raised many questions which has resulted in myself digging through books, articles, and analyzing datasets. 

The recommendations and statements below reflect the opinion of the Committee on Disability Determination for Mental Retardation as summarized in the above referenced book.  They do not necessarily reflect my opinions at this time.

Extracted from the book:
A client must have an intelligence test score that is two or more standard deviations (SD) below the mean (e.g., a score of 70 or below, if the mean = 100 and the standard deviation = 15.
  • Composite score is 70 or below:  If the composite or total test score meets this criterion, then the individual has met the intellectual eligibility component.
  • Composite score is between 71 and 75:  If the composite score is suspected to be an invalid indicator of the person’s intellectual disability and falls in the range of 71-75, a part score of 70 or below can be used to satisfy the intellectual eligibility component.
  • Composite score is 76 or above:  No individual can be eligible on the intellectual criterion if the composite score is 76 or above, regardless of part scores. (p.5)

Only part scores derived from scales that demonstrate high g-loadings—that is, ones that are better representations of general intellectual ability (e.g., crystallized , fluid measures of intelligence)—should be used in place of the composite IQ score when its validity is in doubt.  Many intelligence tests assess several facets of intelligence, but not all facets are equally important or predict life events equally well.  Those intellectual facets that are heavily “g-saturated” provide the best sources for replacing the composite IQ score when its validity is questionable. (p.5)

The use of part scores, most often from the Wechsler measures, introduces an important consideration in the clinical use of intelligence measures for disability determination.  Current scientific conceptions of intelligence focus primarily on fluid and crystallized abilities, with recognition that working or comprehensive memory is also important to overall intellectual functioning.  Many intelligence tests are based on these distinctions.  The Wechsler measures are also moving in this direction, with a focus on factor scores that are analogous to crystallized intelligence (e.g., verbal comprehension index), fluid intelligence (e.g., perceptual organization index), and working/comprehensive memory (e.g., working memory index).  Consequently, the committee has recommended continued use of part scores in eligibility determination, but is advocating use of part scores that are consistent with current scientific thinking. (p.6)

During the next decade, even greater alignment of intelligence tests and the IQ scores derived from them and the Horn-Cattell and Carroll models is likely.  As a result, the future will almost certainly see greater reliance on part scores, such as IQ scores for Gc and Gf, in addition to the traditional composite IQ.  That is, the traditional composite IQ may not be dropped, but greater emphasis will be placed on part scores than has been the case in the past.  As this movement to part scores develops, it will most likely occur first for Gc and Gf, the most central of the second-stratum factors, and then extend to other second-stratum dimensions as they are determined to be useful for differential prediction. (p.94)




Friday, October 2, 2009

Science vs law in evaluating expert scientific testimony: From Vidal (2007) Atkins decision

Science vs law in the court room.

I've been skimming the CA Supreme Court Vidal (2007) Atkins-related decision and found a fascinating discussion of the distinction between science and law (thanks to In the News Blog for directing my focus to this section).  For those who don't want to read the entire PDF document previously posted, below is the relevant text.  Professionals who testify (or who are considering testifying) in Atkins cases, should be aware of the courts role in mediating/deciding expert-testimony based scientific debates.  Scientific debates in the court room are not the same as debates between scholars at conferences, in journal articles, etc.

Underlined/italics in the text below reflect the blogmasters emphasis.
In assessing the role the Full Scale IQ score (or any other single test score) plays in determining mental retardation, we must distinguish between rules of law and diagnostic criteria of psychology. The expert testimony below included a vigorous scientific debate as to whether Vidal’s Full Scale IQ scores should rule out a diagnosis of mental retardation. While one psychologist, McKinzey, gave his opinion that Full Scale IQ scores are, in all circumstances, the “best measure of general intelligence,” two other psychologists, Couture and Widaman, testified that where testing showed an extraordinarily wide divergence between Performance and Verbal IQ scores, the Full Scale measure was not a fully reliable measure. In support of their views, both sides gave scientific, not legal, reasons and cited scientific, not legal, authority

The Court of Appeal sided squarely with McKinzey in this debate over psychological standards, stating flatly that “general intellectual functioning is primarily determined by the defendant’s FSIQ score.” Like the psychologists who testified at the hearing, the lower court majority cited scientific sources (references published by the American Psychiatric Association and the American Association on Mental Retardation) rather than legal authority in support of its view. The Court of Appeal majority erred in thus purporting to resolve a factual question--the best scientific measure of intellectual functioning--as a matter of law. In finding the facts of a particular case, courts and juries untrained in science are sometimes called upon to resolve contested scientific issues, but such factual findings do not establish generally applicable rules of law. The superior court here, for example, found on the basis of Couture’s and Widaman’s testimony that in Vidal’s case his Full Scale IQ scores in the low average to average range did not preclude a finding of mental retardation. In a given case an appellate court might, within its proper role, hold that such a finding was not supported by substantial evidence in the hearing record. But an appellate court cannot convert a disputed factual assertion into a rule of law simply by labeling it a “legal standard,” as the Court of Appeal purported to do here.

Courts also must sometimes evaluate disputed scientific assertions in the course of determining the admissibility of expert scientific testimony. In determining the evidentiary reliability of a new scientific technique, California courts look primarily to the technique’s general acceptance in the relevant scientific community, an approach designed to ensure “ ‘that those most qualified to assess the general validity of a scientific method will have the determinative voice.’ ”(People v. Kelly (1976) 17 Cal.3d 24, 31, italics omitted.) Even under the arguably more searching federal court inquiry described in Daubert v. Merrell Dow Pharmaceuticals, Inc. (1993) 509 U.S. 579, “the focus, of course, must be solely on principles and methodology, not on the conclusions that they generate.” (Id. at p. 595.) The courts’ evidentiary gatekeeping function is thus not a warrant for judicial intervention in genuine scientific debates over substantive principles. In any event, we are not faced here with a question of admissibility of disputed evidence but with the question whether, when both sides of a scientific dispute have been presented by expert testimony, an appellate court may declare the debate’s winner as a matter of law.

The Legislature has mandated that trial courts, in determining mental retardation for Atkins purposes (Atkins, supra, 536 U.S. 304), find whether the individual’s “general intellectual functioning” is significantly impaired (§ 1376, subd. (a)), but has not defined that phrase or mandated primacy for any particular measure of intellectual functioning. The question of how best to measure intellectual functioning in a given case is thus one of fact to be resolved in each case on the evidence, not by appellate promulgation of a new legal rule.


Technorati Tags: , , , , , , , , , , , , , , , , ,