Showing posts with label AP101. Show all posts
Showing posts with label AP101. Show all posts

Friday, May 24, 2013

AP101 Brief #1b: g or not to g in Atkins MR death penalty cases? (part b in series)

This is a prior (OBG) post for those who may have missed it the first time.


Applied Psychometrics (AP) 101 Brief #1b:  g or not to g in Atkins MR death penalty cases (second in a series)

If you have not read the first  post in this series, you should read the first post now.  Then return and resume reading.

As described in the first post, the g-loadings (of the tests or composite scores in an IQ battery) on the first principal-component in principal component analysis (PCA) is a traditional index of the g-ness (aka., saturation of general intellectual ability) of a measure.  Furthermore, g-loadings calculated within a specific intelligence battery only tell you the relative g-ness of measures as defined by that specific collection of measures within that particular IQ battery.  It is when one moves to joint-analysis of intelligence batteries (and the more batteries in the analysis the better) that a more accurate picture of a measures g-ness can be determined.

Having access to the mixed LD/normal university young adult sample reported in the WJ III technical manual (conflict of interest note - I'm a coauthor of the WJ III), I just ran a joint PCA on on this sample of 200.  A description of the sample and instruments administered, abstracted from the WJ III technical manual (click here for a brief WJ III technical manual bulletin summary) can be found by clicking here.   I selected this data set as the adult subjects had all been administered the WAIS-III.  In addition, they had also been administered the CHC-based WJ III Tests of Cognitive Ability (WJ III) and the Gf-Gc based Kaufman Adolescent and Adult Intelligence test (KAIT).  I analyzed the composite scores from the three batteries via the PCA procedures conceptually described in the first post in this series.  Below is a summary of the results.

Interestingly, the top four g-measures are from the KAIT (Fluid and Crystallized composites) and the WJ III (Gc=Comprehension-Knowledge; Gf=Fluid Reasoning).  Even more interesting was the finding that when a more broad array of cognitive ability composites are included in a g-analysis, the WAIS-III VC (Verbal Comprehension Index; also classified as a strong measure of Gc as per CHC theory) is only the 9th strongest g-measure, and even falls behind the WAIS-III Perceptual Organization (PO; primarily Gv and some Gf as per CHC theory) and Working Memory Indexes (WM--Gsm as per CHC theory). 

A hypothesis that has been advanced to explain the differences in how Gc/verbal abilities are measured by the Wechsler verbal scales and Gc/verbal abilities on other cognitive batteries is grounded in Jim Cummins distinction between two types of language proficiency---BICS (Basic Interpersonal Communication Skills) and CALPS (Cognitive Academic Language Proficiency; click here for additional on-line information).  Briefly, BICS is language proficiency in more contextualized everyday language contexts while CALP is the more context-reduced conceptual-linguistic knowledge that occures in a context of semantics, and abstractions.  CALP is more cognitively demanding.  Anyone familiar with the Wechsler verbal subtests of Vocabulary, Comprehension and Similarities knows that they allow for subjects to provide lengthy verbal responses in their own everyday language.  In contrast, verbal items on other IQ tests (e.g., WJ III), require one-word responses and tend to focus more on cognitive processing involving language (e.g., antonyms, synonyms, verbal analogies).  It has been hypothesized that the Wechsler verbal tests and scales are more BICS-influenced while other IQ tests tend to use verbal/Gc test formats that require more CALP.  This might explain the findings reported above--that the WAIS III Verbal Comprehension Index is less cognitively demanding that the Gc/Verbal scales from the WJ III and KAIT.

Why is this finding important?  Because in a number of Atkins decisions, where concerns were raised about the person's WAIS-III Full Scale score being a good estimate of the persons g-ness (general intelligence), considerable stock was placed in the Verbal IQ (of which the Verbal Comprehension Index is now a purer factor measure) as being the best estimate of the person's general intelligence (see posts re: Maldonado and Vidal decisions)

Lets also examine these same data through the lens of multidimensional scaling analysis (specifically, Guttman's Radex Model).  The radex model statistically classifies measures as per two dimensions--cognitive complexity and stimulus content.  As noted in the prior post in this series, cognitively complexity is considered an index of general intelligence (g). For interested readers, a classic article on the use of the MDS radex model in the analysis of intelligence measure was published in Intelligence in 1983 (Marshalek, Lohman & Snow, 1983). Below is a visual-spatial representation of the MDS results in the current university sample, using the same composite measures as reported above (in the PCA g-loading analysis).  The interpretations below are mine.





The two broad cognitive processing continuim interpretaions (X- and Y-axis) are not the focal point of the current discussion.  The most critical finding in the current context, as per the radex model, is how close to the center of the figure a composite measure score is placed.  Measures that are the closest to the center are considered the most cognitively complex.  As measures move further away from the center, they are judged to be less cognitively complex

The results, although using a different method than PCA, produce the same conclusions.  The most cognitively complex measure in this sample (which could be thus interpreted as the best index of cognitive complexity or g-ness) is the KAIT Fluid Intelligence Scale.  The next closest to the center of the figure are the WJ III Gc (Comprehension-Knowledge) and Gf (Fluid Reasoning) clusters.  Again of interest is the location of the WAIS VC composite...it is much less cognitively complex than these other measures, and interestingly, much less cognitively complex than the two similar measures of Gc abilities (WJ III Gc composite; KAIT Crystallized Intelligence composite).  I've also provided my stimulus content hypothesis interpretations of groupings of composties (designated by ovals) from across the batteries (e.g., Processing Speed- WAIS Processing Speed and WJ III Gs or Processing Speed).

The findings in this one sample (which therefore warrants caution in generalization), suggests that the WAIS-III Verbal Comprehension composite, which is the most valid measure of Gc or verbal abilities on the WAIS-III, may NOT be all it is thought to be--when it comes to tapping cognitively complex cognitive processing.  Other Gc or verbal measures from intelligence batteries with adult norms (KAIT; WJ III) were found, in a relative sense, to be much better indicators of a person's g-ness (general intelligence).  The data suggest that, in this sample, even the WAIS-III Working Memory Index score may be a better relative proxy for g-ness than the WAIS verbal composite.

These analyses raise interesting questions about Atkins decisions that have relied either exclusively on the WAIS-R/WAIS-III scores, particularly when part scores (Verbal IQ, Verbal Comprehension; Peformance IQ; Perceptual Organization; etc.) are used instead of the Full-Scale IQ to determine mental retardation, or when the respective WAIS verbal composite is considered better than other potential test global IQ scores (or similar Gc/verbalcomposite scores from other batteries) in making a determination of level of general intellectual functioning.

How can this be?  How can a major scale from the "gold standard" of IQ tests (as it is commonly called in Atkins decisions) be a poorer estimate of general intelligence (g-ness) than most psychologists think?  More importantly, what are the implications for Atkins decisions, when the WAIS-R/III Full Scale score has been questioned as an accurate g-estimate in the face of considerable profile variability, and then the Verbal IQ/Verbal Comprehension Index is used to estimate g-ness (general intelligence)?

In both the Maldonado and Vidal decisions considerable stock was placed on the respective Wechsler verbal composite scores as being the best indicator of general intelligence (for making a determination of mental retardation).  In Maldonado, the reliance of the Wechsler verbal composite trumped a more comprehensive CHC-based IQ battery (BAT-R) administered in Maldonado's native language (Spanish).  Of concern in the Vidal decision, is that he had been administered four different versions of the Wechsler IQ batteries (over many decades), and they consistently revealed a large verbal/nonverbal (performance) IQ split.  Thus, arguments hinged extensively on the Verbal IQ vs the Full Scale IQ.  I'm perplexed why the experts in intelligence and intelligence testing, even without knowing the result of the above g-analysis, did not say "we've got consistent Wechsler V-P split information, I think it would be important to administer more contemporary intelligence tests, or parts of some of these batteries, to find out more information about important g-related cognitive abilities (for the defendant) not measured by the WAIS battery."  The Wechsler batteries had consistently captured Vidals abilities as measured per that battery--wouldn't time have been better spent, and a decision made on a higher quality array of cognitive information, by requesting administration of other IQ tests (or parts of other IQ tests) instead of arguing over old and consistent limited cognitive data? In fact, I, together with Flanagan and Ortiz, published a book in 2000 (The Wechsler Intelligence Scales and Gf-Gc theory:  A contemporary approach to interpretation) that presents procedures for augmenting the various Wechsler batteries to provide for a more comprehensive CHC/Gf-Gc based assessment of a persons intellectual functioning.  This information was also available as early as 1998 (see ITDR by McGrew and Flanagan).  [conflict of interest note - I coauthored these two books which made little in the way of ching-$ for the authors.  They are now both not being printed and none of the authors are receiving any royalties from their sales].

I continue to be baffled/troubled by the over-reliance, and almost god-like stature of the various versions of the WAIS (R/III/IV), in Atkins rulings.  It has been well known (and written about in articles and books; click here) since the early 1990's, that contemporary CHC (aka, Gf-Gc) theory had emerged as the consensus model of intelligence and, more importantly, instruments had been designed (with adult norms) to measure many of the unmeasured or poorly measured CHC abilities not taped by the WAIS-R/III batteries.  If I was an attorney arguing an Atkins case, on either side of the fence, I would seek intellectual testing beyond the so-called "gold standard."  I would want the best possible estimate of g-ness (since this seems to be the crux of the first prong of MR determination in most Atkins cases).

IMHO, the major problem is that of the "inertia of tradition" in intelligence testing, particularly in psychology disciplines that deal with adult populations.  Many practicing psychologists, esp. those working in adult settings whose professional associations and journals have paid less attention to contemporary intelligence theory and test development (less than school and educational psychologists), simply have not kept abreast of these developments. 

How long will Atkins expert intelligence testimony, expert debates, and decisions be made in the face of the Atkins MR IQ Theory-Test gap?  Isn't this simply wrong?  Professionally and ethically shouldn't psychologists who offer judgements in life-or-death decisions hinging on IQ test results be "up to speed" regarding contemporary intelligence theory and instruments?  Should the courts continue to handicapped by the presentation of intelligence test results that are not based on the best evidence from intelligence theory, research, and test development?  The courts are at the mercy of experts who testify, experts who I believe need to be familiar with the cutting edge empirical and theoretical information on the structure of human intelligence and various IQ batteries that are available, beyond the Wechslers.

Given the data presented above, it is possible that the decisions in at least two cases (and I'm sure there are more), may have had a different outcome, or at least an outcome based on a more comprehensive set of intelligence information.  Justice could have been better served via more contemporary intellectual testing practice and interpretation.

Technorati Tags: , , , , , , , , , , , , , , , , ,


Monday, April 30, 2012

IAP AP101 Report # 13: Problems with the 1960 and 1986 Stanford-Binet IQ Scores in Atkins MR/ID Death Penalty Cases




Often in Atkins MR/ID death penalty cases historical and contemporary IQ scores are available for review by psychological experts.  In many cases these scores vary markedly.  The courts frequently wrestle with the issue of determining what the best estimate is of the person’s general intelligence.  A review of many Atkins cases often reveals frequent mention of two “gold standard” IQ tests in reports or testimony—namely, the Stanford-Binet and the Wechsler series.

The purpose of this working paper is to alert psychologists and the courts to two little known (but extremely important) dents in the gold standard status of two versions of the Stanford-Binet—the 1960 SB and the 1986 SB IV. If a Flynn effect adjustment is made to scores from a 1960 SB, the norm date used to calculate the magnitude of the Flynn effect should be 1932…not 1960.  If SB IV scores exist in an individual’s records, experts providing opinions regarding the individual’s general level of intelligence should consider: (a) eliminating the score from consideration, (b) not give the score great weight in formulating an opinion, or (c) at a minimum, provide qualifying statements regarding the validity of the SB IV score as required by the Joint Test Standards.

IAP Applied Psychometrics 101 Report # 13 can be downloaded by clicking here.

Thursday, March 1, 2012

IAP101 Brief #12: Use of IQ component part scores as indicators of general intelligence in SLD and MR/ID diagnosis

   
            Historically the concept of general intelligence (g), as operationalized by intelligence test battery global full scale IQ scores, has been central to the definition and classification of individuals with a specific learning disability (SLD) as well as individuals with an intellectual disability (ID).  More recently, contemporary definitions and operational criteria have elevated intelligence test battery composite or part scores to a more prominent role in diagnosis and classification of SLD and more recently in ID.
            In the case of SLD, third-method consistency definitions prominently feature component or part scores in (a) the identification of consistency between low achievement and relevant cognitive abilities or processing disorders and (b) the requirement that an individual demonstrate relative cognitive and achievement strengths (see Flanagan, Fiorello & Ortiz, 2010).  The global IQ score is de-emphasized in the third-method SLD methods.
            In contrast, the 11th edition of the AAIDD Intellectual Disability: Definition, Classification, and Systems of Supports manual (AAIDD, 2010) placed general intelligence, and thus global composite IQ scores, as central to the definition of intellectual functioning.  This has not been without challenge.  For example, the AAIDD ID definition has been criticized for an over-reliance on the construct of general intelligence and for ignoring contemporary psychometric theoretical and empirical research that has converged on a multidimensional hierarchical model of intelligence (viz., Cattell-Horn-Carroll or CHC theory).
The potential constraints of the “ID-as-a-general-intelligence-disability” definition was anticipated by the Committee on Disability Determination for Mental Retardation, in its National Research Council report “Mental Retardation:  Determining Eligibility for Social Security Benefits” (Reschly, Meyers & Hartel, 2001).  This national committee of experts concluded that “during the next decade, even greater alignment of intelligence tests and the IQ scores derived from them and the Horn-Cattell and Carroll models is likely.  As a result, the future will almost certainly see greater reliance on part scores, such as IQ scores for Gc and Gf, in addition to the traditional composite IQ.  That is, the traditional composite IQ may not be dropped, but greater emphasis will be placed on part scores than has been the case in the past” (Reschly et al., 2002, p. 94).  The committee stated that “whenever the validity of one or more part scores (subtests, scales) is questioned, examiners must also question whether the test’s total score is appropriate for guiding diagnostic decision making.  The total test score is usually considered the best estimate of a client’s overall intellectual functioning.  However, there are instances in which, and individuals for whom, the total test score may not be the best representation of overall cognitive functioning.” (p. 106-107).
            The increased emphasis on intelligence test battery composite part scores in SLD and ID diagnosis and classification raises a number of measurement and conceptual issues (Reschly et al., 2002).  For example, what are statistically significant differences?  What is a meaningful difference?  What appropriate cognitive abilities should serve as proxies of general intelligence when the global IQ is questioned?  What should be the magnitude of the total test score? 
Appropriate cognitive abilities will only be the only issue discussed here.  This issue addresses  which component or part scores are more correlated with general intelligence (g)—that is, what component part scores are high g-loaders?  The traditional consensus has been that measures of Gc (crystallized intelligence; comprehension-knowledge) and Gf (fluid intelligence or reasoning) are the highest g-loading measures and constructs and are the most likely candidates for elevated status when diagnosing ID (Reschly et al., 2002).  Although not always stated explicitly, the third method consistency SLD definitions specify that an individual must demonstrate “at least an average level of general cognitive ability or intelligence” (Flanagan et al., 2010, p.745), a statement that implicitly suggests cognitive abilities and component scores with high g-ness.
Table 1 is intended to provide guidance when using component part scores in the diagnosis and classification of SLD and ID (click on images to enlarge and use the browser zoom feature  to view; it is recommended you click here to access a PDF copy of the table..and also zoom in on it).  Table 1 presents a summary of the comprehensive, nationally normed, individually administered intelligence batteries that possess satisfactory psychometric characteristics (i.e., national norm samples, adequate reliability and validity for the composite g-score) for use in the diagnosis of ID and SLD.



The Composite g-score column lists the global general intelligence score provided by each intelligence battery.  This score is the best estimate of a persons general intellectual ability, which currently is most relevant to the diagnosis of ID as per AAIDD.  All composite g-scores listed in Table 1 meet Jensens (1998) psychometric sampling error criteria as valid estimates of general intelligence.  As per Jensens number of tests criterion, all intelligence batteries g-composites are based on a minimum of nine tests that sample at least three primary cognitive ability domains.  As per Jensens variety of tests criterion (i.e., information content, skills and demands for a variety of mental operations), the batteries, when viewed from the perspective of CHC theory, vary in ability domain coveragefour (CAS, SB5), five (KABC-II, WISC-IV, WAIS-IV), six (DAS-II) and seven (WJ III) (Flanagan, Ortiz & Alfonso, 2007; Keith & Reynolds, 2010).   As recommended by Jensen (1998), the particular collection of tests used to estimate g should come as close as possible, with some limited number of tests, to being a representative sample of all types of mental tests, and the various kinds of test should be represented as equally as possible (p. 85).  Users should consult sources such as Flanagan et al. (2007) and Keith and Reynolds, 2010) to determine how each intelligence battery approximates Jensens optimal design criterion, the specific CHC domains measured, and the proportional representation of the CHC domains in each batteries composite g-score.
Also included in Table 1 are the component part scales provided by each battery (e.g., WAIS-IV Verbal Comprehension Index, Perceptual Reasoning Index, Working Memory Index, and Processing Speed Index), followed by their respective within-battery g-loadings.[1]  Examination of the g-ness of composite scores from existing batteries (see last three columns in Table 1) suggests the traditional assumption that measures of Gf and Gc are the best proxies of general intelligence may not hold across all intelligence batteries.[2] 
In the case of the SB5, all five composite part scores are very similar in g-loadings (h2 = .72 to .79).  No single SB5 composite part score appears better than the other SB5 scores for suggesting average general intelligence (when the global IQ score is not used for this purpose).  At the other extreme is the WJ III where the Fluid Reasoning, Comprehension-Knowledge, Long-term Storage and Retrieval cluster scores are the best g-proxies for part-score based interpretation within the WJ III.  The WJ III Visual Processing and Processing Speed clusters are not composite part scores that should be emphasized as indicators of general intelligence.  Across all batteries that include a processing speed component part score (DAS-II, WAIS-IV, WISC-IV, WJ III) the respective processing speed scale is always the weakest proxy for general intelligence and thus, would not be viewed as a good estimate of general intelligence. 
            It is also clear that one cannot assume that composites with similar sounding names of measured abilities should have similar relative g-ness status within different batteries.  For example, the Gv (visual-spatial or visual processing) clusters in the DAS-II (Spatial Ability), SB5 (Visual-Spatial Processing) are relatively strong g-measures within their respective battery, but the same cannot be said for the WJ III Visual Processing cluster.  Even more interesting are the differences in the WAIS-IV and WISC-IV relative g-loadings for similarly sounding index scores. 
For example, the Working Memory Index is the highest g-loading component part score (tied with Perceptual Reasoning Index) in the WAIS-IV but is only third (out of four) in the WISC-IV.   The Working Memory Index is comprised of the Digit Span and Arithmetic subtests in the WAIS-IV and the Digit Span and the Letter-Number Sequencing subtests in the WISC-IV.  The Arithmetic subtest has been reported to be a factorially complex test which may tap fluid intelligence (Gf-RQ—quantitative reasoning), quantitative knowledge (Gq), working memory (Gsm), and possible processing speed (Gs; Keith & Reynolds, 2010; Phelps, McGrew, Knopik & Ford, 2005).   The factorially complex characteristics of the Arithmetic subtest (which, in essence, makes it function like a mini-g proxy) would explain why the WAIS-IV Working Memory Index is a good proxy for g in the WAIS-IV but not in the WISC-IV. The WAIS-IV and WISC-IV Working Memory Index scales, although named the same, are not measuring identical constructs.

A critical caveat is that the g-loadings cannot be compared across different batteries.  g-loadings may change when the mixture of measures included in the analyses change.  Different "flavors" of g can result (Carroll, 1993; Jensen, 1998). The only way to compare the g-ness across batteries is with appropriately designed cross- or joint-battery analysis (e.g., WAIS-IV, SB5 and WJ III analyzed in a common sample).
The above within and across intelligence battery examples illustrates that those who use component part scores as an estimate of a person’s general intelligence must be aware of the composition and psychometric g-ness of the component scores within each intelligence battery.  Not all component part scores in different intelligence batteries are created equal (with regard to g-ness).  Also, not all similarly named factor-based composite scores may measure the same identical construct and may vary in degree of within battery g-ness.  This is not a new problem in the context of naming factors in factor analysis, and by extension, factor-based intelligence test composite scores, Cliff (1983) described this nominalistic fallacy in simple language—“if we name something, this does not mean we understand it” (p. 120). 




[1] As noted in the footnotes in Table 1, all composite score g-loadings were computed by Kevin McGrew by entering the smallest number (and largest age ranges covered) of the published correlation matrices within each intelligence batteries technical manual (note the exception for the WJ III) in order to obtain an average g-loading estimate.  It would have been possible to calculate and report these values for each age-differentiated correlation matrix for each intelligence battery.  However, the purpose of this table is to provide the best possible average value across the entire age-range of each intelligence battery.  Floyd and colleagues have published age-differentiated g-loadings for the DAS-II and WJ III.  Those values were not used as they are based on the use of the principal common factor analysis method, a method that  analyzes the reliable shared variance among tests.  Although principal factor and principal component loadings typically will order measures in the same relative position, the principal factor loadings typically will be lower.  Given that the imperfect manifest composite scale scores are those that are utilized in practice, and to also allow uniformity in the calculation of the g-loadings reported in Table 1, principal component analysis was used in this work. The same rationale was used for not using the latent factor loadings on a higher-order g-factor in SEM/CFA analysis of each test battery.  Loadings from CFA analyses represent the relations between the underlying theoretical ability constructs and g purged of measurement error.  Also, frequently the final CFA solutions reported in a batteries technical manual (or independent journal articles) allow tests to be factorially complex (load on more than one latent factor), a measurement model that does not resemble the real world reality of the manifest/observed composite scores used in practice.  Latent factor loadings on a higher-order g-factor will often differ significantly from principal component loadings based on the manifest measures, both in absolute magnitude and relative size (e.g., see high Ga loading on g in WJ III technical manual which is at variance with the manifest variable based Ga loading reported in Table 1) 
[2] The h2 values are the values that should be used to compare the relative amount of g-variance present in the component part scores within each intelligence battery.

Tuesday, February 7, 2012

IAP Applied Psychometrics 101 Brief Report # 11: What is the typical IQ and adaptive behavior correlation?


What is the typical relation (correlation) between standardized measures of adaptive behavior (AB)  and measures of intelligence (IQ)?  This is an important question given the role both play in the definition diagnosis of mental retardation (MR) / intellectual disability (ID). 

During the late 1970's and 1980's this was an active area of research.  Numerous studies were published that reported correlations between a wide variety of adaptive behavior scales and intelligence tests.  Probably the best synthesis of this research was provided by Harrison (1987).  Harrison's review included a table of over 40+ correlations.  This is Table 2 in the above referenced and linked article.  Harrison concluded, as have most others who have reviewed the literature, that "the majority of correlations fall in the moderate range" (p.39).  When the correlations with maladaptive measures are excluded from Harrison's table, the correlations range from .03 to .91.  This is a wide range.  Harrison could not identify a specific explanation for the variability or range of correlations.  Harrison speculated that variables might impact the magnitude of the correlations were the specific adaptive behavior or measure of intelligence used and differences in sample variability.

Subsequently the Committee on Disability Determination for Mental Retardation published a National Research Council report (Mental Retardation:  Determining Eligibility for Social Security Benefits; Reschly, Meyers & Hartel, 2001) that also addressed the AB/IQ relation. The report concluded that AB/IQ studies report correlations "ranging from 0 (indicating no relationship) to almost +1 (indicating a perfect relationship).  Data also suggest that the relationship between IQ and adaptive behavior varies significantly by age and levels of retardation, being strongest in the severe and moderate ranges and weakest in the mild range.  There is a dearth of data on the relationship of IQ and adaptive behavior functioning at the mild level of retardation" (p. 8).  Factors identified as moderating the AB/IQ correlation were scale content, measurement of competences versus perceptions, sample variability, ceiling and floor problems of the scales, and level of mental retardation.

Given the above, it is hard to render an objective statement on the approximate typical AB/IQ correlation.  With this in mind, an informal research synthesis was completed and is reported here.

First, only the AB/IQ correlations (IQ/maladaptive correlations were excluded) from Harrison's 1987 table were extracted (n = 43 correlations).  Then, the technical manuals for the current editions of the three most frequently used contemporary adaptive behavior scales were reviewed for additional correlations.  This included the Vineland Adaptive Behavior Scale (Sparrow, Cicchetti & Balla, 2005; n = 2 correlations of .12, .20) and the Adaptive Behavior Scales--II (Harrison & Oakland, 2008; n = 10 correlations ranging from .39 to .67; median = .51).

Although six different correlations were reported in the Scales of Independent Behavior-Revised manual (SIB-R; Bruininks, Woodcock, Weatherman & Hill, 1996), the values were not used as they are inflated estimates when compared to the type of correlations typically reported.  For example, very high correlations of .79, .82 and .91 are reported for certain groups.  A close reading of the tables reveals that the SIB-R correlations with either the WJ or WJ-R intelligence test were calculated on the basis of the W-score growth metric.   By definition, a growth metric includes age variance.  If correlations are reported across wide age groups the correlations convey variance related to the correlation between the AB  and IQ constructs but also contains shared variance due to the influence of general age-base development (age).  Thus, the SIB and SIB-R correlations with IQ, although not wrong and providing different information, are not comparable to all other reported correlations where age variance has been removed (typically by correlating age-based standard scores).  Clear evidence for this point comes from McGrew and Bruininks (1990) who used the same SIB/WJ subject data reported in the SIB and SIB-R manuals, but who removed the W-score confounded age variance prior to the calculation of latent factor correlations (via confirmatory factor analysis) between latent practical intelligence (SIB adaptive behavior) and conceptual intelligence (WJ IQ) factors.  The resulting AB/IQ correlations for three different age groups were .38, .56 and .58--far below the values in the .70 to .92 range.  Thus, the values from McGrew and Bruininks (1990) were included for estimates of the SIB/SIB-R IQ correlations in the current synthesis. 

Finally, latent AB/IQ correlations (as estimated from confirmatory factor analysis models)  of .27 and .39 were included from Ittenbach, Spiegel, McGrew and Bruininks (1992) and Keith, Fehrmann,Harrison and Pottebaum (1987), respectively.  This process resulted in the addition of 17 AB/IQ correlations to the 43 from Harrison, for a total of 60 correlations.

Descriptive statistics for this collection of 60 AB/IQ correlations are as follows: range of correlations from .12 to .90,  a mean of .51 and a median of .48, and a standard deviation of .20.  Below is a figure that includes a frequency polygon (and smoothed normal curve overlay) and a box-whisker plot of the data set.  A review of the box and whisker plot (at the bottom) shows the median correlation (.48) as a vertical line within the rectangle.  The rectangle includes the 50% middle of the distributions of correlations and shows an approximate range of just below .40 to just above .65.  Of particular note is the shape of the frequency polygon and smoothed normal curve.  The shape of the frequency polygon is consistent with a normal curve.  In quantitative research synthesis this type of normal distribution suggests that total data set included in the review is not biased--both studies that are likely under- or overestimates of the "true" population correlation (due to method or sampling factors) are included.  More importantly, the "bunching" up of the majority of the correlations in the middle provide confidence that the median of this distribution is a reasonable unbiased estimate of the populaiton correaltion.  This type of relatively normal distribution suggests that the current collection of 60 AB/IQ correlations is likely a reasonable approximation of the complete set of population AB/IQ correlations.


Based on this informal (and admittedly incomplete review of all possible AB/IQ correlation research) one can conclude that a reasonable estimate of the typical AB/IQ correlation is approximately .50 (mean = .51; median = .48), with most ranging from approximately .40 to .65.  This finding is consistent with Harrison's 1987 conclusion of a "moderate" correlation.  The current analysis continues to reinforce Harrison's (and others) conclusions that adaptive behavior and intelligence are statistically related constructs, but  they are still independent.   An average correlation of .50 indicates that AB and IQ share approximately 25 % common variance (approximately 15% to 40 % common variance if one looks at the range of the 50% middle of the distribution of values).  In practical terms this means that for any individual, standard scores from AB and IQ tests will frequently diverge and not always be consistent.  

Harrison (1987) provides a nice explanation for the primary reasons for the moderate correlation between AB and IQ.  Her quote is reproduced below
Numerous caveats need to be applied to this analysis and report.  The most important are:
  • A comprehensive review of all possible published and unpublished AB/IQ research studies was not completed.  Clearly there are more studies "out there" that could be added to the synthesis. 
  • The analysis makes no attempt to determine if there are moderator effects.  That is, is the typical correlation likely to systematically vary as a function of AB measures, IQ measures, variability in the sample's level of functioning, manifest/measured versus latent variable correlations, level of ability, etc.? 
  •  This has not been peer reviewed.


 It is hoped that this ad hoc update of Harrison's (1987) review, augmented by quantitive organizational methods, will serve to stimulate a formal meta-analysis by others (hint---a nice study or thesis for someone?)




Thursday, September 22, 2011

IAP AP101 Brief # 10: Understanding IQ score differences: Examiner Errors


Why do significant differences in IQ scores often occur between different tests or the same test given at different times? The explanations are many. Previous IAP Applied Psychometric 101 Reports and Briefs have touched on a number of reasons. Click here to view or link to these reports.

In the first AP101 report, which I would recommend reading prior to reading the material below, test administration and or scoring errors (examiner errors) were mentioned as a possible reason for score discrepancies. The brief report below addresses this topic.


Test procedural and administration errors (examiner error)

Despite rigorous graduate training in standardized administration of intelligence tests for most psychologists, the extant research on adherence to standardized administration and scoring procedures has consistently reported (unfortunately) that the frequency of examiner errors occurs with enough regularity, for both novice and experienced psychological examiners, to be a concern.

Ramos, Alfonso and Schermerhorn (2009) summarized the extant research on examiner errors and reported that most research studies reported sufficient average examiner error to produce significant changes in IQ scores for individuals. The most frequent types of errors reported included a failure to record responses, use of incorrect basal and ceiling rules, reporting an incorrect global IQ score, incorrect adding of subtest scores, incorrect assignment of points for specific items, and incorrect calculation of the individuals age. On Wechsler-related studies, Ramos et al.'s review found that studies have reported average error rates from 7.8 to 25.8 errors per test record, almost 90% of examiners making one error, and in one study 2/3 of the test records reviewed resulted in a change in the Full Scale IQ. Examiner errors do not appear instrument specific as Ramos et al’s reported an average error rate of 4.63 errors per test record on the WJ III Tests of Cognitive Abilities.

The importance of verifying accurate administration and scoring is evident in the finding that across experienced psychologists and students in graduate training, ranges of score differences were as high as 25, 22, and 11 points respectively for the WAIS-III Verbal, Performance and Full Scale IQ scores (Ryan & Schnakenberg-Ott, 2003). Despite examiners reporting confidence in their scoring accuracy, Ryan and Schnakenberg-Ott reported average levels of agreement with the standard (accurate) test record of only 26.3% (Verbal IQ), 36.8 % (Performance IQ), and 42.1 % (Full Scale IQ).

This level of examiner error is alarming, particularly in the context of important decision-making (e.g., IQ score-based life-and-death Atkins MR/ID decisions; eligibility for intervention programs; eligibility for social security disability funds). The level of examiner experience does not appear to be an explanatory variable. More recently, when investigating a single subtest (WISC-IV Vocabulary), Erodi, Richard and Hopwood (2009) reported that more errors may be present when evaluating low and high ability subjects.

Numerous test development and professional training and monitoring recommendations have been suggested (see Erodi et al, 2009; Hopwood & Richard, 2005; Kuentzel et al. 2011; Ramos et al., 2009; Ryad & Schnakenberg-Ott, 2003), some that have empirically demonstrated improvement in accuracy (see Kuentzel, Hetterscheidt &Barnett, 2011).

Examiner test administration and scoring errors can be the reason for discrepant IQ-IQ score differences. It is clear that before attempting to interpret any IQ scores, or trying to reconcile IQ-IQ score differences between tests, the first step would be for all examiners to double check their scoring. Another wise step would be to seek independent review of a scored test record by another experienced examiner. In the case of Atkins decisions, attempts should be made to secure copies of the original IQ test records for independent review. If any clear errors are present, they should be corrected and new scores recalculated. Only then should psychologists proceed to draw conclusions about the consistency or differences between scores from different IQ tests or versions of the same test given at different times during an individual’s life-span.

Any intelligence test results used in an Atkin’s hearings must be subject to independent review of the original test protocol (this may be impossible for old historical testing results) to insure against administration or scoring errors that might result in significant differences in the reported IQ score. This is critically important in Atkin’s cases were the courts often use a strict specific-IQ “bright line” cut score to determine the presence of an intellectual disability.

Below are the abstracts from the primary sources for this brief report. Double click on the images to enlarge.
























- iPost using BlogPress from Kevin McGrew's iPad


Generated by: Tag Generator

Wednesday, April 7, 2010

Psychometric PS to Johnston v Florida (2010) denied appeal re: new WAIS-IV scores

This is a follow-up to my brief comments yesterday regarding the Johstone v Fl (2010) denied MR/ID appeal of two days ago.

As mentioned in the decision and my blog comment, the WAIS-III/WAIS-IV tests correlated .94 in a study reported in the WAIS-IV technical manual.  This is a very high correlation...but does NOT mean that the two tests should be expected to provide identical IQ scores.  I discuss these issues in a prior IAP AP101 report.

The tests have different norm dates and thus, the later version (WAIS-IV) would be expected to provide a lower score based on the Flynn effect.  More importantly, as reported in the IAP AP101 report, when one calculates the standard deviation of the difference score (see page 6 of that report) for a correlation of .94, the resulting value is 5.2 (round to 5 for ease of discussion).  This means that, on average, the WAIS-III/WAIS-IV (even if highly correlated at the .94 level) would in the general population be expected to display a range of difference scores from -5 to +5...or a range of 10 IQ points......in 68% of the population.  Please review that prior report for further explanation and discussion.

Technorati Tags: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ,

Monday, March 8, 2010

Why do IQ scores often differ? Check out IQ's Corner "IQ Test CHC DNA Fingerprint" analyses

Understanding IQ score differences possibly related to differential CHC content difference coverage.  See ICDP blog post where it is explained and illustrated.  where new feature is introduced - IQ Test DNA CHC Fingerprints

[Double click on image to enlarge]

Tuesday, January 12, 2010

IAP AP101 Brief # 5:The Wechsler-like IQ subtest scaled score metric: The potential for misuse, misinterpretation and impact on critical life decisions

This is a revised post of a previous post (which has now been deleted).  The earlier post indicated that the report brief described below was in draft form---and I was seeking feedback and comments.  A number of individuals did provide some constructive feedback.  As a result, I revised the report (only slightly) and have posted the final version at the link mentioned below.  Thanks for the feedback.  This is now listed under the IAP AP101 Brief section of the blog sidebar.

Below are the introductory paragraphs to IAP AP Brief #5.  The complete report is available for online viewing or downloading by clicking here.  Enjoy.





I've recently been skimming James Flynn's new book (What is Intelligence:  Beyond the Flynn Effect) to better understand the methodology and interpretation of the Flynn effect. Of particular interest to me (as an applied measurement person) is his analysis of the individual subtest scores from the various Wechsler scales across time. As most psychologists know, Wechsler subtest scaled scores (ss) are on a scale with a mean (M) = 10 and a standard deviation (SD) = 3. The subtest ss range from 1 to 19.  In Appendix 1 of his book, Flynn states "it is customary to score subtests on a scale in which the SD is 3, as opposed to IQ scores which are scaled with SD set at 15. To convert to IQ, just multiply subtest gains by five, as was done to get the IQ gains in the last column."  At first glance, this statement makes it sound as if the transformation of subtest ss to IQ SS is an easy (“just multiply….”; emphasis added by me) and mathematically acceptable procedure without problems. However, on close inspection this transformation has the potential to introduce unknown sources of error into the precision of the transformed SS scores.  It is the goal of this brief technical post to explain the issues involved when making this ss-to- IQ SS conversion.

The ss 1-19 scale has a long history in the Wechsler batteries. For sample, in Appendix 1 of Measurement of Adult Intelligence (Wechsler, 1944), Wechsler described the steps used to translate subtest raw scores to the new ss metric. The Wechsler batteries have continued this tradition in each new revision, although the methodology and procedures to calculate the ss 1-19 values have become more sophisticated over time.   Although the methods used to develop the Wechsler ss 1-19 scale may have become more sophisticated, the resultant underlying scale for each subtest has not…scores still range from 1-19 (M=10; SD=3).  Also, the most recent Stanford-Binet—5th Edition (SB5; Roid, 2003) and Kaufman Assessment Battery for Children-2nd Edition (KABC-II) have both adopted the same ss 1-19 scale for their respective individual subtests.

Why is this relatively crude (to be defined below) scale metric still used in some intelligence batteries when other contemporary intelligence batteries provide subtest scale metrics with finer measurement resolution?  For example, the DAS-II (Elliott, 2007) places individual test scores on the T-scale (M=50; SD=10), with scores that range from 10-90.  The WJ III (McGrew & Woodcock, 2001) places all test and composite scores on the standard score (SS) metric associated with full scale and composite scores (M=100; SD=15).  The critical question to be asked is “are there advantages or disadvantages to retaining the historical ss 1-19 scale or, are their real advantages to having individual test scales with finer measurement resolution (DAS-II; WJ III)?”


......continued............
(complete report available at links in first paragraph of this post)

[Double click on image to enlarge]












Technorati Tags: , ,, , , , , , , , , , , , , , , , , , , , , , , , , ,




Monday, December 14, 2009

The standard error of measurement (SEM): Explanations and facts for Atkins MR/ID death penalty proceedings: IAP AP101 Report #5



A new IAP Applied Psychometrics 101 report (#5) is now available.  The title of the report and abstract is below.  The report can be downloaded by clicking here.

Applied Psychometrics 101 #4:  The Standard Error of Measurement (SEM):  An Explanation and Facts for "Fact Finders" in Atkins MR/ID death penalty proceedings.

Abstract

The standard error of measurement (SEM) is a professionally accepted and scientifically based measurement concept that allows users of psychological test scores to account for the known degree of imprecision in the scores.  Atkins MR/ID cases almost always involve standardized psychological testing in the domains of intelligence (IQ tests) and adaptive behavior (AB).  Scores from IQ and AB measures are fallible—not perfectly reliable.  This report provides an easy to understand explanation of the psychometric concept of SEM augmented by an example based on real-world data.  The report concludes with 8 SEM facts that “fact finders” should understand and internalize when evaluating psychological test data during legal proceedings--Atkins MR/ID death penalty proceedings in particular.
Here is a visual treat/tease from the report:


All prior IAP AP101 reports can be accessed via the Applied Psychometrics 101 (AP101) Reports section of the blog--on the blog sidebar.

Technorati Tags: , , , , , , , , , , , , , , , , , , , , , , , , , , , ,


Monday, November 9, 2009

What does the WAIS-IV measure? CHC analysis and beyond....IAP AP101 Report #2




I've posted a new IAP Applied Psychometrics 101 Report (#2:  What does the WAIS-IV measure?  CHC analysis and beyond) at ICDPs sister blog, IQs Corner.  The focus is on understanding what the structure of the WAIS-IV tests measure collectively.

I plan a follow-up report or brief for ICDP that will present additional analysis that will augment this first report.  This follow-up will be focused more on the interpretation of the WAIS-IV composite indexes, as I've learned that the reporting of IQ test data in Atkins cases only focuses on composite or full-scale scores...and not individual subtests.  This follow-up report or brief will be posted here at ICDP.

You can access the new report under the APPLIED PSYCHOMETRICS 101 (AP101) REPORTS section of this blog.

Stay tunned.

Technorati Tags: , , , , , , , , , , , , , , , , ,


Thursday, October 8, 2009

AP101 Brief #1a: g or not to g in Atkins MR death penalty cases



Applied Psychometrics (AP) 101 Brief #1a:  g or not to g in Atkins MR death penalty cases (first in a series)

Despite whether one believes that general intelligence (g) exists, or not (e.g., John Horn), and ignoring the search for the essence of g (via elementary cognitive tasks measuring reaction time, temporal processing, etc.) at the level of brain mechanisms (e.g., Jensen's neural efficiency hypothesis), it is clear from a reading of most Atkins IQ MR death penalty cases that psychological experts testifying in these cases [primarily because of the emphasis on a "deficit in general intellectual functioning" as the first prong in MR diagnosis in the courts, as per recognized professional association definitions of mental retardation; APA, AAIDD] often argue for different IQ scores as being more accurate estimates of the persons g-ness (IQ) than others.

For example, both in Davis (2009), and especially in Vidal (2007), major arguments focused on whether the Full Scale IQ score from theWAIS-III/IV was the best index of g-ness (and thus mental retardation or mental capacity), or whether one of the part scores (e.g., Verbal IQ, Performance IQ) should be used as the best estimate of the persons g-ness (due to extreme variability in the part scores). My "g-estimate is better than your g-estimate" appears a fundamental point of contention at the core of many Atkins cases,  given the assumption that mental retardation is a global deficit in intelligence (see guest post by Watson for some alternative thoughts and excellent insights on the global vs modular nature of intelligence),

Then, along comes Maldonado (2009) where the g-ness argument, at one juncture, is based on the belief that the Spanish WAIS-III Verbal IQ, which is best interpreted as a CHC measure of crystallized intelligence (Gc), should take precedence over the BAT-R total composite score that is comprised of Gc and six other broad CHC abilities.

"My g-estimate....your g-estimate......this special "nonverbal" g-estimate is more accurate for this individual....that is not a good g-estimate....etc......" back-and-forth arguments beg for empirical scrutiny.  So....buckle up and lets examine some real data.......in search of g-ness.  This is the introduction to a small series of posts that will eventually examine, with empirical data, the relative g-ness of the "gold standard" (WAIS-III/IV) composite scores that are most often debated in these matters.

But first a definition and some methodological background information.  According to the APA Dictionary of Psychology,  general intelligence (the general factor) is:
  • a hypothetical source of individual differences in GENERAL ABILITY (emphasis in original) , which represents individuals' abilities to perceive relationships and to derive conclusions from them.  The general factor is said to be a basic ability that underlies the performance of different varieties of intellectual tasks, in contrast to SPECIFIC ABILITIES (emphasis in original), which are alleged each to be unique to a single task (p. 403).
[Note - some of the the text below comes from Flanagan, McGrew & Oritz (2000).  The Wechsler Intelligence Scales and Gf-Gc theory.  Boston:  Allyn & Bacon.

Intelligence tests have been interpreted often as reflecting a general mental ability referred to as g (Anastasi & Urbina, 1997; Bracken & Fagan, 1990; Carroll, 1993a; French & Hale, 1990; Horn, 1988; Jensen, 1984, 1998; Kaufman, 1979, 1994; Keith, 1997; Sattler, 1992; Sattler & Ryan, 1999; Thorndike & Lohman, 1990).  The g concept was associated originally with Spearman (1904, 1927) and is considered to represent an underlying general intellectual ability (viz., the apprehension of experience and the eduction of relations) that is the basis for most intelligent behavior. The g concept has been one of the more controversial topics in psychology for decades (French & Hale, 1990; Jensen, 1992, 1998; Kamphaus, 1993; McDermott, Fantuzzo, & Glutting, 1990; McGrew, Flanagan, Keith, & Vanderwood, 1997; Roid & Gyurke, 1991; Zachary, 1990).

According to Arend et al., (2003),  Jensen (1998a, 1998b) proposed that cognitive complexity  might represent a fundamental aspect of g an could be quantified based on inspection of the test measures loadings on the first unrotated factor, because complex tasks show higher factor loadings than simple tasks on that factor.  In many respects when psychologists are discussing mental retardation and general intelligence, there is an implicit assumption that low general intelligence (e.g., mental retardation) is reflected most clearly on performance on the most cognitively complex measures (i.e., high g measures). 

As with the controversy surrounding the nature and meaning of g, disagreements exist about how best to calculate and report psychometric g estimates.  Most all methods are based on some variant of principal component, principal factor, hierarchical factor, or confirmatory factor analysis (Jensen, 1998; Jensen & Weng, 1994).  Although a hierarchical analysis is generally preferred (see Jensen, 1998, p. 86), as long as the number of tests factored is relatively large, the tests have good reliability, a broad range of abilities is represented by the tests, and the sample is heterogeneous, (preferably a large random sample of the general population), the psychometric g's produced by the different methods are typically very similar (Jensen, 1998; Jensen & Weng, 1994).  For the interested reader, Jensen’s (1998) treatise on g (The g Factor) is suggested, as it represents the most comprehensive and contemporary integration of the g related theoretical and research literature.

Operationally the determination of high, moderate or low g-ness of tests or composites has typically been based on each measures correlation (aka., factor or principal component loading) with a single common factor, component, or dimension extracted from the correlations among the set of measures in question.  Measures that "load" high on the g-factor are considered to be the better estimates of general intelligence.

Consider the following simple analogy (which is not original...I borrowed the conceptual idea from Cohen et al., 2006).  You have a special pole that posses a special form of  magnetism (general intelligence). You throw a bunch of  metal marbles (which are the test measures), which have different degrees of the same magnetic force, into a box with the pole at the center.  You gently shake the box.  When you open the box, there is one "king" marble at the top of the poll (it has the highest degree of shared magnetism with the strongest part of the pole), followed next by the next strongest....and so on until the metal marble with the least amount of shared magnetic force is at the bottom.  The pole represents g (general intelligence) and the ordering of the metal marbles (the test measures) represents the ordering of the g-ness (degree of shared magnetic force) of the measures.  The "king" test/marble is assigned the highest numerical index, with each succeeding (and lower) test/marble assigned a slightly lower numerical index of g-ness (shared magnetism).

This is what principal component analysis conceptually accomplishes with a collection of IQ test measures.  It statistically orders the various psychometric measures from strong g-loading to low-g-loading.  This is the typical and traditional statistical currency used by psychometericians and psychologists when discussing the degree of g-ness or g-saturation of different measures--those measures most important for establishing an estimate of a person's general intelligence.


The problem with within-battery factor analysis is that it can affect the g-estimates.  For example, a test’s loading [note- g-loadings are most often computed for the individual subests in a test battery, and not the composite scores such as Verbal IQ, processing speed, etc.-- it is the later, the g-ness of composite scores, which appears to be a critical issue in many Atkins cases.  Thus, when reading the this text I will refer to the measures g...which could mean test or composite] on the general intelligence (g) factor will depend on the specific mixture of measures used in the analysis (Gustafsson & Undheim, 1996; Jensen, 1998; Jensen & Weng, 1994; McGrew, Untiedt, & Flanagan, 1996; Woodcock, 1990).  If a single vocabulary measure is combined with nine visual processing measure, the vocabulary measure will most likely display a relatively low g loading because the general factor will be defined primarily by the visual processing measures.  In contrast, if the vocabulary measure is included in a battery of measure that is an even mixture of verbal and visual processing measures, the loading of the vocabulary measure on the general factor will probably be higher.  It is important to understand that measures g loadings, as typically reported, only reflect each measures relation to the general factor within a specific intelligence battery.  Although in many situations a measure g loading will not change dramatically when computed in the context of a different collection of diverse cognitive tests (Jensen, 1998; Jensen & Weng, 1994), this will not always be the case.

Within (internal-validity) vs across (joint; external validity) estimation of test measures g-ness

When measures from different batteries are combined in the joint-battery approach, the battery-bound g  estimates for some measures may be altered significantly.   Flanagan et al. (2000) demonstrated these when they calculated within- and joint-battery g estimates for the WISC-III.  These estimates were derived from a sample of 150 subjects who were administered the WISC-III and WJ III cognitive measures as part of the Phelps validity study reported for the WJ III cognitive technical manual.  Within-battery g estimates were calculated with the WISC-III data based on the first unrotated principal component.  Next the joint-battery factor analysis allowed for an examination of the WISC-III g estimates when calculated together with another intelligencet test battery (WJ III), one that included a broader array of CHC abilitiy measures.

Flanagan et al. (2000) reported that the within- and joint-battery WISC-III g loadings were similar for many of the individual measures.  For example, the within- and joint-battery test g loadings are generally similar (i.e., do not differ by more than .05) for the Similarities (.76 vs .71), Vocabulary (.78 vs .74), Digit Span (.48 vs .49), Block Design (.60 vs .61), Object Assembly (.50 vs .45), and Symbol Search (.57 vs .54) measures.  These six WISC-III measures appear to have similar g characteristics when examined from the perspective of either the WISC-III or CHC (WJ III battery) frameworks.  However, the joint-battery g loadings were noticeably lower than the within-battery g loadings (i.e., lower by .06 or more) for Information  (.77 vs .68), Arithmetic (.70 vs .64), Comprehension (.59 vs .51), Picture Completion (.50 vs .40), Picture Arrangement (.37 vs .31), and Coding (.46 vs .37).   The results suggested that the latter WISC-III measures were relatively weaker g indicators than is suggested by within-battery WISC-III g analysis.

This example demonstrates the potential chameleon nature of test measures g estimates that are calculated within the confines of individual intelligence batteries when compared to those calculated within a comprehensive set of ability measures. 

And, yet to be mentioned is another, older, and for some reasons under-utilized statistical method for examing the g-ness (congitive complexity) of IQ test measures...multidmensional scaling (MDS).  We will save that for the next post in this seires.

To be continued........................

.