Showing posts with label AP101 Reports. Show all posts
Showing posts with label AP101 Reports. Show all posts

Thursday, March 1, 2012

IAP101 Brief #12: Use of IQ component part scores as indicators of general intelligence in SLD and MR/ID diagnosis

   
            Historically the concept of general intelligence (g), as operationalized by intelligence test battery global full scale IQ scores, has been central to the definition and classification of individuals with a specific learning disability (SLD) as well as individuals with an intellectual disability (ID).  More recently, contemporary definitions and operational criteria have elevated intelligence test battery composite or part scores to a more prominent role in diagnosis and classification of SLD and more recently in ID.
            In the case of SLD, third-method consistency definitions prominently feature component or part scores in (a) the identification of consistency between low achievement and relevant cognitive abilities or processing disorders and (b) the requirement that an individual demonstrate relative cognitive and achievement strengths (see Flanagan, Fiorello & Ortiz, 2010).  The global IQ score is de-emphasized in the third-method SLD methods.
            In contrast, the 11th edition of the AAIDD Intellectual Disability: Definition, Classification, and Systems of Supports manual (AAIDD, 2010) placed general intelligence, and thus global composite IQ scores, as central to the definition of intellectual functioning.  This has not been without challenge.  For example, the AAIDD ID definition has been criticized for an over-reliance on the construct of general intelligence and for ignoring contemporary psychometric theoretical and empirical research that has converged on a multidimensional hierarchical model of intelligence (viz., Cattell-Horn-Carroll or CHC theory).
The potential constraints of the “ID-as-a-general-intelligence-disability” definition was anticipated by the Committee on Disability Determination for Mental Retardation, in its National Research Council report “Mental Retardation:  Determining Eligibility for Social Security Benefits” (Reschly, Meyers & Hartel, 2001).  This national committee of experts concluded that “during the next decade, even greater alignment of intelligence tests and the IQ scores derived from them and the Horn-Cattell and Carroll models is likely.  As a result, the future will almost certainly see greater reliance on part scores, such as IQ scores for Gc and Gf, in addition to the traditional composite IQ.  That is, the traditional composite IQ may not be dropped, but greater emphasis will be placed on part scores than has been the case in the past” (Reschly et al., 2002, p. 94).  The committee stated that “whenever the validity of one or more part scores (subtests, scales) is questioned, examiners must also question whether the test’s total score is appropriate for guiding diagnostic decision making.  The total test score is usually considered the best estimate of a client’s overall intellectual functioning.  However, there are instances in which, and individuals for whom, the total test score may not be the best representation of overall cognitive functioning.” (p. 106-107).
            The increased emphasis on intelligence test battery composite part scores in SLD and ID diagnosis and classification raises a number of measurement and conceptual issues (Reschly et al., 2002).  For example, what are statistically significant differences?  What is a meaningful difference?  What appropriate cognitive abilities should serve as proxies of general intelligence when the global IQ is questioned?  What should be the magnitude of the total test score? 
Appropriate cognitive abilities will only be the only issue discussed here.  This issue addresses  which component or part scores are more correlated with general intelligence (g)—that is, what component part scores are high g-loaders?  The traditional consensus has been that measures of Gc (crystallized intelligence; comprehension-knowledge) and Gf (fluid intelligence or reasoning) are the highest g-loading measures and constructs and are the most likely candidates for elevated status when diagnosing ID (Reschly et al., 2002).  Although not always stated explicitly, the third method consistency SLD definitions specify that an individual must demonstrate “at least an average level of general cognitive ability or intelligence” (Flanagan et al., 2010, p.745), a statement that implicitly suggests cognitive abilities and component scores with high g-ness.
Table 1 is intended to provide guidance when using component part scores in the diagnosis and classification of SLD and ID (click on images to enlarge and use the browser zoom feature  to view; it is recommended you click here to access a PDF copy of the table..and also zoom in on it).  Table 1 presents a summary of the comprehensive, nationally normed, individually administered intelligence batteries that possess satisfactory psychometric characteristics (i.e., national norm samples, adequate reliability and validity for the composite g-score) for use in the diagnosis of ID and SLD.



The Composite g-score column lists the global general intelligence score provided by each intelligence battery.  This score is the best estimate of a persons general intellectual ability, which currently is most relevant to the diagnosis of ID as per AAIDD.  All composite g-scores listed in Table 1 meet Jensens (1998) psychometric sampling error criteria as valid estimates of general intelligence.  As per Jensens number of tests criterion, all intelligence batteries g-composites are based on a minimum of nine tests that sample at least three primary cognitive ability domains.  As per Jensens variety of tests criterion (i.e., information content, skills and demands for a variety of mental operations), the batteries, when viewed from the perspective of CHC theory, vary in ability domain coveragefour (CAS, SB5), five (KABC-II, WISC-IV, WAIS-IV), six (DAS-II) and seven (WJ III) (Flanagan, Ortiz & Alfonso, 2007; Keith & Reynolds, 2010).   As recommended by Jensen (1998), the particular collection of tests used to estimate g should come as close as possible, with some limited number of tests, to being a representative sample of all types of mental tests, and the various kinds of test should be represented as equally as possible (p. 85).  Users should consult sources such as Flanagan et al. (2007) and Keith and Reynolds, 2010) to determine how each intelligence battery approximates Jensens optimal design criterion, the specific CHC domains measured, and the proportional representation of the CHC domains in each batteries composite g-score.
Also included in Table 1 are the component part scales provided by each battery (e.g., WAIS-IV Verbal Comprehension Index, Perceptual Reasoning Index, Working Memory Index, and Processing Speed Index), followed by their respective within-battery g-loadings.[1]  Examination of the g-ness of composite scores from existing batteries (see last three columns in Table 1) suggests the traditional assumption that measures of Gf and Gc are the best proxies of general intelligence may not hold across all intelligence batteries.[2] 
In the case of the SB5, all five composite part scores are very similar in g-loadings (h2 = .72 to .79).  No single SB5 composite part score appears better than the other SB5 scores for suggesting average general intelligence (when the global IQ score is not used for this purpose).  At the other extreme is the WJ III where the Fluid Reasoning, Comprehension-Knowledge, Long-term Storage and Retrieval cluster scores are the best g-proxies for part-score based interpretation within the WJ III.  The WJ III Visual Processing and Processing Speed clusters are not composite part scores that should be emphasized as indicators of general intelligence.  Across all batteries that include a processing speed component part score (DAS-II, WAIS-IV, WISC-IV, WJ III) the respective processing speed scale is always the weakest proxy for general intelligence and thus, would not be viewed as a good estimate of general intelligence. 
            It is also clear that one cannot assume that composites with similar sounding names of measured abilities should have similar relative g-ness status within different batteries.  For example, the Gv (visual-spatial or visual processing) clusters in the DAS-II (Spatial Ability), SB5 (Visual-Spatial Processing) are relatively strong g-measures within their respective battery, but the same cannot be said for the WJ III Visual Processing cluster.  Even more interesting are the differences in the WAIS-IV and WISC-IV relative g-loadings for similarly sounding index scores. 
For example, the Working Memory Index is the highest g-loading component part score (tied with Perceptual Reasoning Index) in the WAIS-IV but is only third (out of four) in the WISC-IV.   The Working Memory Index is comprised of the Digit Span and Arithmetic subtests in the WAIS-IV and the Digit Span and the Letter-Number Sequencing subtests in the WISC-IV.  The Arithmetic subtest has been reported to be a factorially complex test which may tap fluid intelligence (Gf-RQ—quantitative reasoning), quantitative knowledge (Gq), working memory (Gsm), and possible processing speed (Gs; Keith & Reynolds, 2010; Phelps, McGrew, Knopik & Ford, 2005).   The factorially complex characteristics of the Arithmetic subtest (which, in essence, makes it function like a mini-g proxy) would explain why the WAIS-IV Working Memory Index is a good proxy for g in the WAIS-IV but not in the WISC-IV. The WAIS-IV and WISC-IV Working Memory Index scales, although named the same, are not measuring identical constructs.

A critical caveat is that the g-loadings cannot be compared across different batteries.  g-loadings may change when the mixture of measures included in the analyses change.  Different "flavors" of g can result (Carroll, 1993; Jensen, 1998). The only way to compare the g-ness across batteries is with appropriately designed cross- or joint-battery analysis (e.g., WAIS-IV, SB5 and WJ III analyzed in a common sample).
The above within and across intelligence battery examples illustrates that those who use component part scores as an estimate of a person’s general intelligence must be aware of the composition and psychometric g-ness of the component scores within each intelligence battery.  Not all component part scores in different intelligence batteries are created equal (with regard to g-ness).  Also, not all similarly named factor-based composite scores may measure the same identical construct and may vary in degree of within battery g-ness.  This is not a new problem in the context of naming factors in factor analysis, and by extension, factor-based intelligence test composite scores, Cliff (1983) described this nominalistic fallacy in simple language—“if we name something, this does not mean we understand it” (p. 120). 




[1] As noted in the footnotes in Table 1, all composite score g-loadings were computed by Kevin McGrew by entering the smallest number (and largest age ranges covered) of the published correlation matrices within each intelligence batteries technical manual (note the exception for the WJ III) in order to obtain an average g-loading estimate.  It would have been possible to calculate and report these values for each age-differentiated correlation matrix for each intelligence battery.  However, the purpose of this table is to provide the best possible average value across the entire age-range of each intelligence battery.  Floyd and colleagues have published age-differentiated g-loadings for the DAS-II and WJ III.  Those values were not used as they are based on the use of the principal common factor analysis method, a method that  analyzes the reliable shared variance among tests.  Although principal factor and principal component loadings typically will order measures in the same relative position, the principal factor loadings typically will be lower.  Given that the imperfect manifest composite scale scores are those that are utilized in practice, and to also allow uniformity in the calculation of the g-loadings reported in Table 1, principal component analysis was used in this work. The same rationale was used for not using the latent factor loadings on a higher-order g-factor in SEM/CFA analysis of each test battery.  Loadings from CFA analyses represent the relations between the underlying theoretical ability constructs and g purged of measurement error.  Also, frequently the final CFA solutions reported in a batteries technical manual (or independent journal articles) allow tests to be factorially complex (load on more than one latent factor), a measurement model that does not resemble the real world reality of the manifest/observed composite scores used in practice.  Latent factor loadings on a higher-order g-factor will often differ significantly from principal component loadings based on the manifest measures, both in absolute magnitude and relative size (e.g., see high Ga loading on g in WJ III technical manual which is at variance with the manifest variable based Ga loading reported in Table 1) 
[2] The h2 values are the values that should be used to compare the relative amount of g-variance present in the component part scores within each intelligence battery.

Thursday, September 15, 2011

Psychometric issues in Atkins MR/ID cases: The IAP AP101 report series




It is clear that many Atkins MR/ID death penalty decisions revolve around psychometric issues that many (but not all) attorneys, judges, and psychologists (who do assessments) are not well versed. Recurring issues in many cases are norm obsolescence (Flynn Effect), standard error of measurement (SEM), practice effects, full-scale vs component part scores, differences between IQs from different tests, to name but a few.

I always think that once I've posted an Applied Psychometrics 101 working paper, all who visit the ICDP blog will easily find the materials. But, I still receive regular phone calls and questions suggesting that my assumption is not correct.

Thus, the purpose of this post is to remind those looking for some information on some of the recurring psychometric issues that a series of reports are available for download. They are listed on the right side of the blogroll, as can bee seen in the image below.(double click on image to enlarge)


To make this info accessible again, I have created a link here that when clicked will provide the readers with all blog posts that reference the reports and provide report-specific links.

I have a couple of new reports "in limbo" and hope to find time in the near future to add more.




- iPost using BlogPress from Kevin McGrew's iPad

Tuesday, June 29, 2010

The Flynn Effect report series: What is the Flynn Effect: IAP AP101 Report #6

A new IAP Applied Psychometrics 101 report (#6) is now available.  The report is the first in the Flynn Effect series, a series of brief reports that will define, explain and discuss the validity of the Flynn Effect (click here to access all prior FE related posts at the ICDP blog) and the issues surrounding the application of a FE "adjustment" for scores based on tests with date norms (norm obsolescence), particularly in the context of Atkins MR/ID capital punishment cases.  The abstract for the brief report is presented below.  The report can be accessed by clicking here.
Norm obsolescence is recognized in the intelligence testing literature as a potential source of error in global IQ scores.  Psychological standards and assessment books recommend that assessment professionals use tests with the most current norms to minimize the possibility of norm obsolescence spuriously raising an individual’s measured IQ.  This phenomenon is typically referred to as the Flynn Effect.  This report is the first in a series of brief reports the will define, explain, and summarize the scholarly consensus regarding the validity of the Flynn Effect.  The series will conclude with an evaluation of the question whether a professional consensus has emerged regarding the practice of adjusting dated IQ test scores for the Flynn Effect, an issue of increasing debate in Atkins MR/ID capital punishment hearings.

Technorati Tags: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ,

Wednesday, April 7, 2010

Psychometric PS to Johnston v Florida (2010) denied appeal re: new WAIS-IV scores

This is a follow-up to my brief comments yesterday regarding the Johstone v Fl (2010) denied MR/ID appeal of two days ago.

As mentioned in the decision and my blog comment, the WAIS-III/WAIS-IV tests correlated .94 in a study reported in the WAIS-IV technical manual.  This is a very high correlation...but does NOT mean that the two tests should be expected to provide identical IQ scores.  I discuss these issues in a prior IAP AP101 report.

The tests have different norm dates and thus, the later version (WAIS-IV) would be expected to provide a lower score based on the Flynn effect.  More importantly, as reported in the IAP AP101 report, when one calculates the standard deviation of the difference score (see page 6 of that report) for a correlation of .94, the resulting value is 5.2 (round to 5 for ease of discussion).  This means that, on average, the WAIS-III/WAIS-IV (even if highly correlated at the .94 level) would in the general population be expected to display a range of difference scores from -5 to +5...or a range of 10 IQ points......in 68% of the population.  Please review that prior report for further explanation and discussion.

Technorati Tags: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ,

Friday, February 5, 2010

AP101 Brief #6: Understanding Wechsler IQ score differences--the CHC evolution of the Wechsler FS IQ score

[Note.  A typo in the original tables used to construct the WAIS figure below has been fixed.  Visual Puzzles on the WAIS-IV had been incorrectly designated as a measure of Gf----it should have been classified Gv.  This has now been changed and the corresponding text also modified.  Sorry for this error.  Changes in the text are so designated below via the strikeover]

Why do the IQ scores for the same individual often differ?

This question often perplexes both users and recipients of psychological reports. In a previous IAP Applied Psychometrics 101 report (AP101 #1:  Understanding IQ score differences) I discussed general statistical information related to the magnitude and frequency of expected IQ score differences for different tests (as a function of the correlation between tests).  In that report I mentioned the following general categories of possible reasons for IQ score differences/discrepancies.
Factors contributing to significant IQ differences are many, and include: (a) procedural or test administration issues (e.g., scoring errors; improper test administration; malingering; age vs grade norms), (b) test norm or standardization differences (e.g., possible errors in the norms; sampling plan for selecting subjects for developing the test norms; publication date of test), (c) content differences, and/or, (d) in the case of group research, research methodology issues (e.g., sample pre-selection effects on reported mean IQs) (McGrew, 1994).
At this time I  return to one of these factors--content differences. This brief report does not focus on content differences between different IQ tests but, instead, focuses on the changing content across the various editions of the two primary Wechsler intelligence batteries (WISC/WAIS). This information should be useful when individuals are comparing IQ scores (for the same person) based on different versions of the Wechsler's .

Of course, content differences will not be the only reason for possible IQ score differences across editions of the Wechsler's for an individual. Other possible reasons may include real changes in intelligence, serious scoring errors present in either one of the two test administration's, the Flynn effect, and other possible factors.   This post focuses only on the changing CHC content of the WISC and WAIS series of intelligence batteries.

As discussed previously in numerous posts, contemporary CHC theory is currently considered the consensus psychometric taxonomy of human cognitive abilities (click here for prior posts and information regarding the theory).  For this current brief report, I reviewed the extant CHC-organized factor analysis literature of the variousWechsler intelligence batteries. I then used this information as per the following steps:

1.  I identified the individual subtests in all editions of the WISC and WAIS batteries that contributed to the respective Full Scale (FS) IQ score for each battery.

2.  Using the accepted authoritative sources re: the CHC analysis of the Wechsler intelligence batteries (Flanagan, McGrew and Ortiz, 2000; Flanagan, Ortiz, and Alfonso, 2007; McGrew and Flanagan, 1998; Woodcock, 1990), I classified each of the above identified subtests as per the broad CHC ability (or abilities) measured by each subtest.  For readers who want a very brief CHC overview (and ability definition cheat-sheet), click here.

3.  I calculated the percentage of each broad CHC ability represented in each batteries respective FS IQ. For example, for the 1974 WISC-R, the FS IQ is calculated by summing the WISC-R scaled scores from 10 of the individual subtests. Four of these 10 subtests (Information, Comprehension, Similarities, and Vocabulary) have all been consistently classified as indicators of broad Gc. Since each of the individual subtests contribute equally to the FS IQ score, Gc represents at least 40%  (4 of 10) of the WISC-R FS IQ. 
  • However, the extant CHC Wechsler research has consistently identified a few tests with dual CHC factor loadings. In particular, both Picture Completion and Picture Arrangement have been consistently reported to load on both the Gv (performance scale) and Gc (verbal scale) on the WISC-R. For tests that demonstrated consistent dual CHC factor loadings, I assigned each broad CHC ability measured as representing 1/2 (0.5) of the test. More precise proportional calculation might have been possible (via the calculation of the average factor loadings across all studies), but for the current purpose I used this  simple and (IMHO) reasonably approximate method.
  • As a result, both the Picture Completion and Picture Arrangement subtests were each assigned a 1/2 (0.5) Gc and 1/2 (0.5) ability classifications. When added together these two 0.5 Gc test classifications sum to 1.0. When combined with the other four clear Gc tests mentioned above, the final Gc test indicator total is 5.  As a result, the total Gc proportional percentage of the WISC-R FS IQ was calculated as 50%.
4.  Although the Wechsler CHC classifications were based on the primary source sources noted above, I did revise some commonly accepted classifications based upon my professional opinion (when supported by empirical research). For example, the Arithmetic subtest has frequently been classified as a measure of Gf, Gsm, and sometimes Gs.   However, when valid factor indicators of Quantitative Knowledge (Gq) have been included in analyses, the Arithmetic subtest consistently displays a robust loading on the Gq factor and only minor loadings on other CHC abilities. I placed greater stock in these studies (e.g., Phelps at al, 2005: Woodcock, 1990) as I deem these to be better designed CHC studies (they included a broader array of CHC ability indicators).  My final determination for Arithmetic was that it is a test that measures both Gq and Gsm.
  • In addition, where appropriate and consistent with published research, I modified a few other commonly accepted CHC Wechsler test classifications to reflect recent research (e.g.., Kaufman et al., 2001; Keith et al., 2006; Keith & Reynolds (in press--CHC abilities and cognitive tests: What we've learned from 20 years of research;  Psychology in the Schools); Lichtenberger & Kaufman, 2001; McGrew, 2009; Tulsky & Price, 2003; plus the factor studies reported in the respective technical manuals of each battery). Referring to the mixed measures of Picture Completion and Picture Arrangement mentioned above, research with the WISC-IV  has suggested that Picture Completion is primarily a measure Gv (Gc factor loading minimal or nonexistent) while Picture Arrangement continues to show significant loadings on both Gv and Gc. Thus, Picture Arrangement was classified as a mixed measure of Gc and Gv for all editions of the WISC. In contrast, in the case of the WISC-IV  Picture Completion was classified as a measure Gv.  
  • It is not possible to describe in detail all of the minor "fine tunings" I did for select Wechsler CHC test classifications. The basis for all are included in the various reference sources cited above. In the final analysis the Wechsler CHC test classifications used in this brief report are those made by myself (Kevin McGrew) based on my integration and understanding of the extant empirical research regarding the CHC abilities measured by individual tests in both the WISC and WAIS series of intelligence batteries.
5.  Finally, I calculated the proportion of CHC abilities represented in the FS IQ scores for all editions of the WISC and WAIS.  These value were tabled and plotted on graphs.  The summary graphs are presented below. [Double click on images to enlarge]





Conclusions/observations:  A review of all information presented (in and across both graphs) produces a number of interesting conclusions and hypotheses. I only present a few at this time. I encourage others to review the documents and provide additional insights or commentary via the comment feature of the blog or on various listserv's where I have posted and FYI message regarding this set of analysis.

1.  Historically, the FS IQ score from the Wechsler batteries, which is typically interpreted as a measure of general intelligence (g), has been heavily weighted towards the measurement of Gc and Gv abilities. This should not be surprising given the original design blueprint specified by David Wechsler (the measurement of intelligence vis-a-vis two different modes of expression).

2.  The WISC series remained constant in the CHC FS IQ composition from 1949 to 1991. Although tests may have been revised or replaced, the differential CHC proportional contribution to the FS IQ was relatively equal across all three editions. Following the 80% combined contribution of Gc and Gv, much smaller contributions to the FS IQ came from measures of Gs (10%) and Gq and Gsm (5% respectively).

3.  The WISC-IV represents a significant change in the general intelligence FS IQ score provided. Gc representation has decreased approximately 20%, Gv representation was cut in half (30 % to 15 %) ,  Gs abilities increased slightly (5 %), and Gq was eliminated. More importantly, there was a fourfold increase in the contribution of the Gsm (from 5% to 20%) and a 20% increase in Gf representation (from 0 to 20%)! Clearly different FS IQ scores may be obtained by the same individual when comparing WISC-IV FS IQ to either WISC-R/WISC-III scores.  More importantly,the difference may be a function of the different mixture of CHC abilities represented in the different editions of the WISC series. 

4.  The first two editions of the WAIS (WAIS and WAIS-R) were identical in differential CHC ability contribution to the FS IQ score. However, starting with the WAIS-III significant changes in the adult Wechsler battery commenced and were later amplified in the WAIS-IV. Both the WAIS-III and WAIS-IV FS IQs reduced the amount of Gc representation by approximately 14% to 15%. The contribution of Gv decreased only slightly (27.3% to 22.7%) from the WAIS-R to WAIS-III, but there was a dramatic reduction (by one half) and then another 2% from the WAIS-III to the WAIS-IV (22.7% to 10% 20%). Offsetting reductions in Gc and Gv over these two editions was a trend towards greater measurement of Gs (has doubled from around 9% from the early two editions to approximately 18% to 20% in the last two editions). Gq FS IQ contribution has remained relatively similar throughout all editions. The most dramatic change, which is also consistent with the WISC series, is an approximate tenfold increase (0 % to 9.1%) in Gf from the WAIS-R to the WAIS-III, which was again doubled in magnitude with the publication of the and WAIS-IV (10% 20%). In general, similar to the WISC series, the adult WAIS series FS IQ has slowly evolved in the CHC abilities represented by the FS IQ. Both Gc and Gv abilities have been systematically reduced concurrently with a significant increases in the contribution of Gs and Gf.

Implications of the CHC evolution of the WISC and WAIS FS IQ scores are many if one attempts to compare a current IQ score from one battery to an older score from a earlier edition of the same battery (or compare an older score from the childrens version to the latest edition of the adult version). Before one can assume that significant changes from a childhood WISC-based IQ to a WAIS-III or WAIS-IV  are due to certain factors (neurological insult; malingering, the Flynn effect, etc.), one should review the above graphs and consider the possibility that the different FS IQ scores may both be valid indicators of functioning but may represent differ CHC mixes (flavors) of general intelligence.

The potential implications and  hypotheses that can be generated with the aid of the above graphs are numerous. For example, Flynn (2006) has suggested that there are problems with the WAIS-III standardization norms given that studies comparing the WAIS-R/WAIS-III scores are not consistent with Flynn effect expectations.  According to Weiss (2007), Flynn is ignoring data that does not fit his theory and instead is using theory to question data (and the integrity of a tests norms). According to Weiss (2007), "the only evidence Flynn provides for this statement is that WAIS-III scores do not fit expectations made based on the Flynn effect. However, the progress of science demands that theories be modified based on new data. Adjusting data to fit theory is an inappropriate scientific method, regardless of how well supported the theory may have been in previous studies." (p.1 from abstract).

I tend to concur with Weiss's arguments that the mere finding that the WAIS-III results were inconsistent with  Flynn effect expectations is insufficient evidence to claim that the a test norms are wrong. If the data don't fit--one may need to retrofit (your theory or hypothesis).  By inspecting the second graph above, one can see that a  viable explanation for the apparent lack of the WAIS-R-to-WAIS-III Flynn effect is that the WAIS-III FS IQ score represents a different proportional composite of CHC abilities. More specifically, the WAIS-III reduced the proportional representation of Gc from 45.5% to 31.8%, decreased the Gv representation by approximately 5%, doubled the impact of Gs, and for the first time ever introduced close to 10% Gf representation. CHC content changes of the FS IQ scores between batteries may be at play.   Can anyone say "comparing apples to apples+oranges?"

And so on.................more comments may be forthcoming.

PS - additional information not included in this original post has now been posted.  Click here.

Technorati Tags: , , , , , , , , , , , , , , , , , , , , ,

Monday, December 14, 2009

The standard error of measurement (SEM): Explanations and facts for Atkins MR/ID death penalty proceedings: IAP AP101 Report #5



A new IAP Applied Psychometrics 101 report (#5) is now available.  The title of the report and abstract is below.  The report can be downloaded by clicking here.

Applied Psychometrics 101 #4:  The Standard Error of Measurement (SEM):  An Explanation and Facts for "Fact Finders" in Atkins MR/ID death penalty proceedings.

Abstract

The standard error of measurement (SEM) is a professionally accepted and scientifically based measurement concept that allows users of psychological test scores to account for the known degree of imprecision in the scores.  Atkins MR/ID cases almost always involve standardized psychological testing in the domains of intelligence (IQ tests) and adaptive behavior (AB).  Scores from IQ and AB measures are fallible—not perfectly reliable.  This report provides an easy to understand explanation of the psychometric concept of SEM augmented by an example based on real-world data.  The report concludes with 8 SEM facts that “fact finders” should understand and internalize when evaluating psychological test data during legal proceedings--Atkins MR/ID death penalty proceedings in particular.
Here is a visual treat/tease from the report:


All prior IAP AP101 reports can be accessed via the Applied Psychometrics 101 (AP101) Reports section of the blog--on the blog sidebar.

Technorati Tags: , , , , , , , , , , , , , , , , , , , , , , , , , , , ,


Monday, November 9, 2009

What does the WAIS-IV measure? CHC analysis and beyond....IAP AP101 Report #2




I've posted a new IAP Applied Psychometrics 101 Report (#2:  What does the WAIS-IV measure?  CHC analysis and beyond) at ICDPs sister blog, IQs Corner.  The focus is on understanding what the structure of the WAIS-IV tests measure collectively.

I plan a follow-up report or brief for ICDP that will present additional analysis that will augment this first report.  This follow-up will be focused more on the interpretation of the WAIS-IV composite indexes, as I've learned that the reporting of IQ test data in Atkins cases only focuses on composite or full-scale scores...and not individual subtests.  This follow-up report or brief will be posted here at ICDP.

You can access the new report under the APPLIED PSYCHOMETRICS 101 (AP101) REPORTS section of this blog.

Stay tunned.

Technorati Tags: , , , , , , , , , , , , , , , , ,


Friday, September 11, 2009

Why IQ test scores can differ: Applied Psychometrics 101 Report #1--Understanding global IQ test correlations



Announcing Applied Psychometrics 101: IQ Test Score Difference Series--#1 Understanding global IQ test correlations. (click here to view and/or download)

Toady I'm announcing the first in what I hope is a series of applied psychometric brief reports. The goal of this project is to explain basic psychometric issues to help professionals and the public better understand psychological measurement, IQ testing, etc. Above is the title of the first report (and a link where it can be accessed). Below is the abstract, followed by some thoughts and questions the report might generate. This report (and future reports) are accessible via a section [(Applied Psychometric 101 (AP101) Reports] on the side bar of this blog.

Abstract
Despite reported evidence of strong concurrent correlations among IQ tests (concurrent validity), different IQ tests often produce different IQ scores for the same individual. This may be due to a number of factors. Prior to discussing the various factors, one must first understand the basic language of typical IQ-IQ comparison research. In the first of this series, IQ-IQ test correlations are explained. Statistically significant high correlations between different IQ tests, although providing strong concurrent validity evidence for tests, do not guarantee similar or identical IQ scores for all individuals tested.
Blogmaster comments

After reading the report, I would encurage readers to come back and reflect on the comments below. I would like to thank Dr. Dale Watson for comments on an earlier draft of the report. Most all of the ideas generated below are thoughts he shared (and that I had been contemplating) after reading the report. I will shortly be adding Dr. Watson to the "experts" blog roll on the blog sidebar.

Some Post "AP101: IQ Score Difference Series--# 1 Understanding global IQ test correlations" thoughts for consideration

I (the blogmaster) assume that most laypersons and, more importantly, agencies that have developed strict prescriptive guidelines for IQ cut scores for service eligibility and/or life or death decisions (e.g., U.S. Supreme Court Atkins ruling that no one with intellectual disabilities/mental retardation can be executed), are unaware of the variability in IQ scores that can arise simply by using different IQ tests (see report). Based on the IAP AP101 report, one should reach the conclusion that the selection of which IQ test to adminster (to determine if an individual is mentally retarded--esp. mild MR) can be a life-or-death decision (i.e., Atkins death penalty cases)! Furthermore, given the adversarial nature of a court hearings/trials and due process hearings, it is clear that a wide variety of questions could arise regarding how to determine which test battery is the "best" measure of intelligence (when different IQ tests used by different psychologists and experts produce significantly different scores). A few examples are listed below:

1. What does “best” mean? Is “best” relative to the purpose for the testing (e.g., best for service eligibility; best for developing instructional education programs; best for making formal legal diagnosis, etc.)? Might certain IQ tests be “best” for certain purposes and other IQ test “best” for other purposes? Is it possible for one test to be “best” for all purposes, all ages, all cultures, etc. ?

2. Could a scenario occur where the courts request that a standard be used to identify the potentially “best” test when IQ-IQ differences are reported by experts? This raises extremely complex questions. For example:
  • Does the popularity of an IQ measure determine which IQ test battery is “best”? Historically the Wechsler series of tests have been considered the “gold standard of IQ tests” largely because of their popularity. Does popularity + more sales = “best?”
  • Can (should) an empirical standard be developed? Is it even possible?
  • Some in the field of intelligence testing have suggested that an IQ tests g (general intelligence) saturation (“g-ness”; amount of variance attributed to the first principal component extracted in principal component analysis—PCA) is a good criterion.
  • Or, is the “best” IQ indicator a composite score that differentially weights the tests in the global IQ score as determined by PCA?
  • Or, is the “best” IQ battery one that only includes tests that have high g-ness?
  • Or, is the “best” IQ battery the one that provides the broadest coverage of the major cognitive abilities established by the most accepted psychometric model of intelligence?
3. Is the amount of g-ness measured in an individual central to the definition of intelligence and/or different diagnostic categories (MR, LD, gifted, etc.). Is g-ness more central to a diagnosis of mental retardation and less (or equally) relevant to a diagnosis of specific learning disability?

4. If it were even possible (which the current author doubts) to establish a consensus on a “best-ness” criterion, would assessment personal be required to administer the so designated test? Who would make the judgment regarding which test battery (or batteries) are the best—would it be in the hands of individual psychologists, professional association, a judge, or…..?

It is hoped the above cited IAP AP101 report has clarified the reality that different IQ tests will often provide different IQ scores for the same individual. IQ-IQ difference scores will occur, and if the IQ tests are properly administered to a cooperative individual, the resultant IQ-IQ score differences are reliable and valid. The “why” of psychometrically sound IQ-IQ score differences is due to a number of possible factors, factors that will be explored in future reports in the Applied Psychometrics 101 series. The potential policy implications, as briefly illustrated by the above set of hypothetical questions, are many, complex, and will not have an easy answer. There may not be a suitable answer and the use of IQ scores in legal and/or adversarial settings may need to change to become more nuanced (i.e., allow for more expert interpretation of the meaning of IQ test scores and IQ-IQ difference scores) and less rigid and prescriptive.

The issues raised in the report do not reflect problems in the state-of-the-art of psychometrically sound IQ tests, but in the use (and misuse) of IQ test scores to make important decisions about individuals and to create public policy and law.

Technorati Tags: , , , , , , , , , , , , , , , , , , ,