An attempt to provide understandable and up-to-date information regarding intelligence testing, intelligence theories, personal competence, adaptive behavior and intellectual disability (mental retardation) as they relate to death penalty (capital punishment) issues. A particular focus will be on psychological measurement, statistical and psychometric issues.
Showing posts with label norms. Show all posts
Showing posts with label norms. Show all posts
Monday, June 15, 2015
WAIS-IV US/Canadian norms controversy---articles for readers to review
I previously provided an FYI post on a hot topic in Canada...claims that the new WAIS-IV Canadian norms were flawed. There are now three articles outlining the different arguments. The three articles, published in JPA, can be found here, here, and here.
I continue to not comment on this controversy given my obvious conflict of interest as a coauthor of the competing WJ-IV.
Kevin McGrew
Labels:
Canada,
norms,
WAIS-IV,
Wechsler batteries
Wednesday, July 11, 2012
AP 101 Brief #14: Demographically adjusted neuropsych (Heaton) norm-based scores inappropriate for MR/ID Dx
Applied Psychometrics 101 Brief #
14: Demographically adjusted neuropsychological
(Heaton) norm-based scores are inappropriate for the diagnosis of MR/ID
Kevin S. McGrew, PhD.
Director
Dale G. Watson, PhD.
Berkeley, CA
Neuropsychological
assessments are sometimes part of psychological evaluations in Atkins MR/ID death penalty cases. These assessments include specialized tests,
often in addition to an age-appropriate IQ battery, that are
specifically designed to assess brain-behavior relations. The neuropsychological-specific
tests (NPST) are used to draw inferences about brain function/dysfunction
and to provide functional implications of neuropsychological test data for a
person’s real-world functioning. NPST
batteries, as well as all the individual tests included in NPST batteries, are not designed or validated to provide
a reliable and valid estimate of a person’s general intelligence (of
course, an exception is the portion of the battery that may include an
individualized measure of general intelligence; e.g., WAIS-IV; WJ III; SB5).
Demographically Adjusted Test Norm
Interpretation is Inappropriate in the Diagnosis of Atkins MR/ID in Capital
Cases
A
test interpretation feature used in some neuropsychological assessments is demographically adjusted norms. The specialized NPST of memory,
sensory-motor function, concept formation, etc. may be reported with these
special demographically adjusted norms.
Also, demographically adjusted norms are sometimes applied to the
individualized measure of general
intelligence included as part of the NPST (see Lange et al., 2006).
The
most well known demographically adjusted norms are the Heaton norms.
As described by a
neuropsychologist in a recent Atkins
cases, Heaton norms are “number crunching, age-corrected, you know,
socioeconomic variable-corrected data,” and as generating “a comprehensive
T-score age-, education, sex-corrected, actually race-corrected, also.” In simple terms, the demographically-adjusted
norms make equation-based statistical adjustments that allow certain NPST
scores for an individual to be compared
against other individuals of the same age and other demographic characteristics
(e.g., gender, race, socio-economic status and level of education. In the context of neuropsychological assessment
to determine whether an individual’s functioning has decreased, such as after a
brain injury or a stroke, demographically adjusted norms may help with the
diagnosis of brain dysfunction and the identification of relative strength- and
weakness-generated interventions.
Siverberg and Millis (2009)
have outlined the clear distinction between using neuropsychological measures
to identify acquired deficits as opposed
to developmental deficiencies. They note:
If the clinician is interested in whether a patient has declined
from their premorbid status, contrasting their obtained raw scores with their
expected premorbid scores (based on age, education, gender, ethnicity, and any
other variables that add to their prediction) is most appropriate. This type of
comparison quantifies impairment—how much examinees’ scores are lowered
relative to their (estimated) preinjury/disease onset baseline. The degree of
impairment is likely most predictive of the patient’s success in returning to
(or continuing) work or other premorbidly engaged-in functional activities with
extraordinary or idiosyncratic cognitive demands. If, in contrast, the
clinician is interested in determining whether the patient’s cognitive
abilities are sufficient for the demands of universal functional tasks (e.g.,
activities of daily living, driving a car, operating a cashier, etc.),
comparing their raw [non-demographically adjusted] scores with general healthy
adult population norms, generating “absolute” scores, is most appropriate (p.
98).
When
used in the context of neuropsychological assessment, certain NPST scores are adjusted
so an individual’s performance is compared
not to the general population but only to others of the same age,
gender, race and level of education.
Norm-referenced testing is at
the heart of psychological assessment for the diagnosis of MR/ID (AAIDD, 2010). The diagnosis of MR/ID requires comparison of
a person’s scores against nationally
representative norms, not a comparison to others of
the same age, gender, race and level of education. An analogous situation would be for a
professional psychologist or lawyer whose intellectual functioning is in the
top 2% of the population as a whole, and who therefore obtains an IQ of 130
when his/her score is compared to nationally representative norms. If his/her score is instead compared only to
those of a group of his/her peers with a similar level of education, he/she may
fall only in the top 16% of that group and so his/her score would be much lower, perhaps 115.
Demographically
adjusted norm scores result in a sliding
reference point that no longer
represents a comparison to the general population, which is the only proper
reference point in the diagnosis of MR/ID. The use of demographically adjusted
norms is inappropriate if such adjusted scores are used to formulate and
inform an opinion regarding level of general intellectual functioning for the
diagnosis of MR/ID.
The
technical adequacy and appropriate use of demographically adjusted NPST scores
is not a settled professional consensus in the field of
neuropsychological assessment. A lack of
consensus in the field is represented by the significantly differing opinions
as articulated by Heaton et al. (1996), Lange,
Chelune, Taylor, Woodward and Heaton (2006), Russell (2005), Sattler (2001),
Sherrill-Pattison, Donders, and Thompson (2000), Strauss,
Sherman and Spreen (2006), Romero et al.
(2009), Yantz,
Gavett, Lynch and McCaffrey (2006).
Particularly
important is the professional consensus-based report produced by the 2008 Multicultural Problem Solving Summit (Byrd et al.,
2010)
of neuropsychologists that produced the document “Challenges in the Neuropsychological Assessment of Ethnic
Minorities: Summit Proceedings” (Romero et al., 2009). This consensus report stated (emphasis added via underline):
Demographic
adjustments to normative data are not validated for the use of
predicting future academic or employment performance, and laws exist to prohibit
the use of race-based norms in employment decisions (p. 767).
There was
consensus among participants that the field would benefit from guidelines for
neuropsychological practice among ethnic and racial minorities…The
guidelines should include a specific focus on appropriate and inappropriate
uses of demographic adjustments, as well as a discussion of the risks of
overpathologizing groups or denying appropriate services, and details of
limitations to the application of various normative standards (p. 767).
Additionally,
a number of authoritative assessment texts used frequently in the graduate
training of psychologists learning to conduct intellectual or
neuropsychological assessments have highlighted the potential problems with
demographically adjusted norm scores (emphasis added via underlining or bold
font) .
Pluralistic
norms[1]
are norms derived for individual groups, such as Euro Americans, African
Americans, Hispanic Americans, Asian Americans, and Native
Americans…Pluralistic norms are potentially dangerous, however, because
they (a) provide a basis for invidious comparisons among different ethnic
groups, (b) may lower expectations of culturally and linguistically
diverse children and reduce their level of aspiration to succeed, (c) may have
little relevance outside of the child’s specific geographic area, and (d)
furnish no information about the complex reasons why some ethnic groups tend to
score lower than others on intelligence tests (p. 661).
There
are two schools of thought regarding how closely matched the norms must be to
the demographic characteristics of the individual being assessed, and these views are diametrically opposed.
These are: (1) that norms should be
as
representative of the general population as possible, and (2) that norms should
approximate, as closely as possible, the unique subgroup to which the
individual belongs. (p. 47).
At
times, it will be paramount to compare the individual to all other persons of
the same age in the general population. Determining a diagnosis of mental retardation or learning
disability would be one example (p. 47).
Of
historical interest is the fact that neuropsychology’s recent shift toward
demographically corrected scores based on race/ethnicity and other variables
has occurred with surprisingly little fanfare or controversy despite
ongoing debate in other domains of psychology. For example, when race-norming was applied to
pre-employment screening in the United States to increase the number of
minorities being chosen as job applicants, the result was the Civil Rights Act
of 1991, which outlawed race-norming
for applicant selection of referral (see Sackett & Wilk, 1994; Gottfredson,
1994; and Greenlaw & Jensen, 1996, for an interesting historical review of
the ill-fated attempt at race-norming the GATB) (p. 50).
There
is also evidence that when “corrective norms” are applied, some demographic
influences remain, and overcorrection may occur, resulting in score
distortion for some subgroups and a risk of increased false negatives
(e.g., Fastenau, 1998) (p. 50).
Importantly,
with regard to the WAIS-III/WMS-III, the Psychological Corporation explicitly
states that demographically adjusted scores are not intended for use in psychoeducational assessment, determination of intellectual deficiency,
vocational assessment, or any other context where the goal is to determine
absolute functional level (IQ or memory) in comparison to the general
population. Rather, demographically adjusted scores are best used for
neurodiagnostic assessment in order to minimize the impact of confounding
variables on the diagnosis of cognitive impairment. That is, they should be
used to infer strengths and weakness relative to a presumed pre-morbid standard
(The Psychological Corporation, 2002). Therefore, neuropsychologists need to
balance the risks and benefits of using within-group norms, and use them
with a full understanding of their implications and the situations in which
they are most appropriate (p. 51).
The
following select quotes from professional neuropsychological publications also
make it clear that the neuropsychological professional jury is still “out”
regarding the methodology and appropriate application of demographically
adjusted NPTS scores (emphasis added
via underline or bold font).
Clinical
neuropsychologists are always starving for good normative data for established
neuropsychological measures. Unfortunately, too many studies (and even
manuals) contain too few subjects and/or their samples are not representative
of the target population on important demographic variables, especially age
and education. This was the problem that Heaton, Grant, and Matthews
attempted to address with Comprehensive Norms for an Expanded Halstead Reitan Battery
(1991). This practical product arose as a direct result of the
authors’ 1986 chapter in an edited book (Heaton, Grant, & Matthews, 1986).
The project represents a substantial effort on the part of the authors, and it
has many commendable qualities. However, the merits are accompanied by
significant shortcomings (p. 444).
Our own informal
inquiries have indicated that many of the scientist-practitioners in clinical
neuropsychology have embraced this book in an uncritical manner. The strong and
unreflective nature of such acceptance of these norms tells us how good an idea
this kind of normative project is, in
the abstract. Unfortunately, this particular product does not contribute as
much as our expectations might lead us to anticipate. The format and marketing
are so convincing that few would comb the introductory pages to analyze the
test selection quirks and statistical/ design problems that abound (p.447).
While we
agree that this initial attempt at providing demographic corrections for
several commonly used tests could have been more statistically sophisticated,
and possibly could have been more user friendly, the evidence seems to indicate
that the norms do have significant advantages for neuropsychological clinical
work and research (p. 457).
It
is generally understood that demographically corrected normative standards are
based upon performances of adults who have developed normally, have typical,
mainstream educational backgrounds, and have no known history of brain injury
or disease. It follows logically from this that such norms should be used with
great caution, if at all, to identify acquired brain dysfunction in patients
who have developmental disorders or other-than-mainstream educational
backgrounds (e.g., special education). For example, it would be inappropriate to “adjust” a mentally
retarded person’s IQ upward because of a low education level, thereby
potentially depriving him/her of social services or mitigating considerations
in criminal prosecution.
As
we have noted, demographically corrected norms are used primarily to identify
the presence and nature of neurobehavioral changes due to known or possible
brain insult (injury or disease). Such norms are generally not the best
choice for characterizing the individual’s absolute level of functioning, or
functioning in relation to the general population of normal adults (Heaton
& Pendleton, 1981) (p. 147).
Heaton,
Grant, and Matthews (1991) published procedures for adjusting raw scores on
various neuropsychological tests according to the individual's age and
education. Despite rather widespread use of these score conversions in both
clinical work and research publications, there have been very few
investigations to evaluate the accuracy or limitations of these score
transformations (p. 181).
Inasmuch
as the HGM method brings about its greatest corrections among persons whose
scores are more likely to be affected by brain impairment, a question must
be raised concerning exactly what the HGM transformations are accomplishing.
It might seem, at least in part, that the corrections are, in fact, correcting
for subtle impairment of brain functions in less-educated and older persons—the
very condition that neuropsychological tests were developed to detect (p.
188).
Concluding Comments
Demographically adjusted (e.g., Heaton norms) scores
are 100% inconsistent with
determination of a person’s general level of intellectual functioning as per
the first prong of the diagnosis of MR/ID.
Demographically adjusted IQ scores, in particular, should not play a
critical role when determining if a person has a deficit in general
intellectual functioning. Furthermore,
interpretation of demographically adjusted NPST test scores (e.g., Halstead
Category Test) to provide convergent validity evidence regarding a person’s
level of general intelligence is not appropriate in the context of MR/ID
diagnosis.
American
Association on Intellectual and Developmental Disabilities. (2010). Intellectual
disability: Definition, classification,
and systems of supports—11th Edition. Washington, DC: Author.
Fastenau, P. S. & Adams, K.
M. (1996). Book review: Heaton, Grant,
and Matthews' Comprehensive Norms: An Overzealous Attempt. Journal of Clinical and Experimental Neuropsychology, 18 (3),
444-448.
Heaton, R. K., Ryan, L., & Grant,
I. (2009). Demographic influences and
use of demographically corrected norms in neuropsychological assessment. In
Igor Grant and Kenneth M. Adams (Eds), Neuropsychological
Assessment of Neuropsychiatric and Neuromedical Disorders, Oxford
University Press US.
Lange, R. T., Chelune, G. J.,
Taylor, M. J., Woodward, T. S., & Heaton, R. K. (2006). Development
of demographic norms for four new WAIS-III/WMS-III indexes. Psychological Assessment, 18 (2), 174 181.
Romero, H. R., Lageman, S. K., Kamath, V., Irani, F., Sim, A., Suarez, P., Manly, J. J., Attix, D. K., & the Summit participants (2009). Challenges in the neuropsychological assessment of ethnic minorities: Summit Proceedings. The Clinical Neuropsychologist, 23, 761-779.
Russell, E. W. (2005). Norming subjects for the Halstead Reitan
battery. Archives of Clinical
Neuropsychology, 20, 479-484.
Sattler, J. (2001). Assessment of
Children: Cognitive Applications—4th
Edition. San Diego, CA: Jerome M. Sattler, Publisher, Inc.
Sherrill-Pattison, S., Donders, J., & Thompson, E.
(2000). Influence of demographic variables on neuropsychological test
performance after traumatic brain injury. The
Clinical Neuropsychologist, 14 (4), 496-503.
Silverberg, N., & Millis, S. (2009). Impairment versus deficiency in
neuropsychological assessment: Implications for ecological validity. Journal
of the International Neuropsychological Society (2009), 15, 94–102.
Strauss, E., Sherman, E. M. S., & Spreen, O. (2006). A Compendium of Neuropsychological Tests:
Administration, Norms, and Competency – 3rd Edition. New York,
NY: Oxford University Press.
Titus, J. B., Retzlaff, P. D., & Dean, R. S. (2002). Predicting
scores of the Halstead Category Test with the WAIS-III. International Journal of Neuroscience, 112, 1099-1114.
Yantz, C. L., Gavett, B. E., Lynch, J. K., & McCaffrey, R. J.
(2006). Potential for the interpretation disparities of Halstead-Reitan
neuropsychological battery performances in a litigating sample. Archives of Clinical Neuropsychology, 21, 809-817.
[1]
The term “pluralistic norms”
refers to the same concept as demographically adjusted norms and was the terminology
used in the 1970’s when procedures for adjusting IQ scores for minority
children based on race and socio-economic status were attempted. Although using a different approach, the theory
behind adjusting IQ scores is conceptually similar to that inherent in
demographically adjusted norms.
[2] It is important to note that
this quote is from the primary author of the Heaton norms. This statement
indicates that the author of the Heaton norms views them as useful in
“clinical” and “research” settings, which I interpret to not cover high stakes
forensic diagnostic purposes such as Atkins
cases.
Sunday, March 27, 2011
IAP Applied Psychometrics 101 Report #10: "Just say no" to averaging IQ subtest scores
Should psychologists engage in the practice of calculating simple arithmetic averages of two or more scaled or standard scores from different subtests (pseudo-composites) within or across different IQ batteries? Dr. Joel Schneider and I, Dr. Kevin McGrew say "no."
Do psychologists who include simple pseudo-composite scores in their reports, or make interpretations and recommendations based on such scores, have a professional responsibility to alert recipients of psychological reports (e.g., lawyers, the courts, parents, special education staff, other mental health practitioners, etc.) of the potential amount of error in their statements when simple pseudo-composite scores are the foundation of some of their statements? We believe "yes."
Simple pseudo-composite scores, in contrast to norm-based scores (i.e., composite scores with norms provided by test publishers/authors--e.g., Wechsler Verbal Comprehension Index), contain significant sources of error. Although they have intuitive appeal, this appeal cloaks hidden sources of error in the scores---with the amount of error being a function of a combination of psychometric variables.
IAP Applied Psychometrics 101 Report #10 addresses the psychometric issues involved in pseudo-composite scores.
In the report we offer recommendations and resources that allow users to calculate psychometrically sound pseudo-composites when they are deemed important and relevant to the interpretation of a person's assessment results.
Finally, understanding the sources of error in simple pseudo-composite scores provides an opportunity for practitioners to understand the paradoxical phenomenon frequently observed in practice where norm-based or psychometrically sound pseudo-composite scores are often higher (or lower) than the subtest scores that comprise the composite. The "total does not equal the average of the parts" phenomenon is explained conceptually, statistically, and via an interesting visual explanation based on trigonometry.

Abstract
The publishers and authors of intelligence test batteries provide norm-based composite scores based on two or more individual subtests. In practice, clinicians frequently form hypotheses based on combinations of tests for which norm-based composite scores are not available. In addition, with the emergence of Cattell-Horn-Carroll (CHC) theory as the consensus psychometric theory of intelligence, clinicians are now more frequently “crossing batteries” to form composites intended to represent broad or narrow CHC abilities. Beyond simple “eye-balling” of groups of subtests, clinicians at times compute the arithmetic average of subtest scaled or standard scores (pseudo-composites). This practice suffers from serious psychometric flaws and can lead to incorrect diagnoses and decisions. The problems with pseudo-composite scores are explained and recommendations made for the proper calculation of special composite scores.
- iPost using BlogPress from my Kevin McGrew's iPad
intelligence IQ tests IQ testing IQ scores CHC intelligence theory CHC theory Cattell-Horn-Carroll human cognitive abilities psychology school psychology individual differences cognitive psychology neuropsychology psychology special education educational psychology psychometrics psychological assessment psychological measurement IQs Corner general intelligence standard scores IQ subtests Wechsler IQ subtests IQ part scores IQ composite scores cross-battery assessment applied Psychometrics psychological measurement
Do psychologists who include simple pseudo-composite scores in their reports, or make interpretations and recommendations based on such scores, have a professional responsibility to alert recipients of psychological reports (e.g., lawyers, the courts, parents, special education staff, other mental health practitioners, etc.) of the potential amount of error in their statements when simple pseudo-composite scores are the foundation of some of their statements? We believe "yes."
Simple pseudo-composite scores, in contrast to norm-based scores (i.e., composite scores with norms provided by test publishers/authors--e.g., Wechsler Verbal Comprehension Index), contain significant sources of error. Although they have intuitive appeal, this appeal cloaks hidden sources of error in the scores---with the amount of error being a function of a combination of psychometric variables.
IAP Applied Psychometrics 101 Report #10 addresses the psychometric issues involved in pseudo-composite scores.
In the report we offer recommendations and resources that allow users to calculate psychometrically sound pseudo-composites when they are deemed important and relevant to the interpretation of a person's assessment results.
Finally, understanding the sources of error in simple pseudo-composite scores provides an opportunity for practitioners to understand the paradoxical phenomenon frequently observed in practice where norm-based or psychometrically sound pseudo-composite scores are often higher (or lower) than the subtest scores that comprise the composite. The "total does not equal the average of the parts" phenomenon is explained conceptually, statistically, and via an interesting visual explanation based on trigonometry.
Abstract
The publishers and authors of intelligence test batteries provide norm-based composite scores based on two or more individual subtests. In practice, clinicians frequently form hypotheses based on combinations of tests for which norm-based composite scores are not available. In addition, with the emergence of Cattell-Horn-Carroll (CHC) theory as the consensus psychometric theory of intelligence, clinicians are now more frequently “crossing batteries” to form composites intended to represent broad or narrow CHC abilities. Beyond simple “eye-balling” of groups of subtests, clinicians at times compute the arithmetic average of subtest scaled or standard scores (pseudo-composites). This practice suffers from serious psychometric flaws and can lead to incorrect diagnoses and decisions. The problems with pseudo-composite scores are explained and recommendations made for the proper calculation of special composite scores.
- iPost using BlogPress from my Kevin McGrew's iPad
intelligence IQ tests IQ testing IQ scores CHC intelligence theory CHC theory Cattell-Horn-Carroll human cognitive abilities psychology school psychology individual differences cognitive psychology neuropsychology psychology special education educational psychology psychometrics psychological assessment psychological measurement IQs Corner general intelligence standard scores IQ subtests Wechsler IQ subtests IQ part scores IQ composite scores cross-battery assessment applied Psychometrics psychological measurement
Generated by: Tag Generator
Thursday, December 2, 2010
IQ test battery publication timeline: Atkins MR/ID Flynn Effect cheat sheet
As I've become involved in consulting on Atkins MR/ID death penalty cases, a frequent topic raised is that of norm obsolescence (aka, the Flynn Effect). When talking with others I often have trouble spitting out the exact date of publication of the various revisions of tests, as I keep track of more than just the Wechsler batteries (which are the primary IQ tests in Atkins reports). I often wonder if others question my expertise...but most don't realize that there are more IQ batteries out there than just the Wechsler adult battery....and, in particular, a large number of child normed batteries and other batteries spanning childhood and adulthood. Thus, I decided to put together a cheat sheet for myself..one that I could print and have in my files. I put it together in the form of a simple IQ battery publication timeline. Below is an image of the figure. Double click on it to enlarge.
An important point to understand is that when serious discussions start focusing on the Flynn effect in trial's, most often the test publication date is NOT used in the calculation of how obsolete a set of test norms are. Instead, the best estimate of the year the test was normed/standardized is used, which is not included in this figure (you will need to locate this information). For example, the WAIS-R was published in 1981...but the manual states that the norming occurred from May 1976 to May 1980. Thus, in most Flynn effect discussions in court cases, the date of 1978 (middle of the norming period) is typically used. This makes recall of this information difficult for experts who track all the major individually administered IQ batteries.
Hope this helpful...if nothing else...you must admit that it is pretty :) Click on image to view.

- iPost using BlogPress from my Kevin McGrew's iPad
intelligence intelligence testing Atkins cases ICDP blog psychology school psychology neuropsychology Forensic psychology criminal psychology criminal justice death penalty capital punishment ABA IQ tests IQ scores adaptive behavior AAIDD mental retardation intellectual disability Flynn effect
An important point to understand is that when serious discussions start focusing on the Flynn effect in trial's, most often the test publication date is NOT used in the calculation of how obsolete a set of test norms are. Instead, the best estimate of the year the test was normed/standardized is used, which is not included in this figure (you will need to locate this information). For example, the WAIS-R was published in 1981...but the manual states that the norming occurred from May 1976 to May 1980. Thus, in most Flynn effect discussions in court cases, the date of 1978 (middle of the norming period) is typically used. This makes recall of this information difficult for experts who track all the major individually administered IQ batteries.
Hope this helpful...if nothing else...you must admit that it is pretty :) Click on image to view.
- iPost using BlogPress from my Kevin McGrew's iPad
intelligence intelligence testing Atkins cases ICDP blog psychology school psychology neuropsychology Forensic psychology criminal psychology criminal justice death penalty capital punishment ABA IQ tests IQ scores adaptive behavior AAIDD mental retardation intellectual disability Flynn effect
Wednesday, June 30, 2010
The Flynn Effect report series: Is the Flynn Effect a Scientifically Accepted Fact? IAP AP101 Report #7
Another new IAP Applied Psychometrics 101 report (#7) is now available. The report is the second in the Flynn Effect series, a series of brief reports that will define, explain and discuss the validity of the Flynn Effect (click here to access all prior FE related posts at the ICDP blog) and the issues surrounding the application of a FE "adjustment" for scores based on tests with date norms (norm obsolescence), particularly in the context of Atkins MR/ID capital punishment cases. The abstract for the brief report is presented below. The report can be accessed by clicking here.
Report # 1 (What is the Flynn Effect) can be found by clicking here.
This report is the second in a series of brief reports the will define, explain, and summarize the scholarly consensus regarding the validity of the Flynn Effect (FE). This brief report presents a summary of the majority of FE research (in tabular form of n=113 publications) which indicates (via a simple “vote tally” method) that despite no consensus regarding the possible causes of the FE, it is overwhelming recognized as a fact by the scientific community. The series will conclude with an evaluation of the question whether a professional consensus has emerged regarding the practice of adjusting dated IQ test scores for the Flynn Effect, an issue of increasing debate in Atkins MR/ID capital punishment hearings.
Technorati Tags: psychology, forensic psychology, forensic psychiatry, neuropsychology, intelligence, school psychology, psychometrics, educational psychology, IQ, IQ tests, IQ scores, adaptive behavior, adaptive functioning, intellectual disability, mental retardation, MR, ID, criminal psychology, criminal defense, criminal justice, ABA, American Bar Association, Atkins cases, death penalty, capital punishment, AAIDD, Atkins MR/ID listserv, ICDP blog, Flynn Effect, norm obselescence, Flynn Effect Series, IAP Applied Psychometric reports
Tuesday, June 29, 2010
The Flynn Effect report series: What is the Flynn Effect: IAP AP101 Report #6
A new IAP Applied Psychometrics 101 report (#6) is now available. The report is the first in the Flynn Effect series, a series of brief reports that will define, explain and discuss the validity of the Flynn Effect (click here to access all prior FE related posts at the ICDP blog) and the issues surrounding the application of a FE "adjustment" for scores based on tests with date norms (norm obsolescence), particularly in the context of Atkins MR/ID capital punishment cases. The abstract for the brief report is presented below. The report can be accessed by clicking here.
Technorati Tags: psychology, forensic psychology, forensic psychiatry, neuropsychology, intelligence, school psychology, psychometrics, educational psychology, IQ, IQ tests, IQ scores, adaptive behavior, adaptive functioning, intellectual disability, mental retardation, MR, ID, criminal psychology, criminal defense, criminal justice, ABA, American Bar Association, Atkins cases, death penalty, capital punishment, AAIDD, Atkins MR/ID listserv, ICDP blog, Flynn Effect, norm obselescence, Flynn Effect Series, IAP Applied Psychometric reports
Norm obsolescence is recognized in the intelligence testing literature as a potential source of error in global IQ scores. Psychological standards and assessment books recommend that assessment professionals use tests with the most current norms to minimize the possibility of norm obsolescence spuriously raising an individual’s measured IQ. This phenomenon is typically referred to as the Flynn Effect. This report is the first in a series of brief reports the will define, explain, and summarize the scholarly consensus regarding the validity of the Flynn Effect. The series will conclude with an evaluation of the question whether a professional consensus has emerged regarding the practice of adjusting dated IQ test scores for the Flynn Effect, an issue of increasing debate in Atkins MR/ID capital punishment hearings.
Technorati Tags: psychology, forensic psychology, forensic psychiatry, neuropsychology, intelligence, school psychology, psychometrics, educational psychology, IQ, IQ tests, IQ scores, adaptive behavior, adaptive functioning, intellectual disability, mental retardation, MR, ID, criminal psychology, criminal defense, criminal justice, ABA, American Bar Association, Atkins cases, death penalty, capital punishment, AAIDD, Atkins MR/ID listserv, ICDP blog, Flynn Effect, norm obselescence, Flynn Effect Series, IAP Applied Psychometric reports
Tuesday, May 4, 2010
MUST READ: Atkins best practice and standard recommendations (McVaugh & Cunningham, 2009)
I've been toying with the idea of jotting down a list of suggested "best practice" recommendations and suggested professional standards based on the mass of literature that I've been reading since starting the ICDP blog. Every time I think I should start, I have been paralyzed by the sheer scope of the task....reading and taking notes from all relevant literature sources would take massive time...and I have no grad. assistants or employees. I was thus thrilled when the following article arrived in my email inbox today.
Although I may not agree 100% with everything these authors state, I must say that with regard to what I have read to date, this article is probably the best single and solid source on suggested best practice and professional standard recommendations for the assessment and Dx of MR/ID in Atkins cases. The authors present 20 different recommended guidelines covering a large number of the critical issues in assessment and DX of MR/ID in a legal context (e.g., practice effects, SEM, Flynn Effect, retrospective assessment of AB, adaptive behavior domains, different state statutes, etc.).
This is a MUST read for all mental health professionals and folks in the legal profession who are involved in Atkins cases. I think this document could serve as a foundational starting point for any group working on the development of standards and practice recommendations in Atkins MR/ID cases.
Kudos to the authors for the excellent work. I plan to reread numerous times, and may add "my variations on a theme" to some of the specific guidelines (when time permits).
MacVaugh, G. & Cunningham, M. (2009). Atkins v. Virginia: Implications and recommendations for forensic practice. The Journal of Psychiatry the Law, 37, 131-187 (click here to view)
Abstract
Technorati Tags: psychology, forensic psychology, forensic psychiatry, neuropsychology, intelligence, school psychology, psychometrics, educational psychology, IQ, IQ tests, IQ scores, adaptive behavior, adaptive functioning, intellectual disability, mental retardation, MR, ID, criminal psychology, criminal defense, criminal justice, ABA, American Bar Association, Atkins cases, death penalty, capital punishment, AAIDD, Atkins MR/ID listserv, ICDP blog, standards, cultural issues, professional standards, best practice, Daubert standard, clinical judgment
Although I may not agree 100% with everything these authors state, I must say that with regard to what I have read to date, this article is probably the best single and solid source on suggested best practice and professional standard recommendations for the assessment and Dx of MR/ID in Atkins cases. The authors present 20 different recommended guidelines covering a large number of the critical issues in assessment and DX of MR/ID in a legal context (e.g., practice effects, SEM, Flynn Effect, retrospective assessment of AB, adaptive behavior domains, different state statutes, etc.).
This is a MUST read for all mental health professionals and folks in the legal profession who are involved in Atkins cases. I think this document could serve as a foundational starting point for any group working on the development of standards and practice recommendations in Atkins MR/ID cases.
Kudos to the authors for the excellent work. I plan to reread numerous times, and may add "my variations on a theme" to some of the specific guidelines (when time permits).
MacVaugh, G. & Cunningham, M. (2009). Atkins v. Virginia: Implications and recommendations for forensic practice. The Journal of Psychiatry the Law, 37, 131-187 (click here to view)
Abstract
In 2002, the United States Supreme Court held in the landmark case of Atkins v. Virginia that the execution of individuals who have mental retardation is unconstitutional. Following the Atkins holding, courts in death penalty jurisdictions have relied heavily upon mental health professionals in making a determination of whether or not capital offenders have mental retardation. The determination of mental retardation in death penalty cases, however, presents complex challenges for both courts and mental health professionals. In addition, there is variability in how death penalty states define mental retardation and in the assessment methods used by mental health professionals to diagnose mental retardation in such cases. The purpose of this article is to (a) describe how statutes in death penalty jurisdictions have operationalized the various clinical definitions of mental retardation, (b) discuss issues confronting examiners in assessing and diagnosing mental retardation in Atkins cases, and (c) provide recommendations for forensic practice.
Technorati Tags: psychology, forensic psychology, forensic psychiatry, neuropsychology, intelligence, school psychology, psychometrics, educational psychology, IQ, IQ tests, IQ scores, adaptive behavior, adaptive functioning, intellectual disability, mental retardation, MR, ID, criminal psychology, criminal defense, criminal justice, ABA, American Bar Association, Atkins cases, death penalty, capital punishment, AAIDD, Atkins MR/ID listserv, ICDP blog, standards, cultural issues, professional standards, best practice, Daubert standard, clinical judgment
Friday, November 20, 2009
Malingering in Atkins MR/ID DP cases: State-of-the art, malinger by proxy, and voodoo psychometrics
A quick reading of a small sample of Atkins MR/ID death penalty court decisions makes it clear that the issue of malingering is often a critical component of expert testimony.
The APA Dictionary of Psychology defines malingering as:
I am not an expert on the state-of-the-art of the psychometric integrity of various malingering measures used to purportedly detect defendant malingering. Clearly in capital punishment cases there is the possibility of a strong motivation to score low on IQ tests or standardized measures of adaptive behavior -- lower scores may make the difference between execution or life in prison without parole. Not being an expert in this area of forensic assessment, I'm going to try provide information from high quality sources re: the state-of-the art of malingering assessment. Also, when appropriate, I will point out situations of inappropriate (unethical?) malingering assessment methods when they are obvious. This current post contains a sampling of interesting malingering issues, research, and an example of inappropriate malingering assessment. Click here for prior posts re: malingering issues, research and references.
What does the research say about malingering assessment in the context of intellectual disability determination?
I've found the scholarly work of Salekin and Doane to be of particular value in providing an evaluation of the research in this area. Below are two recent journal articles by Salekin and Doane. I believe the abstracts/summaries speak for themselves.
Doane, B., & Salekin, K. L. (2009). Susceptibility of current adaptive behavior measures to feigned deficits. Law and Human Behavior, 33, 329-343.
Abstract
Salekin, K. L., & Doane, B. (2009). Malingering intellectual disability: The value of available measures and methods. Applied Neuropsychology, 16, 105-113.
Abstract
Another interesting topic is malingering resulting from examiner bias. I find the concept of "malingering by proxy" very interesting. Below is a discussion of this phenomenon as described by Schlesinger:
Schlesinger, L. B. (2003). A case study involving competency to stand trial: Incompetent defendant, incompetent examiner, or "malingering by proxy" ? Psychology, Public Policy, and Law, 9 (3/4), 381-399.
In two recent Atkins court decisions in the state of Oklahoma (see Salazar, 2005 and Lambert, 2005), the states prosecution psychologist (Dr. Prosecution Psychologist - Dr. PP) testified re: malingering based, in part, on non-normed, non-standardized malingering measures that Dr. PP had developed himself, and one which he named after his secretary (in an attempt to mask the purpose of the test to the defendant). The two "instruments" in question were the non-standardized Blackwell Memory Test and the Oklahoma Spelling Test. Apparently the Blackwell Memory test was modelsx after other formal instruments that use a "forced choice symptom validity" test format. Similar to prior voodoo psychometric activities commented on at this blog, I'm dumb founded that a professional psychologist testifying in an Atkins hearing, or any other clinical or forensic setting, would attempt to assess a psychological construct (viz., malingering) via the development of their own special instrument that did not undergo the professional accepted and required test development procedures (as clearly spelled out the the Joint Test Standards). This activity clearly violates a number of professional standards. Below are at least two (and I'm sure there are more when one examines all relevant professional codes of ethics/standards) from the Joint Test Standards:
Technorati Tags: psychology, forensic psychology, criminal psychology, criminal justice, neuropsychology, school psychology, educational psychology, ABA, American Bar Association, criminal defense lawyers, MR, ID, mental retardation, intellectual disability, IQ, IQ tests, IQ scores, intelligence, adaptive behavior, malingering, Joint Test Standards

The APA Dictionary of Psychology defines malingering as:
the deliberate feigning of an illness or disability to achieve a particular desired outcome (e.g., financial gain or escaping responsibility, punishment, impresonment, or military duty) (p.551)
I am not an expert on the state-of-the-art of the psychometric integrity of various malingering measures used to purportedly detect defendant malingering. Clearly in capital punishment cases there is the possibility of a strong motivation to score low on IQ tests or standardized measures of adaptive behavior -- lower scores may make the difference between execution or life in prison without parole. Not being an expert in this area of forensic assessment, I'm going to try provide information from high quality sources re: the state-of-the art of malingering assessment. Also, when appropriate, I will point out situations of inappropriate (unethical?) malingering assessment methods when they are obvious. This current post contains a sampling of interesting malingering issues, research, and an example of inappropriate malingering assessment. Click here for prior posts re: malingering issues, research and references.
What does the research say about malingering assessment in the context of intellectual disability determination?
I've found the scholarly work of Salekin and Doane to be of particular value in providing an evaluation of the research in this area. Below are two recent journal articles by Salekin and Doane. I believe the abstracts/summaries speak for themselves.
Doane, B., & Salekin, K. L. (2009). Susceptibility of current adaptive behavior measures to feigned deficits. Law and Human Behavior, 33, 329-343.
Abstract
The current study examined the susceptibility of the Adaptive Behavior Assessment System—2nd edition (ABAS-II; Harrison & Oakland, 2003) and the Scales of Independent Behavior—Revised (S1B-R; Bruininks, Woodcock, Weatherman, & Hill, 1996) to the feigning of adaptive functioning deficits. Using four different instruction sets, the authors evaluated whether the provision of diagnostic information (a form of coaching) improved participants’ ability to simulate adaptive deficits commensurate with a diagnosis of mental retardation. The authors found that the ABAS-II was quite vulnerable to believable manipulation by raters, while the SIB-R was not. In fact, exaggeration on the SIB-R was easily detected regardless of the information provided. Implications regarding the use of these measures in Atkins mental retardation evaluations are discussed.
Salekin, K. L., & Doane, B. (2009). Malingering intellectual disability: The value of available measures and methods. Applied Neuropsychology, 16, 105-113.
Abstract
Atkins v. Virginia (2002) is a case that has changed the landscape in relation to the assessment of malingering in a legal context. This landmark decision abolished the death penalty for defendants found to have intellectual disability (ID; formally known as mental retardation), but limitations in our assessment techniques lead to questions regarding the veracity of ID claims. In fact, Justice Scalia noted with clarity that concerns exist regarding the ability of individuals to feign ID and to do so successfully. At the time of writing, little empirical research has been completed, but that which exists demonstrates an overall lack of validity for traditional measures of cognitive malingering for use with this population. This manuscript provides an overview of the utility of many of the traditional measures of malingering for use with an ID population and serves as a call for research in this very important area.Summary
In closing, review of the research in the assessment of malingered ID demonstrates that effort tests and indices of cognitive malingering are not working with this population, and that true cases can be misidentified as malingered. Some would say that the inclusion of multiple measures of malingering and the interpretation of all of the data together, rather than tests in isolation, provide control for diagnostic error. But to date, we have no data to suggest that either of these techniques is protective and more importantly, we have no data on how a juror or a judge might be impacted by even the slightest mention of malingering. Though untested, these authors posit that it is very unlikely that a defense expert will succeed in supporting an Atkins claim if there is even a hint that malingering may have occurred.
Another interesting topic is malingering resulting from examiner bias. I find the concept of "malingering by proxy" very interesting. Below is a discussion of this phenomenon as described by Schlesinger:
Schlesinger, L. B. (2003). A case study involving competency to stand trial: Incompetent defendant, incompetent examiner, or "malingering by proxy" ? Psychology, Public Policy, and Law, 9 (3/4), 381-399.
The most blatant kind of examiner bias, however, is seen mostly in forensic cases: the deliberate, conscious intent to distort or misrepresent findings for partisan purposes. This sort of conduct is a breach of professional ethics (Committee on Ethical Guidelines for Forensic Psychologists, 1991), unlike the involuntary forms of bias resulting from patient attributes.
There is yet another variety of examiner bias that is not an automatic act of impaired judgment arising from patient demographics, nor is it an intentional falsification of results. Here, the forensic psychologist finds in the defendant (nonexistent) signs, symptoms, or disorders that were initially suggested by the referring attorney. External incentives (such as economic gain) are typically absent. The effect, which could be called “malingering by proxy,” derives from the forceful opinions of the legal advocate, which can be quite contagious. My impression is that this form of examiner bias is not an uncommon occurrence in forensic work, where the structure of relationships leaves the clinician particularly vulnerable to such (nonconscious) infection.
The genesis of this form of examiner bias begins when the clinician is first approached about the case. Most forensic referrals come from a lawyer who attempts to recruit the consultant for the defense or the prosecution team. For instance, an attorney might call and say:
- Hello Dr. Z; I was referred to you by a psychiatrist, Dr. Y. She told me you had worked with her on many cases. Your colleague regards you highly and said you are one of the top forensic psychologists in the area. I’d like to retain your services for help with a client I represent. Dr. Y. saw my client yesterday and thought he was mentally retarded. My law partner and I just came back from the county jail, and he seemed retarded to the both of us. We all think he is incompetent to stand trial. Can I count on you to be part of the defense team? By the way, don’t worry about your fee; my client’s family is very supportive of him, and they’ll be sure to pay you promptly.
After an introduction like this, some consultants may find it difficult to disregard the flattery or to challenge members of a “team” they are about to join. However, if forensic psychologists are not careful at this point, they could succumb to a form of examiner bias that could jeopardize the entire evaluation before they have even met the defendant.Junk science malingering assessment--from actual cases
In two recent Atkins court decisions in the state of Oklahoma (see Salazar, 2005 and Lambert, 2005), the states prosecution psychologist (Dr. Prosecution Psychologist - Dr. PP) testified re: malingering based, in part, on non-normed, non-standardized malingering measures that Dr. PP had developed himself, and one which he named after his secretary (in an attempt to mask the purpose of the test to the defendant). The two "instruments" in question were the non-standardized Blackwell Memory Test and the Oklahoma Spelling Test. Apparently the Blackwell Memory test was modelsx after other formal instruments that use a "forced choice symptom validity" test format. Similar to prior voodoo psychometric activities commented on at this blog, I'm dumb founded that a professional psychologist testifying in an Atkins hearing, or any other clinical or forensic setting, would attempt to assess a psychological construct (viz., malingering) via the development of their own special instrument that did not undergo the professional accepted and required test development procedures (as clearly spelled out the the Joint Test Standards). This activity clearly violates a number of professional standards. Below are at least two (and I'm sure there are more when one examines all relevant professional codes of ethics/standards) from the Joint Test Standards:
Standard 1.4: If a test is used in a way that has not been validated, it is incumbent on the user to justify the new use, collecting new evidence if necessaryUnbelievable.
Standard 11.2. When a test is to be used for a purpose for which little or no documentation is available, the user is responsible for obtaining evidence of the test's validity and reliability for this purpose.
Technorati Tags: psychology, forensic psychology, criminal psychology, criminal justice, neuropsychology, school psychology, educational psychology, ABA, American Bar Association, criminal defense lawyers, MR, ID, mental retardation, intellectual disability, IQ, IQ tests, IQ scores, intelligence, adaptive behavior, malingering, Joint Test Standards
Subscribe to:
Posts (Atom)




