Key points are not available for this paper at this time.
Abstract We constructed a shell (blueprint) for generating science performance assessments, and evaluated the characteristics of the assessments produced with it. The shell addressed four tasks: Planning, HandsOn, Analysis, and Application. Two parallel assessments were developed, Inclines (IN) and Friction (FR). Two groups of fifth graders who differed in both science curriculum experience and socioeconomic status took the assessments consecutively in either of two sequences, IN FR or FR IN. We obtained high interrater reliabilities for both assessments, statistically significant score differences due to assessment administration sequence, and a considerable task-sampling measurement error. For both assessments, the magnitude of score variation due to the hands-on task indicated that it tapped a kind of knowledge not addressed by the other three tasks. Although IN and FR were similar in difficulty, they correlated differently with an external measure of science achievement. Moreover, measurement error differed depending on assessment administration sequence. The results indicate that shells can produce reliable assessments, but do not solve the task-sampling variability problem or insure assessment exchangeability. We conclude that future shell research should focus on: (a) increasing shell precision, (b) improving shell usability, and (c) determining what specifications must be provided by the shell to ensure that the assessments generated by different developers are comparable.
Solano‐Flores et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: