In late 2025, the Michigan Department of Education reviewed K-3 reading screeners against the state’s Reading Foundational Skills (RFS) requirements. Of the 10 assessments submitted, 3 earned approval as valid and reliable K-3 screening and progress monitoring tools. The other 7 did not. The cards below summarize MDE’s findings rubric by rubric, so you can see exactly where each tool fell short. Vendor responses are shown in italic with an orange bar to separate them from MDE’s own commentary.
You can browse the compiled data in spreadsheet form here: MI Screener Reviews Compiled (Google Sheet).
All information on this page is sourced directly from the State of Michigan Department of Education’s K-12 Literacy and Dyslexia Law page and is a direct compilation of the content listed in the MDE’s K-3 Valid and Reliable Screening and Progress Monitoring Assessment List (PDF, March 2026). What I’ve done is download every consensus document linked from that single PDF and compile them into one searchable view.
Accurate as of publication date: May 7, 2026. I’ll do my best to keep this current as MDE adds or updates assessments, but no promises that it’s always up to the minute, so always double-check against the official sources above. I’m not a lawyer, and this isn’t legal or assessment-selection advice, just a parent-friendly compilation of public data.
MetScreening Overall
Met screening requirements
- 100% of screening criteria met General assessment information:
- Administration format: Individual, large group, computer-administered (whole class)
- Scoring: Automatically (computer-scored). The vendor noted that “Teachers score paper-based assessments using Amira’s scoring guidelines and manually enter results into the platform for tracking and reporting.”
- Administration time:
- Grade K: 15-18 minutes
- Grade 1: 15-20 minutes
- Grade 2: 15-20 minutes
- Grade 3: 15-20 minutes
MetElements
- Phonemic awareness: Measured by Phonological Awareness (K-3)
- Rapid automatized naming: Measured by Rapid Automatized Naming (K-3)
- Letter-sound correspondence: Measured by Letter Sounds (K-2)
- Single-word reading: Measured by Word Identification Fluency (K-3)
- Nonsense-word reading: Measured by Pseudoword Identification (K-3)
- Oral passage reading: Measured by Oral Reading Fluency (1-3)
- Elements that may be included (optional):
- Retelling: Measured by Listening Comprehension/Retell (K-3)
- Cloze reading procedure: Measured by Cloze (1-3)
- Answering questions about a reading passage: Measured by Reading Comprehension (1-3)
MetClassification Accuracy
- External criterion measure for classification accuracy analyses: NWEA MAP Reading (Fall, Winter, and Spring)
- Entered "0" for all cells across student demographics categories for the Screening Classification Accuracy Sample table at grade 3, including all “Unknown” rows.
- The vendor noted that “Demographic subgroup results for Grades K-2 were generated as part of a focused evaluation study. At the time of that analysis, Grade 3 data were not included in the scope of the study. Amira is currently working with partner districts to expand this disaggregation to include Grade 3, ensuring subgroup-level reporting across all early literacy grades in future analyses.”
MetReliability
- Disaggregated reliability analyses were not provided for students in grade 3 for any group.
- The vendor noted that “Amira Learning is committed to continuing to conduct comprehensive reliability analyses disaggregated by key demographic subgroups for all grades. This ongoing initiative will provide reliability estimates across gender, ethnicity, English language proficiency, socioeconomic status, and IEP status to ensure equitable assessment performance for all students. We will continue to collect data systematically and conduct analyses as demographic information becomes available. To facilitate this work, we will partner with the MI DOE to obtain demographic subgroup information, enabling robust reliability estimates for each student population.”
MetValidity
- External criterion measure(s) for Validity Analyses: NWEA MAP Growth Reading, i-Ready Reading Diagnostic
- Disaggregated validity analyses were not provided for the Socioeconomic Status group.
- The vendor noted that “Amira has already conducted extensive subgroup reliability analyses by SES and is in the process of incorporating SES- disaggregated validity metrics into its reporting pipeline.”
MetBias Analyses
- Bias analyses were conducted for Gender, Ethnicity, English Language Proficiency, and IEP status, demonstrating no differences in classification accuracy between at- risk and not-at-risk students across these groups. Bias analyses were not provided for the Socioeconomic Status group. The vendor noted that “…a dedicated bias analysis for socioeconomic status (SES) has not yet been conducted. Amira Learning has reported reliability data disaggregated by SES, demonstrating strong measurement consistency across economic subgroups… As new data become available, these subgroup analyses will be conducted using both classification accuracy methods and Differential Item Functioning (DIF) to ensure the tool remains valid, unbiased, and equitable for all learners in Michigan and nationally.”
MetAdditional Considerations
Met screening requirements
- 100% of Additional Considerations criteria were met. The threshold for approval is 90%.
n/aSupp: Classification Accuracy
N/A Did not submit supplemental screening information
n/aSupp: Reliability
N/A Did not submit supplemental screening information
n/aSupp: Validity
N/A Did not submit supplemental screening information
n/aSupp: Bias Analyses
N/A Did not submit supplemental screening information
Progress Monitoring Measures
n/aPhonological Awareness
- Number of alternate measures: 75 tasks for all grades
- Reliability: Expectations were met for reliability for Grades K, 2, and 3, and not met for Grade 1.
- Confidence Interval Lower Bound was not provided for Test-Retest reliability at Grade 1.
- The vendor is committed to conducting comprehensive reliability analyses disaggregated by key demographic subgroups for Grades K-3.
- Reliability of Slope: Expectations were not met for reliability of slope for Grades K-3.
- Reliability of slope analyses were not provided. The vendor-provided explanation and plans for future analyses described validity, not reliability of slope.
- External criterion measure(s) for Validity Analyses: NWEA MAP Reading Phonological Awareness
- Validity: Expectations were met for validity for Grades K-1, and not met for Grades 2-3.
- Validity data were not provided for students at Grades 2-3. The vendor noted that additional analyses will be conducted during 2025-26 school year.
- Disaggregated validity analyses were not provided. The vendor noted that future research will address this.
n/aLetter Knowledge
- Number of alternate measures: 26
- Reliability: Expectations were met for reliability for Grades K-3.
- The vendor is committed to conducting comprehensive reliability analyses disaggregated by key demographic subgroups for Grades K-3.
- Reliability of Slope: Expectations were met for reliability of slope for Grades K-3.
- Reliability of slope analyses were not provided, but the vendor provided an explanation and future plans for conducting analyses.
- External criterion measure(s) for Validity Analyses: NWEA MAP Phonics/Word Recognition
- Validity: Expectations were not met for validity for Grades K-3.
- Confidence Interval Lower Bounds were not provided for Predictive validity at Grades K-1.
- Validity data were not provided for students at Grades 2-3. The vendor noted that additional analyses will be conducted during 2025-26 school year.
- Disaggregated validity analyses were not provided. The vendor noted that future research will address this.
n/aAlphabetic Decoding
- Number of alternate measures: 20
- Reliability: Expectations were met for reliability for Grades K-3.
- The vendor is committed to conducting comprehensive reliability analyses disaggregated by key demographic subgroups for Grades K-3.
- Reliability of Slope: Expectations were met for reliability of slope for Grades K-3.
- Reliability of slope analyses were not provided, but the vendor provided an explanation and future plans for conducting analyses.
- External criterion measure(s) for Validity Analyses: none
- Validity: Expectations were not met for validity for Grades K-3.
- Validity data were not provided for students at Grades K-3. The vendor noted that additional analyses will be conducted during 2025-26 school year.
- Disaggregated validity analyses were not provided. The vendor noted that future research will address this.
MetPseudoword Identification
- Number of alternate measures: 15 words per grade
- Reliability: Expectations were met for reliability for Grades K-3.
- The vendor is committed to conducting comprehensive reliability analyses disaggregated by key demographic subgroups for Grades K-3.
- Reliability of Slope: Expectations were met for reliability of slope for Grades K-3.
- Reliability of slope analyses were not provided, but the vendor provided an explanation and future plans for conducting analyses.
- External criterion measure(s) for Validity Analyses: NWEA MAP Phonics/Word Recognition
- Validity: Expectations were met for validity for Grade 1, and not met for Grades K, 2, and 3.
- Response included a description of validity analyses for RAN, not Pseudoword Identification. However, responses to all other questions related to validity analyses refer to Pseudoword Identification.
- Validity data were not provided for students at Grades K and 2-3. The vendor noted that additional analyses will be conducted during 2025-26 school year.
- Disaggregated validity analyses were not provided. The vendor noted that future research will address this.
MetFluency
- Number of alternate measures: 20+ per grade level
- Reliability: Expectations were met for reliability for Grades K-3.
- The vendor is committed to conducting comprehensive reliability analyses disaggregated by key demographic subgroups for Grades K-3.
- Reliability of Slope: Expectations were met for reliability of slope for Grades K-3.
- Reliability of slope analyses were not provided, but the vendor provided an explanation and future plans for conducting analyses.
- External criterion measure(s) for Validity Analyses: NWEA MAP Oral Reading Fluency
- Validity: Expectations were met for validity for Grade 1, and not met for Grades K, 2, and 3.
- Validity data were not provided for students at Grades K and 2-3. The vendor noted that additional analyses will be conducted during 2025-26 school year.
- Disaggregated validity analyses were not provided. The vendor noted that future research will address this.
n/aCloze
- Number of alternate measures: 20 items per grade
- Reliability: Expectations were met for reliability for Grades 1-3.
- The vendor is committed to conducting comprehensive reliability analyses disaggregated by key demographic subgroups for Grades 1-3.
- Reliability of Slope: Expectations were met for reliability of slope for Grades 1- 3.
- Reliability of slope analyses were not provided, but the vendor provided an explanation and future plans for conducting analyses.
- External criterion measure(s) for Validity Analyses: none
- Validity: Expectations were not met for validity for Grades 1-3.
- Validity data were not provided for students at Grades 1-3. The vendor noted that additional analyses will be conducted during 2025-26 school year.
- Disaggregated validity analyses were not provided. The vendor noted that future research will address this.
n/aReading Comprehension
- Number of alternate measures: 55
- Reliability: Expectations were met for reliability for Grades 1-3, and not met for Grade K.
- Confidence Interval Lower Bound Test-Retest reliability at Grade K fell below the preferred value of 0.60.
- The vendor is committed to conducting comprehensive reliability analyses disaggregated by key demographic subgroups for Grades K-3.
- Reliability of Slope: Expectations were met for reliability of slope for Grades K-3.
- Reliability of slope analyses were not provided, but the vendor provided an explanation and future plans for conducting analyses.
- External criterion measure(s) for Validity Analyses: none
- Validity: Expectations were not met for validity for Grades K-3.
- Validity data were not provided for students at Grades K-3. The vendor noted that additional analyses will be conducted during 2025-26 school year.
- Disaggregated validity analyses were not provided. The vendor noted that future research will address this.
MetAmira ISIP Progress Monitoring
reading fluency, Retelling, Cloze reading procedure, · Number of alternate measures: 20 Answering questions about
- Administration time is estimated to be approximately 12-22 minutes but is a reading passage noted as being grade-configurable.
- Reliability: Expectations were met for reliability for Grades K-3.
- Disaggregated reliability analyses were not provided for Gender or Ethnicity.
- Reliability of Slope: Expectations were met for reliability of slope for Grades K-3.
- External criterion measure(s) for Validity Analyses: i-Ready Reading Diagnostic, NWEA MAP Reading
- Validity: Expectations were met for validity for Grades K-3.
- Disaggregated validity analyses were not provided for Gender or Ethnicity.
Not MetScreening Overall
Did not meet screening requirements
- 33% of screening criteria were met. The threshold for approval is 100%. General assessment information:
- Administration format: Individual, small group (n=4), large group (n=20), computer-administered
- Scoring: “Reading Assistant leverages machine learning and AI to automatically score every student interaction.”
- Administration time:
- Grade K: 10-15 minutes
- Grade 1: 10-15 minutes
- Grade 2: 10-15 minutes
- Grade 3: 10-15 minutes
Not MetElements
- The vendor did not respond within 24 hours to follow-up clarifications for RFS questions 2.1, 2.2, 2.3, and 2.4.
- The vendor was asked to provide a detailed explanation regarding the ways in which EPS Reading Assistant Universal Screener and Assessment is the same and different from the Amira Learning Program, given their statement that “Reading Assistant Universal Screener and Assessment is a branded version of Amira Learning program.”
- Phonemic awareness: Measured by Blending, Segmentation & Deletion, Substitution (K); Blending, Segmentation & Elision, Substitution (1-3). The vendor did not provide the number of items students are typically presented with for each of the four phonemic awareness tasks. The response did not include information about differences/changes in tasks that vary by grade or time of year.
- Rapid automatized naming: Measured by RAN (K-3). The vendor did not clarify if all forms of RAN (numbers, colors, and objects) are scored the same and considered to be equivalent
- Letter-sound correspondence: Measured by Word Identification (K-3). The response did not include information about 1)
differences/ changes in tasks that vary by grade or time of year, 2) how long students have to respond, and 3) the number of items students are typically presented with for each of the tasks.
- Single-word reading: The vendor used inconsistent names when referring to measures of single-word reading. The response did not include information about 1) differences/ changes in tasks that vary by grade or time of year, 2) how long students have to respond, and 3) the number of items students are typically presented with for each of the tasks.
- Nonsense-word reading: The vendor used inconsistent names when referring to measures of nonsense-word reading (1-3, K can be added). The response did not include information about 1) differences/ changes in tasks that vary by grade or time of year, 2) how long students have to respond, 3) the number of items students are typically presented with for each of the tasks, and 4) how scores are calculated for each correct sound vs. the whole pseudo-word.
- Oral passage reading: The response did not include sufficient information to determine if ORF (1-3, K can be added) met the element requirements. The response was missing information about
- the approximate length of each passage and each of the three passage parts, and 2) a description of the rules that cause the program to determine when a student is struggling to read text fluently and subsequently adjust the text to a lower level, as well as how the program shifting the text reading level influences a student’s score.
- Elements that may be included (optional):
- Retelling: Not measured
- Cloze reading procedure: Not measured
- Answering questions about a reading passage: The vendor used inconsistent names when referring to measures to answer questions about a reading passage (1-3). The response was missing information about 1) what students see or hear during the subtest, 2) how students are asked to respond, and 3) the number of items students are typically presented with for the comprehension task.
Not MetClassification Accuracy
- External criterion measure for classification accuracy analyses: The response to item 4.3 names NWEA MAP Growth assessment and Hasbrouck & Tindal WCPM norms (but does not specify grade levels for each), which conflicts with the response to 4.1 and 4.2. Hasbrouck and Tindal norms are also not referenced elsewhere in the submission.
- Percentile ranks provided in response to item 4.6 were inconsistent, one referencing the 30th percentile and another the 20th
percentile. The vendor did not clarify which percentile cutpoint was used for their spring criterion measure.
- AUC values were entered as > or = rather than a precise value.
- Inconsistencies across responses to items 4.1, 4.2, 4.3, 4.4, 4.5, and 4.6, making it unclear which percentile cutpoint was used for the spring criterion measure.
- In addition, many ARM scores provided were inexplicably lower than expected for the time of year and grade given the scores for other grades and times of year.
MetReliability
- The vendor did not include a justification for inter-rater reliability as selected in 5.1, and did not explain whether they have compared the program scoring using machine learning with human scoring.
- Parallel Forms Reliability as entered in item 5.1, was not described, including how it was similar or different to test-retest and internal consistency.
- Responses were missing a description of the analysis procedures for inter-rater reliability and Parallel Forms Reliability as entered in item 5.1.
- The vendor indicated that they did conduct disaggregated reliability analyses, but did not provide the specific results of disaggregated reliability analyses.
Not MetValidity
- External criterion measure(s) for Validity Analyses: MAP Growth Reading Assessment
- The response for item 6.1 states that classification accuracy data are based on MAP Growth at the 20th percentile. The vendor did not clarify if they have updated analyses aligned with the RFS requirement of analyses based on the 30th-40th percentiles for approval.
- The vendor did not provide the results of disaggregated validity analyses. The response referred to disaggregated reliability analyses.
MetBias Analyses
- The vendor did not clarify why this statement is only made for English Learners: “Reliability and validity metrics for ELLs are reported separately” and why other group metrics are not reported separately.
Not MetAdditional Considerations
- The vendor did not respond to the last three parts of question 8.4 and did not distinguish between reports that are suited for administrators vs. teachers.
- The vendor did not provide links to sample user reports.
- The responses did not describe compatibility with the MDE Early Literacy Benchmark Assessments.
n/aSupp: Classification Accuracy
N/A Did not submit supplemental screening information
n/aSupp: Reliability
N/A Did not submit supplemental screening information
n/aSupp: Validity
N/A Did not submit supplemental screening information
n/aSupp: Bias Analyses
N/A Did not submit supplemental screening information
Not MetScreening Overall
Did not meet screening requirements
- 50% of screening criteria were met. The threshold for approval is 100%. General assessment information:
- Administration format: The vendor states, “The i-Ready Diagnostic, an adaptive assessment, can be administered in a group setting. The additional Literacy Task is administered in a one-one-one (individual) setting.”
- Scoring: The vendor states, “Manually, automatically, and by entry of scores in the student’s digital record in the user interface after administration • diagnostic is auto scored, additional literacy task is scored by teacher- can be digitally entered.”
- Administration time:
- Grade K: 26-36 minutes
- Grade 1: 26-51 minutes
- Grade 2: 42-75 minutes
- Grade 3: 42-75 minutes
Not MetElements
- The i-Ready Assessment submission included the i-Ready Diagnostic and i- Ready Literacy Task for Fluency. The vendor states, “Students first complete our adaptive i-Ready Diagnostic assessment and then complete one additional Literacy Task fluency assessment. Although there is only one single additional assessment administered beyond the adaptive Diagnostic, the content of that single assessment varies across grades and times of year to ensure our measure of fluency is developmentally appropriate and aligned to the most current research on early literacy screening.”
- Phonemic awareness: Element requirements were not met per the description provided for Diagnostic (K-1). It is unclear what students see and hear in all ways that phonemic awareness is assessed. For example, it is unclear how segmenting is assessed by providing students with answer choices.
- Rapid automatized naming: Measured by Fluency Task (K).
- Letter-sound correspondence: Element requirements were not met per the description provided for Diagnostic (K-2). For K and 1 specifically, it is unclear which type of Phonics (PH) task most students encounter when they start the school year, and then how the adaptive nature of the assessment works for presenting other items.
- Single-word reading: Measured by Fluency Task: Word Recognition (Fall of 1). Meets the element definition only for the Fluency task at first grade, not for the Diagnostic at any grade level.
- Nonsense-word reading: Element requirements were not met per the description provided for Diagnostic (K-2). The vendor states, “Approximate number of items/words students are presented with: Students are presented with 12 items in the PH domain, which will focus on a progression of phonics skills depending on student performance, from letter-sound recognition to letter-sound correspondence to decoding of real and nonsense words to encoding of real words.” It is unclear whether all students, especially those with low Phonics (PH) skills, would encounter nonsense word tasks as required by MCL 380.1280f.
- Oral passage reading: Measured by Fluency Task (1 winter and spring only, 2-3).
- Elements that may be included (optional):
- Retelling: Measured by Fluency Task (1 winter and spring only, 2-3)
- Cloze reading procedure: Not measured
- Answering questions about a reading passage: Measured by Diagnostic (2-3). Answering questions about a reading passage: Meets the element definition for grades 2-3. At K-1, Diagnostic can be a measure of listening comprehension rather than primarily reading comprehension. It is unclear if all first graders would need to complete reading comprehension tasks in addition to or in place of listening comprehension.
Not MetClassification Accuracy
- External criterion measure for classification accuracy analyses: Dynamic Indicators of Basic Early Literacy Skills® (DIBELS) NEXT (K-2), Smarter Balanced Assessment (SBA) ELA/Literacy test for grades 3-8 (3)
- Vendors are asked to provide classification accuracy data for cutpoints between the 30th and 40th percentile on a spring criterion measure in order to identify students who display early signs of reading difficulties. The provided classification data were based on cut scores between the 10th and 20th percentiles.
- Vendors are asked to provide the number of students in their classification analyses sample, disaggregated by five groups. The vendor entered “n/a” for the following groups: Native Hawaiian or Other Pacific Islander, Two or More Ethnicities, Students Designated as Economically Disadvantaged, Students Not Designated as Economically Disadvantaged, Socioeconomic Status Unknown, Students with an IEP, Students without an IEP, IEP Status Unknown, Students designated as English Learners, Students not designated as English Learners, and English Language Proficiency Unknown. The vendor outlines a plan for future analyses without specifying a timeline.
MetReliability
Met screening requirements
- 100% of Overall/Composite Score Reliability criteria were met. The threshold for approval is 92%.
Not MetValidity
- External criterion measure for validity analyses: Concurrent-i-Ready Grade 3 and ACAP, predictive-Grade 3 and 4 SBA. It appears that some responses in for the concurrent and predictive validity measures were entered in the incorrect spaces in the RFS response.
- Disaggregated Validity: The vendor stated, “As part of our data collection and analysis efforts for concurrent and predictive validity studies, Curriculum Associates will generate validity coefficients, both overall and disaggregated by student group for groups with sufficient sample size, beginning with research conducted in 2026.”
MetBias Analyses
- The vendor explained the Differential Item Functioning (DIF) process but did not share the results or provide interpretive statements.
MetAdditional Considerations
- The vendor did not describe the compatibility between i-Ready and Smarter Balanced Interim Assessments, MAP Suite, and Star Assessments outside of a hyperlinked crosswalk. Vendors are asked to provide complete responses within the RFS. Michigan’s Early Literacy Benchmark Assessment was not addressed in the vendor’s response.
n/aSupp: Classification Accuracy
N/A Did not meet overall/composite screening requirements
n/aSupp: Reliability
N/A Did not meet overall/composite screening requirements
n/aSupp: Validity
N/A Did not meet overall/composite screening requirements
n/aSupp: Bias Analyses
N/A Did not meet overall/composite screening requirements
MetScreening Overall
Met screening requirements
- 100% of screening criteria met General assessment information:
- Administration format: Individual, Small group, and Large group (any size group). The vendor noted that "The mCLASS observational platform allows teachers to capture student responses on their mobile device, and provides automatic scoring of the assessments. On the scoring screen, the teacher captures the student performance directly.”
- Scoring: Computer scored/automatically
- Administration time:
- Grade K: 2-4 minutes
- Grade 1: 3-5 minutes
- Grade 2: 5-8 minutes
- Grade 3: 5-8 minutes
MetElements
Phonemic awareness: Measured by Phonemic Segmentation Fluency (K-1) Rapid automatized naming: Measured by Letter Naming Fluency (K-1) Letter-sound correspondence: Measured by Nonsense Word Fluency (K-3) Single-word reading: Measured by Word Reading Fluency (K-3) Nonsense-word reading: Measured by Nonsense Word Fluency (K-3) Oral passage reading: Measured by Oral Reading Fluency (1-3) Elements that may be included (optional): Retelling: Not measured Cloze reading procedure: Measured by Maze (2-3) Answering questions about a reading passage: Not measured
MetClassification Accuracy
- External criterion measure for classification accuracy analyses: DIBELS Next Composite Score (K), Iowa Total Reading Score (grades 1-3)
- For third grade in the fall, the sensitivity of the cutpoint is 0.71, and the specificity is 0.69. The 0.69 specificity value is the only classification accuracy value that fell below the required range, so the assessment still meets screening approval requirements (one or fewer unmet classification accuracy values).
- The vendor did not explain why there were zero Kindergarten American Indian or Alaska Native students in the sample, and did not explain why there were zero Kindergarten and third-grade Asian students in the sample. The vendor entered “N/A” in grade-level sample size cells for many groups (Native Hawaiian or Other Pacific Islander, Ethnicity Other, Two or More Ethnicities, Students Designated as Economically Disadvantaged, Students Not Designated as Economically Disadvantaged, Socioeconomic Status Unknown, Students with an IEP, Students without an IEP, IEP Status Unknown, Students designated as English Learners, Students not designated as English Learners, English Language Proficiency Unknown).
- The vendor provided the following explanation: “We are currently planning a study to strengthen the reliability, validity, and classification accuracy evidence of DIBELS 8th Edition with a larger, more diverse sample, particularly with respect to race/ethnicity, EL status, and students receiving Special Education services. Our recruitment efforts will include purposefully reaching out to school districts serving large proportions of students from these traditionally underrepresented groups for inclusion in the sample. Participation in the study will require that district/schools be willing to share the aforementioned demographic data. As part of these data collection efforts we will provide participating schools with a template to provide English language proficiency data for participating ELs to allow us to explore the possibility of differential performance on DIBELS 8th Edition for ELs with varying levels of English language proficiency. This study is planned for school year 2025-2026, with analyses to be conducted in Summer/Fall of 2026 and documentation of the results to be released soon thereafter.”
MetReliability
Met screening requirements
- 100% of Overall/Composite Score Reliability criteria were met. The threshold for approval is 92%.
MetValidity
Met screening requirements
- 100% of Overall/Composite Score Validity criteria were met. The threshold for approval is 90%.
MetBias Analyses
Met screening requirements
- 100% of Overall/Composite Score Bias Analysis criteria were met. The threshold for approval is 100%.
MetAdditional Considerations
- The vendor did not directly describe compatibility with Smarter Balanced Interim Assessments, MAP Suite, Star Assessments, or Early Literacy Benchmark Assessments. Rather, the response highlighted the features of mCLASS with DIBELS 8th Edition compared to i-Ready, specifically, and computer-administered assessments in general, such as the approved Benchmark Assessments.
n/aSupp: Classification Accuracy
N/A Did not submit supplemental screening information
n/aSupp: Reliability
N/A Did not submit supplemental screening information
n/aSupp: Validity
N/A Did not submit supplemental screening information
n/aSupp: Bias Analyses
N/A Did not submit supplemental screening information
Progress Monitoring Measures
n/aPhonemic Segmentation Fluency (PSF)
- Number of alternate measures: 20
- Reliability: Expectations were met for reliability for Grades K-1.
- The vendor plans to conduct disaggregated reliability analyses of PSF data by 1) gender, 2) ethnicity, 3) English language proficiency, 4) socioeconomic status, and 5) IEP status using operational data during the Fall of 2026 with complete results available in the Winter of 2026.
- Reliability of Slope: Expectations were met for reliability of slope for Grades K-1.
- External criterion measure(s) for Validity Analyses: “The concurrent validity of DIBELS 8 was evaluated by correlating its subtests with external criterion measures given in the same benchmark period. These measures included DIBELS Next composite scores, the Phonological Processing-2nd Edition (CTOPP-2) composite scores, and Iowa Assessment Total Reading and Word Analysis raw scores. Predictive correlations: Depending on the grade, DIBELS 8 measures were correlated with end of year administrations of the DIBELS Next Composite, the Total Reading and Word Analysis scores from the Iowa Assessment, and the CTOPP-2 symbolic and non-symbolic composite scores.”
- Validity: Expectations were not met for validity for Grades K-1.
- Coefficients for concurrent and predictive validity of the progress monitoring measures fell below the preferred value of 0.60 for Grades K-1.
- The vendor is “currently planning a study to strengthen the reliability, validity, and classification accuracy evidence of DIBELS 8th Edition with a larger, more diverse sample, particularly with respect to race/ethnicity, EL status, gender, and students receiving Special Education services. Recruitment efforts will include purposefully reaching out to school districts serving large proportions of students from these traditionally underrepresented groups for inclusion in the sample. This study is planned for the 2025-2026 school year, with analyses to be conducted in Summer and Fall of 2026 and documentation of the results to be released soon thereafter.”
MetNonsense Word Fluency (NWF)
- Number of alternate measures: 20
- Reliability: Expectations were met for reliability for Grades K-3.
- The vendor plans to conduct disaggregated reliability analyses of NWF data by 1) gender, 2) ethnicity, 3) English language proficiency,
- socioeconomic status, and 5) IEP status using operational data during the Fall of 2026 with complete results available in the Winter of 2026.
- Reliability of Slope: Expectations were met for reliability of slope for Grades K-3.
- The Reliability of Slope 95% CI Lower Bound was less than the preferred value of 0.40 at third grade, but was above 0.40 for grades K, 1, and 2.
- External criterion measure(s) for Validity Analyses: “The concurrent validity of DIBELS 8 was evaluated by correlating its subtests with external criterion measures given in the same benchmark period. These measures included DIBELS Next composite scores and Iowa Assessment Total Reading and Word Analysis raw scores. Predictive correlations: Depending on the grade, DIBELS 8 measures were correlated with end of year administrations of the DIBELS Next Composite and the Total Reading and Word Analysis scores from the Iowa Assessment.”
- Validity: Expectations were met for validity for Grades K-3.
- The vendor is “currently planning a study to strengthen the reliability, validity, and classification accuracy evidence of DIBELS 8th Edition with a larger, more diverse sample, particularly with respect to race/ethnicity, EL status, gender, and students receiving Special Education services. Recruitment efforts will include purposefully reaching out to school districts serving large proportions of students from these traditionally underrepresented groups for inclusion in the sample. This study is planned for the 2025-2026 school year, with analyses to be conducted in Summer and Fall of 2026 and documentation of the results to be released soon thereafter.”
MetWord Reading Fluency (WRF)
- Number of alternate measures: 20
- Reliability: Expectations were met for reliability for Grades K-3.
- The vendor plans to conduct disaggregated reliability analyses of WRF data by 1) gender, 2) ethnicity, 3) English language proficiency,
- socioeconomic status, and 5) IEP status using operational data during the Fall of 2026 with complete results available in the Winter of 2026.
- Reliability of Slope: Expectations were met for reliability of slope. for Grades K-1, and not met for grades 2-3.
- The Reliability of Slope 95% CI Lower Bound was less than the preferred value of 0.40 at Grades 2-3.
- External criterion measure(s) for Validity Analyses: The concurrent validity of DIBELS 8 was evaluated by correlating its subtests with external criterion measures given in the same benchmark period. These measures included DIBELS Next composite scores and Iowa Assessment Total Reading and Word Analysis raw scores. Predictive correlations: Depending on the grade, DIBELS 8 measures were correlated with end of year administrations of the DIBELS Next Composite and the Total Reading and Word Analysis scores from the Iowa Assessment.
- Validity: Expectations were met for validity for Grades K-3.
- The vendor is currently “planning a study to strengthen the reliability, validity, and classification accuracy evidence of DIBELS 8th Edition with a larger, more diverse sample, particularly with respect to race/ethnicity, EL status, gender, and students receiving Special Education services. Recruitment efforts will include purposefully reaching out to school districts serving large proportions of students from these traditionally underrepresented groups for inclusion in the sample. This study is planned for the 2025-2026 school year, with analyses to be conducted in Summer and Fall of 2026 and documentation of the results to be released soon thereafter.”
MetOral Reading Fluency(ORF)
- Number of alternate measures: 20
- Reliability: Expectations were met for reliability for Grades 1-3.
- The vendor plans to conduct disaggregated reliability analyses of ORF data by 1) gender, 2) ethnicity, 3) English language proficiency,
- socioeconomic status, and 5) IEP status using operational data during the Fall of 2026 with complete results available in the Winter of 2026.
- Reliability of Slope: Expectations were met for reliability of slope for Grades 1-2, and not met for Grade 3.
- The Reliability of Slope 95% CI Lower Bound was less than the preferred value of 0.40 at Grade 3.
- External criterion measure(s) for Validity Analyses: “The concurrent validity of DIBELS 8 was evaluated by correlating its subtests with external criterion measures given in the same benchmark period. These measures included DIBELS Next composite scores and Iowa Assessment Total Reading and Word Analysis raw scores. Predictive correlations: Depending on the grade, DIBELS 8 measures were correlated with end of year administrations of the DIBELS Next Composite and the Total Reading and Word Analysis scores from the Iowa Assessment.”
- Validity: Expectations were met for validity for Grades 1-3.
- The vendor is “currently planning a study to strengthen the reliability, validity, and classification accuracy evidence of DIBELS 8th Edition with a larger, more diverse sample, particularly with respect to race/ethnicity, EL status, gender, and students receiving Special Education services. Recruitment efforts will include purposefully reaching out to school districts serving large proportions of students from these traditionally underrepresented groups for inclusion in the sample. This study is planned for the 2025-2026 school year, with analyses to be conducted in Summer and Fall of 2026 and documentation of the results to be released soon thereafter.”
MetMaze
- Number of alternate measures: Ten forms are available for Maze, as the vendor does not recommend using this PM more than monthly.
- Reliability: Expectations were met for reliability for Grades 2-3.
- The vendor plans to conduct disaggregated reliability analyses of Maze data by 1) gender, 2) ethnicity, 3) English language proficiency, 4) socioeconomic status, and 5) IEP status using operational data during the Fall of 2026 with complete results available in the Winter of 2026.
- Reliability of Slope: Expectations were met for reliability of slope for Grades 2-3.
- Reliability of Slope results are not available for Maze. The vendor provided a rationale related to the fewer number of Maze alternate forms.
- External criterion measure(s) for Validity Analyses: “The concurrent validity of DIBELS 8 was evaluated by correlating its subtests with external criterion measures given in the same benchmark period. These measures included DIBELS Next composite scores and Iowa Assessment Total Reading and Word Analysis raw scores. Predictive correlations: Depending on the grade, DIBELS 8 measures were correlated with end of year administrations of the DIBELS Next Composite and the Total Reading and Word Analysis scores from the Iowa Assessment.”
- Validity: Expectations were met for validity for Grade 2, and not met for Grade 3.
- Coefficients for concurrent (0.59) and predictive validity (0.59) of Maze progress monitoring fell just below the preferred value of 0.60 for Grade 3.
- The vendor is “currently planning a study to strengthen the reliability, validity, and classification accuracy evidence of DIBELS 8th Edition with a larger, more diverse sample, particularly with respect to race/ethnicity, EL status, gender, and students receiving Special Education services. Recruitment efforts will include purposefully reaching out to school districts serving large proportions of students from these traditionally underrepresented groups for inclusion in the sample. This study is planned for the 2025-2026 school year, with analyses to be conducted in Summer and Fall of 2026 and documentation of the results to be released soon thereafter.”
Acadience Reading K-6 (with Acadience Rapid Automatized Naming and Acadience Decodable Text)
Not MetScreening Overall
Did not meet screening requirements
- 50% of screening criteria were met. The threshold for approval is 100%. General assessment information:
- Administration format: Individual, Small group, and Large group. The vendor noted that "Paper-based administration is available for all measures; assessors may use web application for digital scoring, but the student experience is not on computer."
- Scoring: Manually (by hand) and Automatically (computer-scored)
- Administration time:
- Grade K: 5-9 minutes
- Grade 1: 7-10 minutes
- Grade 2: 6-8 minutes
- Grade 3: 11 minutes
Not MetElements
- Phonemic awareness: Measured by First Sound Fluency (K), Phoneme Segmentation Fluency (K-1)
- Rapid automatized naming: Measured by Rapid Automatized Naming (K- 1)
- Acadience RAN was identified as not being included in the assessment’s overall/composite score. Therefore, the Rapid automatized naming element would need to be addressed via the Supplemental Screening Assessment questions (11.1-14.4). The
responses provided for these questions included data for both Acadience RAN and Acadience Decodable Text together.
- Letter-sound correspondence: Measured by Nonsense Word Fluency Correct Letter Sounds (K-2)
- Single-word reading: Measured by Nonsense Word Fluency Whole Words Read (K-1), Decodable Text (K-1), Oral Reading Fluency (1)
- Three subtests (Nonsense Word Fluency Whole Words Read, Decodable Text, Oral Reading Fluency) were identified as addressing the required Single-word reading element across grades K-1. In the response to Q2.2, the vendor cited one study as a rationale for not having a measure of single-word reading that meets the element definition for Single-word reading as stated in the RFS Glossary: “Assessment in which students read lists of stand-alone real words (i.e., not in context or connected text).” However, the review team determined that the element requirement was not met.
- Nonsense-word reading: Measured by Nonsense Word Fluency (K-2)
- Oral passage reading: Measured by Oral Reading Fluency (1-3)
- Elements that may be included (optional):
- Retelling: Measured by ORF Retell (1-3)
- Cloze reading procedure: Not Measured There was inconsistency across responses for questions 2.1- 2.3 regarding the inclusion of Maze in addressing the Cloze reading procedure element.
- Answering questions about a reading passage: Not Measured
Not MetClassification Accuracy
- External criterion measure for classification accuracy analyses:
- Fall: Acadience Reading Composite Score Middle of Year Benchmark Goal
- Winter: Acadience Reading Composite Score End of Year Benchmark Goal
- Spring: GRADE
- The criterion measures used for the Fall and Winter analyses were determined to be over-aligned to the assessment.
- No Confidence Interval Lower and Upper Bounds for the AUC estimate were provided for the Spring analyses.
MetReliability
- The Confidence Interval Lower Bound for the internal consistency (alpha) median coefficient at grade 3 was below the threshold of 0.60. However, all
grade levels met the reliability requirement by having a minimum of 2 acceptable forms of reliability provided.
Not MetValidity
- External criterion measure(s) for Validity Analyses: DIBELS 8th Edition, GRADE, CTS, CST
- The criterion measures used for the Predictive and Concurrent Validity analyses at grades 2-3 were unclear.
- Specific sample sizes for the analyses were not provided.
- The provided plan for conducting future disaggregated analyses was vague.
MetBias Analyses
- The vendor noted that a study is ongoing and results are anticipated in Fall 2026.
MetAdditional Considerations
- There appear to be no currently available disaggregation options for assessment reports. The vendor noted that "At this time, there are no data visualizations that include demographic information in the application. However, data views and reports that include demographic information are planned for future development."
- It is unclear if the following portion of the response to question 8.8 refers specifically to Michigan’s Early Literacy and Mathematics Benchmark Assessments: “Acadience aligns well with these tools, offering developmentally appropriate measures and empirically validated benchmarks for grades K-6.”
n/aSupp: Classification Accuracy
- The Supplemental Screening data provided for questions 11.1-14.4 are intended to cover only one additional required element. The response to question 2.3 indicated that the Rapid automatized naming element is not included in the composite/overall score and would thus be eligible for addressing in questions 11.1-14.4. However, the submitted responses included information for both Acadience RAN and Acadience Decodable Text.
n/aSupp: Reliability
N/A Did not meet overall/composite screening requirements
n/aSupp: Validity
N/A Did not meet overall/composite screening requirements
n/aSupp: Bias Analyses
N/A Did not meet overall/composite screening requirements
Not MetScreening Overall
Did not meet screening requirements
- 50% of screening criteria were met. The threshold for approval is 100%. General assessment information:
- Administration format: Individual. The vendor noted that "Each subtest is administered directly by the examiner; no subtest is administered directly by a computer."
- Scoring: Manually (by hand) and Automatically (computer-scored). The vendor noted that "The NLM Reading and NLM Listening has AI technology integrated if the user chooses to use it (this is entirely optional)."
- Administration time: 5-10 minutes for all grades K-3
MetElements
- Phonemic awareness: Measured by Dynamic Decoding Measures (DDM) Phonemic Awareness (K-3), DDM Phoneme Manipulation (K-3)
- Based on the response to question 2.2, it was unclear how these two subtests may change across grade levels.
- Rapid automatized naming: Measured by Rapid Automatized Naming (K- 3)
- Letter-sound correspondence: Measured by DDM Orthographic Mapping (K-3)
- The Letter Names task included in the DDM Orthographic Mapping subtest did not align with the Letter-sound correspondence element definition provided in the RFS Glossary: “The knowledge of the relationship between letters and the sounds they represent.” However, the Letter Sounds task included in the DDM Orthographic Mapping subtest meets the definition.
- Single-word reading: Measured by DDM Orthographic Mapping (K-3)
- Based on the response to question 2.2, it was unclear how this subtest may change across grade levels.
- Nonsense-word reading: Measured by DDM Decoding Inventory (K-3)
- Oral passage reading: Measured by Narrative Language Measures (NLM) Reading: Decoding Fluency (1-3)
- Elements that may be included (optional):
- Retelling: Measured by NLM Listening NLM Retell (K-3), NLM Reading NLM Retell (1-3)
- Cloze reading procedure: Measured by NLM Listening: NLM Questions-Inferential Vocabulary (K-3), NLM Reading: NLM Questions-Inferential Vocabulary (1-3) The provided description of the task (Inferential Vocabulary) addressing the optional Cloze reading procedure element did not align with the definition as stated in the RFS Glossary: “An objective reading assessment that deletes words in a designed reading passage.”
- Answering questions about a reading passage: Measured by NLM Listening: NLM Questions-Factual, Inferential Vocabulary, Inferential Reasoning (K-3), NLM Reading: NLM Questions: Factual, Inferential Vocabulary, Inferential Reasoning (1-3) The NLM Listening: NLM Questions task did not align with the Answering questions about a reading passage element definition provided in the RFS Glossary: “An assessment in which students respond to questions presented after reading a passage.” However, the NLM Reading: NLM Questions task meets the definition.
- The assessment is not available in Spanish for grades K-3.
Not MetClassification Accuracy
- External criterion measure for classification accuracy analyses: MAP, DIBELS, Acadience, CELF-P, TOWRE, TILLS, PELI, Bus Story, narrative and expository writing samples, narrative language samples (analyzed using SALT), Eligibility for special education
- In the responses to questions 4.1-4.2, multiple criterion measures were identified and described for each grade level. While noting that the descriptions provided were thorough and demonstrated non- overalignment with the screening assessment, it was unclear how the criterion measures were connected to the data provided in the tables for questions 4.4-4.6. For example, the values entered for the AUC Confidence Interval Lower Bounds across all grades and times were sufficient, but it was difficult to determine how to interpret them.
- Regarding additional data collection in the response to question 4.8 (sample representativeness), the vendor noted that “…ongoing national data analysis (N > 10,000 students) will provide even stronger empirical support for subgroup stability in predictive accuracy.”
Not MetReliability
- The descriptions of analyses in questions 5.2-5.4 and the data provided in the table for question 5.5 revealed inconsistencies and missing data:
- Test-retest reliability was selected for question 5.1 and data were provided in the table for question 5.5, but this analysis was not addressed in the responses to questions 5.2 or 5.4.
- Confidence Interval Upper and Lower Bounds were not provided for all reported analyses (and, if applicable, their lack was not noted or justified in the response to question 5.4).
- Administration fidelity was reported as a form of reliability; while acknowledging the importance of fidelity, it may be better categorized as a measure of assessor reliability.
- The representativeness of the sample across all performance levels was not directly addressed in the response to question 5.3.
- In response to question 5.6, the vendor noted that “We examined reliability by ethnicity, English language proficiency, socioeconomic status (SES), and IEP status for both the CUBED-3. Inter-rater reliability and alternate-form reliability coefficients for the CUBED-3 NLM Listening and NLM Reading subtests did not vary meaningfully across subgroups. No significant differences were observed in reliability estimates as a function of ethnicity, SES, or English language proficiency status, indicating that the scoring and administration procedures are robust to examiner bias and perform equivalently across diverse student groups.” It was unclear whether significant differences were observed as a function of IEP status.
Not MetValidity
- Multiple criterion measures were identified and described in response to question 6.1, and the vendor noted in response to question 6.3 that “To provide a stable summary estimate, the median coefficient from each study was computed by grade, and the mean of those medians is reported here. This convergence approach minimizes the influence of outliers and reflects the central tendency of observed relationships across multiple datasets.” For the first column of cells in the table for question 6.4, “Reference Standard” was entered for all grades and analyses. It was unclear which combinations of criterion measure(s) identified in question 6.1 were applicable for each analysis at each grade level, which prevented the reviewers from accurately scoring validity items.
- The response to question 6.2 referred reviewers to the response provided for question 5.3. The note in the row above regarding the representativeness of the sample across all performance levels for question 5.3 is also applicable for question 6.2.
- The response to question 6.6 referenced classification accuracy analyses disaggregated by student groups, rather than disaggregated validity analyses.
MetBias Analyses
Met screening requirements • 100% of Overall/Composite Score Bias Analyses criteria were met. The threshold for approval is 100%.
MetAdditional Considerations
- The links to the sample reports provided in the responses to questions 8.3 (for teachers) and 8.4 (for administrators) were identical; the title of the sample report indicated it was for administrators.
n/aSupp: Classification Accuracy
N/A Did not submit supplemental screening information
n/aSupp: Reliability
N/A Did not submit supplemental screening information
n/aSupp: Validity
N/A Did not submit supplemental screening information
n/aSupp: Bias Analyses
N/A Did not submit supplemental screening information
FastBridge (earlyReading, CBMreading, and aReading)
Not MetScreening Overall
Did not meet screening requirements
- 50% of screening criteria met. The threshold for approval is 100%. General assessment information:
- Administration format: Individual, Group (any size), Computer-administered
- Scoring: Manually (by hand) and automatically (computer-scored)
- Administration time:
- Grade K: 6-8 minutes
- Grade 1: 6-20 minutes
- Grade 2: 20-40 minutes
- Grade 3: 20-40 minutes
Not MetElements
- Phonemic awareness: Measured by earlyReading Onset Sounds (K), earlyReading Word Segmenting (K-1)
- Rapid automatized naming: Measured by earlyReading Letter Names (K)
- In response to question 2.2, the earlyReading Letter Names subtest was not described in sufficient detail to determine its alignment with the Rapid automatized naming element definition provided in the RFS Glossary: “A screening assessment component in which students orally name items (i.e., letters, numbers, colors, objects) as quickly as possible within a prescribed time limit.”
- Letter-sound correspondence: Measured by earlyReading Letter Names (K), earlyReading Letter Sounds (K)
- The Letter Names task was not described in sufficient detail to determine its alignment with the Letter-sound correspondence element definition provided in the RFS Glossary: “The knowledge of the relationship between letters and the sounds they represent.” However, the Letter Sounds task meets the definition.
- Single-word reading: Measured by earlyReading Sight Words (K-1)
- Nonsense-word reading: Measured by earlyReading Nonsense Words (K- 1)
- In describing the earlyReading Nonsense Words subtest in response to question 2.2, the vendor noted that "Decodable Words can be substituted for Nonsense Words when needed." The Nonsense-word reading element definition provided in the RFS Glossary states: “Assessment in which students read a list of standalone non-words with decodable patterns appropriate for the target age and grade range.” Real words cannot be substituted for nonsense words.
- Oral passage reading: Measured by CBMreading (1-3)
- Elements that may be included (optional):
- Retelling: Not measured
- Cloze reading procedure: Not measured
- Answering questions about a reading passage: Measured by aReading (2-3)
- In question 1.4, the vendor noted that “FastBridge is a single assessment that has multiple subtests by grade level to create a composite score. In grade K, earlyReading is the required subtest, in grade 1, earlyReading and CBMreading combine to create one composite score, and in Grades 2-3, aReading and CBMreading are administered as multiple measures within the same assessment.” It was unclear how this structure aligned with the reported data for the remainder of the submission, particularly when certain subtests were not specifically noted as being included in the analyses (e.g., the omission of aReading in the discussion of the reliability analyses).
MetClassification Accuracy
- External criterion measure for classification accuracy analyses: GRADE (K-1), NWEA MAP Growth (2-3)
- The response to question 4.2 did not specifically address the appropriateness of using the GRADE and NWEA MAP Growth assessments as criterion measures for the classification accuracy analyses.
- It was not completely clear if the classification accuracy analyses were conducted based on a single overall/composite score or for each subtest of the FastBridge suite separately. In response to question 4.3, the vendor noted that “Cut points were selected by optimizing sensitivity and then balancing sensitivity with specificity using methods presented in Silberglitt and Hintze (2005).”
- Entered "0" for all cells across student demographics categories for the Screening Classification Accuracy Sample table. The vendor noted that “…Renaissance’s Research team plans to target key demographic groups for a special study that would include the FastTrack assessments (i.e., CBMreading, earlyReading, and aReading). The plan includes identification of potential districts willing to participate, and identification of an external criterion measure, which includes the Star Assessment Suite. Demographic data including gender, race/ethnicity, English Language Learner status, free and reduced lunch status, and IEP status will likely be requested as part of the study to collect additional data on classification accuracy. We are committed to getting the requested data and are in the process of performing additional analyses to address these requirements."
Not MetReliability
- It was unclear if the reliability analyses were conducted based on a single overall/composite score or for each subtest of the FastBridge suite separately. In question 5.2, the vendor noted reliability analyses conducted for CBMReading and earlyReading, but aReading was not addressed.
Not MetValidity
- It was unclear if the validity analyses were conducted based on a single overall/composite score or for each subtest of the FastBridge suite separately. In question 6.3, the vendor noted validity analyses conducted for CBMReading and earlyReading (and RAN), but aReading was not addressed.
- In the table for question 6.4, Aimsweb was included as the criterion measure for Concurrent validity at grades 2-3, but the measure was not described or justified in the responses to questions 6.1 and 6.3.
MetBias Analyses
- In question 7.2, the vendor noted that “Differential validity and differential classification accuracy methods by gender and race/ethnicity and grade and season were used. Validity and classification accuracy were based on the concurrent relationship between CBMreading and an adaptive measure of reading standards, aReading.” Bias analyses for the earlyReading subtest were not addressed, and the vendor did not report any bias analyses conducted on the basis of student socioeconomic status, English language proficiency, or IEP status.
MetAdditional Considerations
- In question 8.8, the vendor did not provide a description of the compatibility between FastBridge and each of the following:
- i-Ready
- Michigan’s Early Literacy and Mathematics Benchmark Assessments
n/aSupp: Classification Accuracy
- The vendor entered “See the response above” for all items in the Supplemental Screening sections. Because all requirements for the Primary Screening sections were not met, the Supplemental Screening sections were not reviewed in full.
n/aSupp: Reliability
N/A Did not meet overall/composite screening requirements
n/aSupp: Validity
N/A Did not meet overall/composite screening requirements
n/aSupp: Bias Analyses
N/A Did not meet overall/composite screening requirements
MetScreening Overall
Met screening requirements
- 100% of screening criteria met General assessment information:
- Administration format: Computer-administered, with the option to hand-score recordings for relevant subtests
- Scoring: All student responses are machine-scored automatically
- Administration time: 20-30 minutes for all grades K-3
MetElements
- Phonemic awareness: Measured by a (K-3) Phonological Awareness Progression that includes: "Rhyme Completion, Counting Syllables, Onset- Rime Blending, Initial Sound Matching, Blending Phonemes, Phoneme Counting, Phoneme Addition/Deletion, and Phoneme Substitution. Although the element definition was met, students are not required to produce speech sounds for phonemic awareness tasks. Students listen, then choose a response on the screen.
- Rapid automatized naming: Measured by Rapid Automatized Naming (K- 3)
- Letter-sound correspondence: Measured by a (K-3) Phonics and Word Recognition Progression that includes: Letter-Sound Fluency, Build Words- One Letter, Word Families-Initial Letter, Decoding-CVC, Build Words-CVC, Decoding-Single Syllable, Build Words-Single Syllable.
- Single-word reading: Measured by the same (K-3) Phonics and Word Recognition Progression as the Letter-sound correspondence element. Although the element definition was met, students are not required to read words aloud. They hear a word and then choose the matching picture or word on the screen.
- Nonsense-word reading: Measured by Nonsense Words Fluency (K-3)
- Oral passage reading: Measured by Oral Reading Passages (1-3)
- Elements that may be included (optional):
- Retelling: Not measured
- Cloze reading procedure: Not measured
- Answering questions about a reading passage: Measured by Literal Comprehension questions after Oral Reading Passages (1-3)
MetClassification Accuracy
- External criterion measure for classification accuracy analyses: NWEA MAP Growth
- One or more students in the classification accuracy sample had unknown gender or unknown ethnicity. However, the overall grade level samples were very large, and the RFS did not prompt follow-up explanations for students in the unknown groups.
MetReliability
Met screening requirements
- 100% of Overall/Composite Score Reliability criteria were met. The threshold for approval is 92%.
MetValidity
- External criterion measure(s) for validity analyses: Concurrent MAP Growth Reading RIT and Spring MAP Growth Reading RIT
MetBias Analyses
- Bias analyses have been conducted for Gender and Ethnicity, revealing no evidence of bias. Additional bias analyses for Socioeconomic Status, English Language Proficiency, and IEP Status are planned, with anticipated results available by January 2028.
MetAdditional Considerations
Met screening requirements
- 100% of Additional Considerations criteria were met. The threshold for approval is 90%.
n/aSupp: Classification Accuracy
N/A Did not submit supplemental screening information
n/aSupp: Reliability
N/A Did not submit supplemental screening information
n/aSupp: Validity
N/A Did not submit supplemental screening information
n/aSupp: Bias Analyses
N/A Did not submit supplemental screening information
Progress Monitoring Measures
MetProgress Monitoring Measure 1: Phonological Awareness
- Number of alternate measures: “The Phonological Awareness progress monitoring assessment does not rely on a fixed set of alternate forms. Item pools are sufficient to support monthly progress monitoring in addition to seasonal screening.” There are 431 items in the Phonological Awareness item pool. The vendor did not specify how many items are available across a continuum of difficulty levels or by grade level.
- Reliability: Expectations were met for reliability for Grades K-3.
- “Disaggregated reliability results were not yet available for Phonological Awareness progress monitoring tests. NWEA anticipates that adequate sample sizes and test data for these subgroup analyses could be available within approximately one school year of data collection under Michigan’s implementation of early literacy screening requirements, with results expected as early as the end of 2027.”
- Reliability of Slope: Expectations were met for reliability for Grades K-3. Complete research is not required for approval.
- “We have not conducted analyses for sensitivity/reliability of slopes for the three progress monitoring tests because administration intervals were not sufficiently standardized across classrooms, creating heterogeneous time gaps that confound rate estimates. NWEA anticipates that the data needed for these analyses could be available within approximately one school year of continued data collection under Michigan’s implementation of early literacy screening requirements, with results expected as early as the end of 2027.”
- External criterion measure(s) for Validity Analyses: Concurrent: MAP Growth Reading RIT; Predictive: Spring MAP Growth Reading RIT
- Validity: Expectations were met for validity for Grades K-3.
- The vendor did not describe the representativeness of students across all performance levels for each validity analysis conducted.
- “Disaggregated validity results were not yet available for Phonological Awareness progress monitoring tests. Based on current operational administration volumes and continued efforts to obtain student demographic information, NWEA anticipates that adequate sample sizes and test data for these subgroup validity analyses could be available within approximately one school year of data collection under Michigan’s implementation of early literacy screening requirements, with results expected as early as the end of 2027.”
MetProgress Monitoring Measure 2: Phonics and Word Recognition
- Number of alternate measures: “The Phonics and Word Recognition progress monitoring assessment does not rely on a fixed set of alternate forms. Item pools are sufficient to support monthly progress monitoring in addition to seasonal screening.” There are 612 items in the Phonics and Word Recognition item pool. The vendor did not specify how many items are available across a continuum of difficulty levels or by grade level.
- Reliability: Expectations were met for reliability for Grades K-3.
- “Disaggregated reliability results were not yet available for Phonics and Word Recognition progress monitoring tests. NWEA anticipates that adequate sample sizes and test data for these subgroup analyses could be available within approximately one school year of data collection under Michigan’s implementation of early literacy screening requirements, with results expected as early as the end of 2027.”
- Reliability of Slope: Expectations were met for reliability for Grades K-3. Complete research is not required for approval.
- “We have not conducted analyses for sensitivity/reliability of slopes for the three progress monitoring tests because administration intervals were not sufficiently standardized across classrooms, creating heterogeneous time gaps that confound rate estimates. NWEA anticipates that the data needed for these analyses could be available within approximately one school year of continued data collection under Michigan’s implementation of early literacy screening requirements, with results expected as early as the end of 2027.”
- External criterion measure(s) for Validity Analyses: Concurrent: MAP Growth Reading RIT; Predictive: Spring MAP Growth Reading RIT
- Validity:
- The vendor did not describe the representativeness of students across all performance levels for each validity analysis conducted.
- “Disaggregated validity results were not yet available for Phonics and Word Recognition progress monitoring tests. Based on current operational administration volumes and continued efforts to obtain student demographic information, NWEA anticipates that adequate sample sizes and test data for these subgroup validity analyses could be available within approximately one school year of data collection under Michigan’s implementation of early literacy screening requirements, with results expected as early as the end of 2027.”
MetProgress Monitoring Measure 3: Oral Reading Fluency
- “After each passage students are asked six comprehension questions, which are selected-response items presented with audio and scored dichotomously. In addition to the oral reading metric SWCPM, number correct out of six comprehension questions are reported for each passage. These values are not synthesized. Only scaled words correct per minute are plotted in the progress graph.”
- Number of alternate measures: “The Oral Reading progress monitoring is structured around pools of passages stratified by Lexile bands. Within each Lexile band, multiple passages are available, ranging from 21 to 44 passages per band. During administration, students are assigned to the Lexile band aligned with their grade and current Lexile level, and one passage is randomly selected from within the designated Lexile band. Scaled words correct per minute is the reported outcome, which normalizes reading rate (WCPM)with respect to empirical passage difficulty (scaled WCPM).”
- Reliability: Expectations were met for reliability for Grades 1-3.
- “Disaggregated reliability results were not yet available for Oral Reading Fluency progress monitoring tests. NWEA anticipates that adequate sample sizes and test data for these subgroup analyses could be available within approximately one school year of data collection under Michigan’s implementation of early literacy screening requirements, with results expected as early as the end of 2027.”
- Reliability of Slope: Expectations were met for reliability for Grades 1-3. Complete research is not required for approval.
- “We have not conducted analyses for sensitivity/reliability of slopes for the three progress monitoring tests because administration intervals were not sufficiently standardized across classrooms, creating heterogeneous time gaps that confound rate estimates. NWEA anticipates that the data needed for these analyses could be available within approximately one school year of continued data collection under Michigan’s implementation of early literacy screening requirements, with results expected as early as the end of 2027.”
- External criterion measure(s) for Validity Analyses: Concurrent: DIBELS Next/Acadience Oral Reading Fluency scores; Predictive: Spring MAP Growth Reading RIT
- Validity:
- The vendor did not describe the representativeness of students across all performance levels for each validity analysis conducted.
- “Disaggregated validity results were not yet available for Oral Reading Fluency progress monitoring tests. Based on current operational administration volumes and continued efforts to obtain student demographic information, NWEA anticipates that adequate sample sizes and test data for these subgroup validity analyses could be available within approximately one school year of data collection under Michigan’s implementation of early literacy screening requirements, with results expected as early as the end of 2027.”
MAP Reading Fluency (primary screener) and MAP Growth (secondary screener)
Not MetScreening Overall
Did not meet screening requirements
- 100% of MAP Reading Fluency (primary screener) screening criteria were met. The threshold for approval is 100%.
- MAP Growth (supplemental screening measure) screening sections were not rated due to insufficient additional element coverage beyond MAP Reading Fluency.
- See the Consensus Scoring Document for MAP Reading Fluency, which did meet screening requirements. General assessment information:
- Administration format: Computer-administered, with the option to hand-score recordings for relevant subtests
- Scoring: All student responses are machine-scored automatically
- Administration time: 65-75 minutes (20-30 minutes for MAP Reading Fluency, 45 minutes for MAP Growth) for all grades K-3
MetElements
- Phonemic awareness: Measured by a Phonological Awareness Progression that includes: "Rhyme Completion, Counting Syllables, Onset- Rime Blending, Initial Sound Matching, Blending Phonemes, Phoneme Counting, Phoneme Addition/Deletion, and Phoneme Substitution (K-3). Although the element definition was met, students are not required to produce speech sounds for phonemic awareness tasks. Students listen, then choose a response on the screen.
- Rapid automatized naming: Measured by Rapid Automatized Naming (K- 3)
- Letter-sound correspondence: Measured by a Phonics and Word Recognition Progression that includes: Letter-Sound Fluency, Build Words- One Letter, Word Families-Initial Letter, Decoding-CVC, Build Words-CVC, Decoding-Single Syllable, Build Words-Single Syllable. (K-3)
- Single-word reading: Measured by the same Phonics and Word Recognition Progression as the Letter-sound correspondence element (K-3). Although the element definition was met, students are not required to read words aloud. They hear a word and then choose the matching picture or word on the screen.
- Nonsense-word reading: Measured by Nonsense Words Fluency
- Oral passage reading: Measured by Oral Reading Passages (1-3)
- Elements that may be included (optional):
- Retelling: Not measured
- Cloze reading procedure: Not measured
- Answering questions about a reading passage: Measured by Literal Comprehension questions after Oral Reading Passages (1-2)
MetClassification Accuracy
- External criterion measure for classification accuracy analyses: NWEA MAP Growth
- One or more students in the classification accuracy sample had unknown gender or unknown ethnicity. However, the overall grade level samples were very large, and the RFS did not prompt follow-up explanations for students in the unknown groups.
MetReliability
Met screening requirements
- 100% of Overall/Composite Score Reliability criteria were met. The threshold for approval is 92%.
MetValidity
Met screening requirements
- 100% of Overall/Composite Score Validity criteria were met. The threshold for approval is 90%.
MetBias Analyses
- Bias analyses have been conducted for Gender and Ethnicity, revealing no evidence of bias. Additional bias analyses for Socioeconomic Status, English Language Proficiency, and IEP Status are planned, with anticipated results available by January 2028.
MetAdditional Considerations
Met screening requirements
- 100% of Additional Considerations criteria were met. The threshold for approval is 90%.
Not MetSupp: Classification Accuracy
- The Request for Submissions Question 2.4 option to Continue and provide additional data states: “The overall/composite score for the screening assessment does not cover all of the required elements for all grade levels. For the remaining sections of the RFS related to screening, I will enter research results and data specific to the assessment's overall/composite score. At the end of the RFS, I will enter one additional set of research results and data related to a required element not covered in the overall/composite score to ensure complete coverage of all required elements.” However, all required elements are addressed through MAP Reading Fluency alone, which was submitted separately for review. MAP Growth adds another measure for the element of Answering Questions about a Reading Passage, which is also covered by MAP Reading Fluency. Given the lack of additional coverage of required and optional elements, administering all of MAP Growth would add significant testing time (45 minutes), but not provide additional insight into students’ reading skills.
n/aSupp: Reliability
- MAP Growth screening reliability was not rated due to insufficient additional element coverage beyond MAP Reading Fluency.
n/aSupp: Validity
- MAP Growth screening validity was not rated due to insufficient coverage of additional elements beyond MAP Reading Fluency.
n/aSupp: Bias Analyses
- MAP Growth screening bias analyses were not rated due to insufficient coverage of additional elements beyond MAP Reading Fluency.
Star Assessments (Star Early Literacy, Star Reading, Star CBM Reading)
Not MetScreening Overall
Did not meet screening requirements
- 67% of screening criteria met. The threshold for approval is 100%. General assessment information:
- Administration format: Individual, Group (any size), Computer-administered
- Scoring: Automatically (computer-scored)
- Administration time:
- Grade K: 1-15 minutes
- Grade 1: 1-20 minutes
- Grade 2: 1-20 minutes
- Grade 3: 1-20 minutes
MetElements
- Phonemic awareness: Measured by Star CBM Phoneme Segmentation (K- 1)
- Rapid automatized naming: Measured by Star CBM Letter Naming (K)
- Letter-sound correspondence: Measured by Star CBM Letter Naming and Letter Sounds (K)
- The Letter Naming task did not align with the Letter-sound correspondence element definition provided in the RFS Glossary: “The knowledge of the relationship between letters and the sounds they represent.” However, the Letter Sounds task meets the definition.
- Single-word reading: Measured by Star CBM Sight and High-Frequency Words (1)
- In the response to question 2.3, the vendor noted that the Star CBM Sight and High-Frequency Words subtest is not included in the overall/composite score. Data for this subtest were included in the Supplemental Screening sections of the submission (questions 11.1- 14.4), but it was not reviewed in full at this time; refer to the note for Supplemental Screening Score: Classification Accuracy below.
- Nonsense-word reading: Measured by Star CBM Receptive Nonsense Words (K), Star CBM Expressive Nonsense Words (1)
- Oral passage reading: Measured by Star CBM Passage Oral Reading (1-3)
- Elements that may be included (optional):
- Retelling: Not measured
- Cloze reading procedure: Measured by Star Early Literacy/Star Reading (2-3) The description provided for the Star Early Literacy/Star Reading subtest in addressing the Cloze reading procedure element was vague. In question 2.2, the vendor noted that “The vocabulary and comprehensive measures include cloze product items where words are deleted from sentences and reading passages. Students are expected to select the word that best fits the sentence or reading passage.”
- Answering questions about a reading passage: Measured by Star Early Literacy/Star Reading (1-3)
- In question 1.3, the vendor noted that “For the Michigan submission, we are offering Star Early Literacy and Star Reading for the optional skills: cloze reading procedure and answering questions about passages.” The required elements (except for Single-word reading, noted above) appeared to be met by the Star CBM Reading Overall Risk score. It was unclear if Star CBM Reading and Star Early Literacy/Star Reading were able to be combined into a single overall/composite score.
Not MetClassification Accuracy
- External criterion measure for classification accuracy analyses:
- Star Early Literacy: FastBridge, Wonders Reading Assessment
- Star Reading: DIBELS ORF, FL Assessment of Student Thinking, Smarter Balanced
- Star CBM Reading: Star Early Literacy, Star Reading, DIBELS ORF, NWEA Measures of Academic Process In response to question 4.2, the vendor noted that “Renaissance has looked at correlations and classification accuracy with other assessments, its own measures, and various state tests as strategies for providing validity evidence in support of its measures. Star assessments do not share items across products, and separate analyses have been performed with different samples to look at reliability and validity. All Star tools were developed independently and the products are not designed to be over-aligned with each other.”
- For the results of the classification accuracy analyses provided for questions 4.4-4.6, Star CBM Reading did not meet acceptable statistical requirements for grade 1:
- AUC Confidence Interval Lower and Upper Bounds were not provided.
- The provided value for specificity was less than 0.70.
- The data provided for questions 4.4-4.6 contained some missing values for the AUC Confidence Interval Lower and Upper Bounds:
- Star CBM: Grade 1 (noted above)
- Star Early Literacy: Grade 1 (not addressed) and Grade 2 (entered as “n/a”)
- Two samples by grade level were noted as having sizes below 100:
- Star Reading: Grade 1
- Star CBM (Passage Oral Reading): Grade 2
- Regarding additional analyses for classification accuracy, the vendor noted that “Our estimated timeline for collecting demographic data will start at the beginning of the 2026-2027 school year. We expect to make good progress with recruitment and data collection over the school year and will provide an update in summer 2027.”
MetReliability
- Regarding additional disaggregated reliability analyses, the vendor noted that “Our estimated timeline for collecting demographic data will start at the beginning of the 2026-2027 school year. We expect to make good progress with recruitment and data collection over the school year and will provide an update in summer 2027.”
Not MetValidity
- Although the quantitative results of the validity analyses provided in the table for question 6.4 met the minimum threshold for approval, several identified criterion measures (ACT Aspire, CAT-5, and Smarter Balanced) were not described or justified in questions 6.1, 6.3, or 6.5. It was therefore difficult to determine the applicability of the data provided in question 6.4 to the submitted assessment.
- For question 6.4, Predictive validity data were not provided at grade K. However, the requirement of one form of validity (Concurrent) with acceptable results was met.
- Regarding additional disaggregated validity analyses, the vendor noted that “Our estimated timeline for collecting demographic data will start at the beginning of the 2026-2027 school year. We expect to make good progress with recruitment and data collection over the school year and will provide an update in summer 2027.”
MetBias Analyses
- The vendor noted that the DIF analyses described in questions 7.2-7.3 applied to Star Early Literacy and Star Reading.
- Regarding additional bias analyses, the vendor noted that “Our estimated timeline for collecting demographic data will start at the beginning of the 2026-2027 school year. We expect to make good progress with recruitment and data collection over the school year and will provide an update in summer 2027.”
MetAdditional Considerations
- In question 8.8, the vendor did not provide a description of the compatibility between Star Assessments and each of the following:
- i-Ready
- Michigan’s Early Literacy and Mathematics Benchmark Assessments
- The vendor noted that “Renaissance would be happy to pursue a validation study, with Curriculum Associates’ permission. As for the Michigan Department of Education’s Early Literacy Benchmark Assessments, Renaissance would be happy to link to these assessments. We will work with MDE on this project.”
n/aSupp: Classification Accuracy
- The vendor provided responses to the four Supplemental Screening sections regarding the task that addressed the required Single- word reading element: Star CBM Reading Sight and High- Frequency Words. Because all requirements for the Primary Screening sections were not met, the Supplemental Screening sections were not reviewed in full.
n/aSupp: Reliability
N/A Did not meet overall/composite screening requirements
n/aSupp: Validity
N/A Did not meet overall/composite screening requirements
n/aSupp: Bias Analyses
N/A Did not meet overall/composite screening requirements
Where to go next: how a screener can be “accurate” and still miss your kid, the test cheat sheet library, or five questions to ask your school about its screener.