Inputs, Outputs, & Outcomes Deep Dive

Inputs, Outputs, and Outcomes

Inputs are resources and activities invested in a particular program or strategy that are usually knowable at the beginning of a cycle and that are a measure of effort applied. In a school system, inputs are things like staff, books, programs, facilities, buses, and everything else that we invest resources in prior to the start of school.

Outputs are the result of a particular set of inputs that is usually knowable in the midst of a cycle and that is a measure of the implementation of the program or strategy. In a school system, outputs are things like formative assessment scores, interim assessment scores, benchmark results, grades, quarterly financials, and everything else that measures our implementation of our inputs throughout the course of the school year.

Outcomes are the impact of the program or strategy that is usually knowable at the end of a cycle and that is a measure of the effect on the intended beneficiary. In a school system, outcomes are things like graduation rates, summative assessment scores, end of year audit results, end of year retention rates and everything else that measures the final result of all the inputs and outputs at the end of the school year.

All inputs, outputs, and outcomes in a school system come in one of two varieties: adult or student. As an example, adult outcomes are results at the end of the school year that are measures of what adults know or are able to do. Student outcomes are results at the end of the school year that are measures of what students know and are able to do. It is critical to remember that school systems only exist for one reason: to improve student outcomes. Not for adult outcomes, not for adult inputs, not for student outputs. School systems only exist to improve student outcomes. All other activities should be in service of this sole reason for existing.

Example

An example of a desired student outcome is having more students scoring 3 or better on their AP exams. That is a measure of what students know or are able to do by the end of the cycle; it’s a student outcome. But outcomes don’t simply occur. They are the result of inputs by staff that lead to outputs that then can lead to outcomes.

An input that might help lead to the aforementioned outcome is making more seats available in AP courses. This is an input in that it’s the investing of resources at the beginning of the school year. Outputs that help lead to the desired outcome might be the increase in the number/percentage of students enrolled in or on track in the AP courses. These are outputs in that they are measures of the fidelity of implementation of the input and they are knowable throughout the school year.

Details

Inputs vs Outputs vs Outcomes: A Conceptual Framework

Within the cycle of school systems, there is a value chain that begins with adult deployment of resources, continues through implementation day in and day out, and ends with student measurable results. This is analogous to how a journey begins with planning for the journey and allocating resources, proceeding on the journey, and culminates with arrival at a destination. For the purpose of coding how school boards use their time during public school board meetings, I refer to these three distinct phases in the school system cycle of education creation as inputs, outputs, and outcomes.

Inputs describe measures of resources or strategies that are observable at the beginning of a cycle. In a school system, examples include books, buses, buildings, staff, plans, and any other resource or strategy measure that is generally known before the first day of school begins.

Outputs describe measures of implementation that are observable in the midst of a cycle. In a school system, examples include attendance, budget updates, usage milestones, quizzes, mid-year surveys, interim assessments, and any other implementation measure that is generally knowable throughout the school year.

Outcomes describe measures of results that are observable at the end of a cycle. In a school system, examples include end-of-year average daily attendance, whether or not the financials balanced, teacher retention rates, summative assessments, and any other measures of school system results generally knowable at the end of the school year.

Adult Outcomes vs Student Outcomes

Of equal importance to being able to distinguish between inputs, outputs, and outcomes, is noticing whether the active party is a student or an adult. It’s very different to be discussing adult outcomes — measures of what adults know or are able to do at the end of the cycle — than to be discussing student outcomes — measures of what students know or are able to do at the end of the cycle.

Adult inputs are resources and strategies observable at the beginning of a cycle that adults deploy. Examples of adult inputs include budget allocations, staffing assignments, adopted curricula/programs, professional-learning calendars, bus routes/fleet, facilities readiness, and other resources/strategies put in place before school starts.

Student inputs are resources and strategies observable at the beginning of a cycle that students deploy. Examples of student inputs include students’ own course selections made before the term, student selection of learning pathways, opting into tutoring/enrichment cohorts at the outset, student-created plans for their own learning, student-led restorative practices plans, and setting a personal study routine/plan at the start of the year.

Adult outputs are measures of what adults have done or implemented in the midst of the cycle. Examples of adult outputs include quarterly financials and tax-collection reports, administrative/board corrective-action plan updates, implementation milestones, attendance to staff trainings, and interim budget execution reports.

Student outputs are measures of what students have done or implemented in the midst of the cycle. Examples of student outputs include classroom attendance, behavior incident rates, assignment/grade checks, benchmark and interim assessment results, on-track/credit-accumulation updates, and mid-year student surveys.

Adult outcomes are measures of what adults know or are able to do at the end of the cycle. Examples of adult outcomes include end-of-year audit results, whether the budget balanced, teacher/leader retention rates, certification/compliance completion, and end-of-year safety/compliance findings.

Student outcomes are measures of what students know or are able to do at the end of the cycle. Examples of student outcomes include graduation rates, summative assessment results, AP/IB successful pass rates, college persistence, and other end-of-year measures of what students know or can do.

Early vs Mid vs Late

Another series of distinctions concerning inputs, outputs, and outcomes relates to when they occur chronologically at the system level. For example some inputs arrive earlier in the process – like planning activities – than other inputs. Some inputs that arrive in the middle of other inputs – like the acquisition of resources. Some inputs arrive after most other inputs - like the initiation of activities that have been planned. This same phenomenon holds for outputs and outcomes.

Early inputs are plans for resources and strategies that occur prior to other inputs. Examples of early inputs include annual academic plan drafts, staffing plans, and budgets (planning).

Mid inputs are the acquisition of planned resources that occur in the midst of early and late inputs. Examples of mid inputs include issuing RFPs and purchase orders, hiring to fill the staffing plan, selection of resources, and procurement of curriculum/devices/fleet (resources).

Late inputs are the initiation of activities that using the planned resources and strategies and that occur after other inputs. Examples of late inputs include launching tutoring, opening newly staffed courses, and beginning professional development cycles as the year starts (activities).

Early outputs are measures of participation in implementation prior to other outputs. Examples of early outputs include student participation in classes, staff participation in professional development, staff participation in planned strategies (participation).

Mid outputs are measures of implementation fidelity in the midst of early and late outputs. Examples of mid outputs include walkthrough/rubric ratings, completion of required implementation steps, and other fidelity of implementation indicators reported during the year (fidelity).

Late outputs are in-process result measures of implementation that occur after other outputs have taken place. Examples of late outputs include interim assessment performance, end-of-term on-track indicators, and quarter/semester grade distributions (midterms).

Early outcomes are measures of results that are knowable immediately at the conclusion of the cycle before other outcomes are known. Examples of early outcomes include end of course performance, end of year attendance rates, balanced budgets, student/staff/family satisfaction surveys, and internal summative assessment results (results).

Mid outcomes are measures of the impact of results that are knowable in between early outcomes and late outcomes. Examples of mid outcomes include external summative assessment results (state or federal), annual audit findings, and system accreditation ratings (impact).

Late outcomes are measures of the impact on the long term vision knowable after all other outcomes are known. Examples of late outcomes include post-secondary success, successful health outcomes throughout early adulthood, successful career outcomes throughout early adulthood, and demonstrations of maximization (vision).

Diagnostic, Formative, Interim, & Summative Assessments

Closely related to the concepts of inputs, outputs, and outcomes are the concepts of diagnostic, formative, interim, and summative. Where inputs / outputs / outcomes can refer to any of the processes within an organization, diagnostic / formative / interim / summative focus specifically on methods of assessment and associated metrics. These terms are most commonly used in academic contexts.

Diagnostic assessments are those given before the beginning of an instructional cycle or at the very beginning of an instructional cycle that measure what knowledge and skills students have already accumulated. Generally administered in the opening weeks of each school year, diagnostic assessments provide schools and teachers with clarity regarding what topics need to be given more or less time. Diagnostic metrics data generally serves as an input used to organize and structure instructional decisions for the rest of the first semester.

Formative assessments are those given in the midst of an instructional cycle that measure student learning, narrowly tailored to specific student expectations. Generally created and/or administered on an ongoing basis by teachers or campuses for the purpose of refining instructional practice. Formative metrics are the data provided by formative assessments or other input evaluation tools and generally serve as data used to evaluate and refine the fidelity of implementation. Formative assessments and metrics are types of outputs but are likely not aligned to the summative assessment or metrics.

Interim assessments are those given in the midst of an instructional cycle that measure a cross-section of knowledge or skills. Generally administered up to two or three times per year by campuses or school systems for the purpose of predicting summative performance. Interim metrics are the data provided by interim assessments or other outcome predictive tools and generally serve as data used to predict summative results by tracking quantifiable progress at critical and ongoing junctures during a cycle. Interim assessments and metrics are types of outputs and are explicitly aligned to the summative assessment or metrics.

Summative assessments are those given at the end of an instructional cycle that measure a cross-section of knowledge or skills over the course of an instructional cycle or school year. Generally administered at the end of a curricular unit, school year, or the transition to a new schooling experience by school systems or states for the purpose of evaluating school or school system effectiveness. Summative metrics are the data provided by summative assessments or other measurements of the final results in a school system. Summative metrics are used to track outcomes and to evaluate whether the school system/chief executive was successful. Summative metrics are generally a type of outcome – and when measuring what students know or are able to do, are a type of student outcome.

Norm-referenced & Criterion-referenced

Another set of assessment-related concepts is the distinction between norm-referenced and criterion-referenced. These are two different ways of interpreting what the assigned rating will be after the raw scores are collected.

If assessment results are considered criterion-referenced, this means that answers given by the student are being compared to a set of fixed standards. A key element of criterion-referenced assessments is that every student whose raw score is above a certain threshold is assigned the same rating. For example, if a score of 9 right answers out of 10, 90%, is considered an ‘A’ rating, then everyone who scored 9/10 will earn an ‘A’. Criterion-referenced assessment results are more ideal for identifying whether or not the students have a mastery of a predefined set of knowledge and/or skills.

If assessment results are considered norm-referenced, this means that answers given by the student are being compared to the answers provided by the other students. A key element of norm-referenced assessments is that it is possible to have an improved raw score but not have an improved assigned rating. For example, if a student improved their score by 5 points but all other students improved their score by 10 points, the student’s assigned rating might actually go down. Norm-referenced assessment results are more ideal for identifying where students are in relation to each other.

Proficiency, Growth, & Comparison

Yet another set of distinctions critical to this work is the nature of expectation construction. Proficiency is about having expectations that are relative to a predefined competency level. Proficiency asks whether or not a student is at a particular level, usually expressed something like “on grade level” or “passing”. An example of a proficiency goal might be, “the percentage of 3rd grade students demonstrating on grade level proficiency in reading will increase from 40% in June 2021 to 70% by June 2026”.

Growth is about having expectations that are relative to a predefined starting point. Growth asks whether or not a student has demonstrated improvement from the starting point to an ending point, usually expressed something like “grew by more than a grade level” or “demonstrated adequate/expected growth”. An example of a growth goal might be, “the percentage of 3rd grade students who began the year below grade level in reading and who grew by at least 1.25 grade levels will increase from 40% in June 2021 to 70% by June 2026”.

Comparison is about having expectations that are relative to a control group. Comparison asks whether or not a student has demonstrated growth or proficiency relative to a different group of students, usually expressed something like “closing the gap”. An example of a comparison goal might be, “the percentage of 3rd grade students who were two grade levels behind in reading who closed the gap between themselves and the overall 3rd grade average will increase from 40% in June 2021 to 70% by June 2026”.

Static Goals & Cohort Goals

Another set of choices in goal setting is whether to have goals that track a specific group of students over time, cohort goals, or that track a fixed indicator over time, static goals. By default, most people unknowingly utilize static goals (in the same way that most people, absent thinking through the options, default to proficiency goals or default to assuming that assessment scoring is criterion-referenced).

The benefit of static goals is that they provide an insight into how a specific area of the organization is improving over time by measuring different students using the same assessment. The benefit of cohort goals is that they provide an insight into how a defined group of students is improving as it moves through the system by measuring the same students using different assessments.

Milestones, Interim Goals / Interim Guardrails, & Goals / Guardrails

Closely related to the concepts of formative, interim, and summative are the concepts of milestones, interim goals / interim guardrails, and goals / guardrails. Where formative / interim / summative focus specifically on methods of assessment and associated metrics, milestones, interim goals / interim guardrails, & goals / guardrails focus specifically on project management and progress monitoring. These terms are most commonly used in managerial or governance contexts.

Milestones are dates by when specific inputs are implemented – by when project deliverables are due. By their nature, milestones are descriptive of staff actions that have been taken and, as such, are how an organization’s managerial team track progress on implementation. Milestones are created by managerials teams and are not a governance instrument.

Interim goals and interim guardrails are measures of progress toward a defined goal or guardrail. By their nature, interim goals and interim guardrails are output metrics that are predictive of the goals / guardrails and that are influenceable by the chief executive. Interim goals and interim guardrails are defined by the chief executive, in consultation with the board, and are the data the chief executive uses to routinely report on performance to the board.

Goals and guardrails are policy statements created by the board, in consultation with the chief executive, that are intended to capture the vision and values, respectively, of the community that the board represents. Goals are SMART while guardrails are not.

Selecting Goal / Interim Goal Targets

When determining the targets for goals and interim goals, the management team should evaluate at least three things:

5yr Average Growth: Identify the previous five years summative percentage of year-over-year growth for the targeted student outcome. Then average those five numbers together and that’s the minimum average growth to beat. The analysis assumes a reasonably static trendline so 1) if circumstances have disrupted that, adjust accordingly and 2) this should be considered the bare minimum threshold of performance and any targets should outperform this.

Degree of Difficulty: Some tasks are easier than others, so this should always be taken into account. In general, I assume english is harder for students to learn than math. I assume lower grades are easier to catch up than higher grades. I assume a significant jump to get off the very floor of performance (ie: going from below 0% proficiency to 10% proficiency) is easier than the final jump to reach the very ceiling of performance (ie: going from 90% proficiency to 100% proficiency). I assume that helping students remain on grade level is easier than catching students up who are significantly below grade level. The more difficult the circumstances, the smaller the target I’d set; the easier the circumstances, the larger the target I’d set.

Resource Commitment: Whatever the average growth over the previous five years is for a given student outcome, I assume that as the growth baseline. And I assume the amount invested in strategies explicitly focused on improving that student outcome as the investment baseline. The general rule I use is that the multiple of the desired improvement above the baseline growth will determine the required new investment – the multiple times investment baseline. So if we’ve been growing at 1% per year and spending $100, then if I want to grow at 4% per year I need to be prepared to spend $400. If there isn’t a willingness or ability to invest at least that amount to improve performance, then the desired target is too high.

Questions To Ask When Evaluating Interim Metrics

Knowing whether or not interim metrics are sufficiently meaningful is always challenging. Here is a set of questions that staff, board members, and coaches should review prior to their being accepted:

Interim Metric Requirements

Are all of the interim metrics – for both goals and guardrails – SMART? This means they each have:

A starting point and an ending point

A starting date and an ending date

A specific measurement

In the case of interim goals, a specific student population

Are the interim metrics leading indicators relative to the goal/guardrail rather than lagging?

Are there exactly three interim metrics for each goal and for each guardrail?

If all three of the interim metrics are accomplished, will that likely ensure accomplishment of the goal?

Are all of the interim metrics influenceable by the superintendent? Does the superintendent have at least 80% authority over the implementation of the items being measured?

Are each of the interim metrics updateable multiple times per year?

Are each of the interim metrics outputs (mid-cycle results, knowable in the midst of the cycle) rather than inputs (resources/strategies, knowable at the beginning of the cycle)?

Interim Metric Considerations

What is the degree of correlation between the interim metric and the goal metric?

Can the data be monitored on the school board’s monitoring calendar within 30 to 60 days of when it is collected?

Do any of the interim metrics utilize data sourced from external to the school system (like CGCS’ academic KPIs or operational KPIs)?

Do any of the interim metrics rely on data that is historically unreliable or highly variable between individual schools?

Do the interim metrics fully address the content of the goal/guardrail?

Are each of the interim metrics strong outputs (measures of implementation quality) rather than weak outputs (measures of participation)?

Are there significant unintended consequences that need to be considered?

Does each interim metric have only one data set rather than multiple?

Has the management team evaluated each interim metric for unintended consequences?

Is the metric predictive both downward and upward?

Is the interim metric the data that management actually uses for ongoing decision making?

Is there curriculum that explicitly teaches the material covered by the interim goals and does evidence exist that the curriculum is being implemented in schools?

Do all of the interim metrics appear as part of the cabinet members’ annual evaluations? (see Cascading)

Regarding both upward and downward predictability, there are many measures (particularly adult inputs/outputs) that will appear to be predictive because as they decline, the likelihood of the goal being accomplished will also decline. For example, if you starve your 3rd graders, their caloric intake will decline and their performance will certainly decline as well. This gives you two potential interim metrics: calories provided and calories consumed. These metrics’ decline would appear to predict the goal’s decline. But if you start feeding students double the calories needed, that doesn’t mean that you’d expect to see a doubling in performance. Just because something appears predictive downward doesn’t mean it will also be predictive upward. This is one of the main reasons that interim goals have to be student outputs rather than adult inputs, adult outputs, or even student inputs – almost all of these will be necessary but insufficient to improve student outcomes, so even though they’d be predictive downward, they’re almost never predictive upward.

Regarding the last last two questions, monitoring done well simply reveals managerial decision making. If monitoring reports contain no information that management is actually using on a routine basis but instead is just data to show the school board for the sake of this process, monitoring of that data will be worthless. If management is having to invent each monitoring report from whole cloth each month for the next 60 months, that is a sign of a profound misunderstanding about the continuous improvement process at best, gross incompetence in managerial leadership at worst.

Is Assessment Harmful To Students?

Categorically, emphatically, and unapologetically, no. Assessment is like any other tool — a screwdriver, a level, some pliers — in that it is both critically essential to effective construction (in this case constructing knowledge rather than buildings) and it is capable of being abused. Assessments used abusively certainly have the capacity to be harmful. But anyone who suggests that all assessment is inherently harmful is profoundly misinformed and harbors beliefs that will be harmful to children. Assessment plays several key roles in the learning process but two of the most vital are long term memory creation and instructional practice improvement.

Assessment: Long Term Memory Creation

Part of the neurological science of learning involves students migrating information out of short term working memory and into long term memory. One strategy for supporting this migration is assessment. When students are taught something on Monday of week 1 and then quizzed on it on Friday of week 1 (this is usually formative assessment), the process of migration begins. Then when the same material is assessed again at the end of the semester period (this is often interim assessment), the process continues. When students are assessed on the same material again at the end of the year, requiring them to recall information learned months earlier (this is usually summative assessment), the associations with that learning are strengthened and reinforced in the brain’s physical structures. In this way, appropriately designed and spaced assessment directly contributes to the learner’s ability to master content.

Assessment: Instructional Practice Improvement

If there is any one place where the magic of education is most likely to occur, it’s in the interactions between the learner and the educator. Study after study suggests that of all the factors school systems control, quality of instruction has the largest impact on student performance. As instructional quality continuously improves, so too are student outcomes more likely to improve. And one of the most consistent paths for teachers to improve the quality of their instruction is to teach, assess what was taught, then reteach, then reassess. This continuous improvement cycle pushes teachers to constantly evaluate what worked and what did not and make adjustments based on it. But this process is impossible to conduct without assessment data – usually formative assessment data. Improving instructional practice is a vital key to improving student outcomes, but where there is no assessment, there will almost certainly be no improvement in instructional practice.

Is Standardized Assessment Harmful To Students?

Categorically, emphatically, and unapologetically, no. Standardized assessment is like any other tool — a screwdriver, a level, some pliers — in that it is both critically essential to effective construction (in this case constructing knowledge rather than buildings) and it is capable of being abused. Standardized assessments used abusively certainly have the capacity to be harmful. But anyone who suggests that all standardized assessment is inherently harmful is profoundly misinformed and harbors beliefs that will be harmful to children. Standardized assessment plays several key roles in the learning process but two of the most vital are performance transparency and system practice improvement.

Standardized Assessment: Performance Transparency

A commonly asked question is whether or not parents have the right to know how their child’s school is performing compared to other schools. If the answer is “no”, parents don’t have the right to understand the relative performance of their child’s school, then standardized assessment is less necessary. But if this is a right parents should have, then standardized assessment is essential. If the math teacher at one school gives students A’s for the same work that the math teacher at another school gives her students B’s, letter grades are no longer fair indicators of relative performance. One way of addressing unfairness in the system like this is to administer a common assessment at both schools, apply a common scoring rubric across both schools, and then routinely train teachers on the use of these assessments. In other words, to make it fair, you’d have to make the assessment standardized.

This transparency isn’t just needed in K-12 performance. Many professions – medicine, technology, engineering, construction, truck driving, and more – rely on standardized assessment as part of their systems for determining who can enter the field as a means of promoting fairness and ensuring performance. This is used to protect all of us — ensure doctors in multiple hospitals all give the right medicine, engineers design bridges across the state that can all carry the right load, system administrators design servers across an organization that all protect our data, etc. When the transparency of performance across multiple sites matters, well-formed standardized assessment is an appropriate response.

Standardized Assessment: System Practice Improvement

Transparency isn’t the only good reason to have standardized data. It’s also immensely useful for helping complex systems identify outliers and make needed corrections. But if the data isn’t apples to apples, it’s harder to know what is noise and what is signal.

It’s worth noting that standardized assessment can be formative, interim, or summative in nature and whichever of these is the case has a significant impact on how the data is appropriately used. It doesn’t make sense to use formative data to predict performance on a summative assessment; it doesn’t make sense to use summative data to try to inform weekly changes. Generally speaking, system-level practice improvement will rely on aggregated interim and summative data more so than formative data.

Is Standardized Assessment Biased?

Yes, because all assessments are biased – standardized or otherwise. Every assessment was created by someone who deployed aspects of their world view, their language, their understanding of the material into the assessment. Asking if standardized assessments are biased is the wrong question. What to ask instead: 1) is the bias that is present harmful to the collection of accurate data, and 2) what have we done to identify and address harmful bias? As an example, here’s a question from a 3rd grade math assessment:

If 300 crayons are added to 100 crayons, how many crayons are there?

Applying the wrong question actually creates more potential for harm. When asked, “is this question biased?” most people suggest it is not. That is asking the wrong question and that is the wrong answer to the wrong question. This question is clearly biased. The better question to ask is, “for whom is this question biased and is the bias that is present harmful to the collection of accurate data?” If children are being assessed who have no familiarity with crayons – as can be the case with asylees or children experiencing similar traumas – this is certainly a biased question. Does that bias harm the collection of accurate data? It depends on context. For the general population, probably not (the question doesn’t require the test taker to know what a crayon is, only that it’s an item being counted – but only “probably not” because not knowing what a crayon is may slow a child down as they try to figure that out, which could have an impact on their overall performance). But if the assessment is only going to be administered to asylees, then yes, the question has a high enough likelihood of harm that I’d consider omitting it. Context matters.

Of equal importance is taking time to ensure that assessments have been vetted for bias. Each item in an assessment should be reviewed by multiple teams in an effort to identify any inappropriate items whether regarding grade level of the standards being assessed, grade level of the language in the questions, developmental appropriateness of the material, or bias.

It’s worth noting that while all assessment has the capacity for harmful bias in its design, standardized assessments – when well designed – have a heightened capacity to protect against bias due to the more structured administration of the assessment – a protection harder to achieve in non-standardized assessments.

What Are Inappropriate Uses Of Standardized Assessments?

The simple answer: using assessment data for what it wasn’t designed to be used for. Academic assessments are not designed to tell you a child’s worth, or their value. Anyone who asserts otherwise is using standardized assessment data in an inappropriate manner. And as mentioned, another example of misusing data is to use standardized summative assessment data as if it’s common formative assessment data, or vice versa.

An additional topic on data use revolves around what types of decisions are made based on the data. It is useful to consider standardized assessment data when planning resource allocation and when determining whether the school system is effectively serving specific student groups. But because standardized assessment data can be correlated with socioeconomic phenomena, it is problematic to use it in ways that limit students’ access to opportunity – whether for access to courses, grade levels, or institutions.

Another potential danger deals with item types. How knowledge and skill are assessed is quite meaningful. There are generally three different ways that assessment can be conducted — three different item types. They are selected response, constructed response, and performance task. When most people think of standardized assessment, our minds automatically go to the scantron sheets of old. And that is an example of an item type. It is incredibly efficient, but what it gains in speed it lacks in, well, everything else. Scantrons are a classic example of a selected response item type.

Constructed response item types ask the student to create their own answer rather than selecting an answer from a list. And performance task item types ask the student to do something that is a demonstration of their knowledge or skill. Both of these tend to be evaluated using a rubric. A related concept, portfolio assessment, is similar to a performance task in that it is looking at a collection of previous work the student has completed. All of these can be part of standardized assessments or non-standardized assessments.

A reason that item types can be so meaningful is that how assessment is designed often influences how educators design and deliver instruction. If the standardized summative assessment is entirely a selected response, fill in the bubble style assessment, there is an increased likelihood that educators will lean in that direction with their formative and interim assessment practices. Increasingly, education leaders are waking up to this danger and leaning toward standardized summative assessment that relies more on constructed response and performance task item types, as well as portfolio assessments.

Does Constant Test Prepping Improve Performance On Standardized Assessments

No. This is a myth perpetuated by, not surprisingly, the test prep industry. What does support improved performance is a small amount of preparation regarding the item types that will be used in the assessment and the general design of the assessment. This familiarity helps a test taker be better positioned to demonstrate what they know on the assessment. That’s a very good thing. But it only takes a few hours to do this. Endless test prepping that begins weeks and months before an assessment is time wasted and is a practice that is actively harmful to children. What improves performance on standardized assessments? When students are taught the assessed material to a depth of knowledge that allows them to use the material in new ways. There are no shortcuts.

Project Management vs Progress Monitoring

Project management is about tracking the inputs – the day-to-day actions – of staff to determine whether they are likely to deliver the desired outputs. Project management is asking if the tactics being used to accomplish the strategy are being implemented effectively such that the milestones will be met. A common result of effective project management practices is revision of implementation plans and practices. Project management is a management duty, not a governance duty.

Progress Monitoring is about tracking the outputs – the interim metrics – of staff to determine whether they are likely to deliver the desired outcomes/results/goals/guardrails. Progress Monitoring is asking if the strategy being used to accomplish the results is actually likely to do so. A common result of effective progress monitoring is revision of organizational strategy. Progress monitoring happens at both the managerial and governance levels (though with significant differences in scope).

Arguments Against Measurement

Often, people will make arguments against measuring anything. While this position, taken to the extreme, is harmful and undermines our ability to serve our most educationally vulnerable children, the underlying pain behind it is not without merit. Two similar ideas summarize the strongest argument against measurement – and simultaneously animate several of our most ardent recommendations regarding the composition of effective metrics.

Campbell’s Law states, “The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor.”

Goodhart’s Law states, “Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.”

In short: the moment you start measuring something – and in particular, once you attach meaning or consequence to the measure – the utility of the data collected by the measure begins to decline. THIS ISN’T TALKING ABOUT CHEATING — that’s a separate phenomenon entirely and is a function of poor leadership, not assessment. And it’s not even just about teaching to the test, it’s just about the collapse of meaning after a while. These are sociologically and statistically observable and naturally occurring functions of measurement, not arguments for destroying the unbelievable good for children that standardized assessments can provide. While these collapsing effects are not instantaneous, they are real and warrant project management and performance management systems that are calibrated to fight against these tendencies.

Key strategies built into effective systems to combat these realities include changing interim metrics every one to three years, relying on strong outputs (for example, “how many students are on-track”) for interim metrics which are much harder to game than inputs (for example, “how many of x did we do”) or weak outputs (for example, “how many people attended/participated”), and relying on metrics that are already explicitly used for decision making rather than novel data that no one actually uses or tracks.

Measuring Progress

Once SMART goals about student outcomes have been adopted, effective goal monitoring requires four main ingredients: monitoring calendar, monitoring report, superintendent participation, and board member participation.

Effective Monitoring Calendars

Before boards can begin effective monitoring, they should adopt a 36-60 month schedule that describes which goals will be monitored during which month. The board will typically have the superintendent draft a calendar since the administration knows when student performance data is freshly available throughout the year. Nevertheless, it remains the board’s monitoring calendar, not the superintendents. Qualities to look for include:

It should span the entire length of the goals – if the goals are five year long, the calendar should be five years long as well

It should include all of the board’s goals and guardrails

It often includes all board trainings, board-led community trainings, board-led community listenings, board self evaluations, board-led superintendent evaluations, and statutory votes

It should schedule each goal to be monitored at least four times throughout the year, and each guardrail at least one time per year (on 12 month cal)

It should schedule one or two interim goals to be monitored each month, no less and definitely no more than three

It can schedule as many interim guardrails to be monitored during a month as the board wants

It should never suggest that goal monitoring reports be placed on the consent agenda, but guardrail monitoring reports may be on consent

It should clarify that boards will monitor goals during every month of the year that the board meets

It’s often helpful for the board to put all of its monitoring and non-monitoring items on a single calendar so that it’s easy to track the board’s priorities while also keeping its other obligations in perspective. Such a calendar can still typically be 36-60 months in span given the cyclical nature of district activities. Here’s an example:

Effective Monitoring Reports

Here are four qualities to ask about the 1-5 page monitoring report before the board can begin progress monitoring (if the answer to any of these is “no”, hand the report back to the Superintendent and have them complete it before proceeding – likely at the next regularly scheduled board meeting):

The Goal: Does it clearly show which specific goal / interim goal is being monitored?

The Data: Does it clearly show data for the 3 previous reporting periods (preferably on a line graph)? Does it clearly show the current reporting period? Does it clearly show the target reporting periods (annual targets and deadline target)?

The Interpretation: Does it clearly show the Superintendent’s understanding of system performance relative to the goal?

The Evidence & Plan: Does it clearly show supporting documentation that evidences the Superintendent’s understanding of system performance? If the school system is not at target or the Superintendent’s understanding of system performance indicates implementation is not on track, does the monitoring report clearly describe systemic root causes, strategic responses (including rationale), and any needed next steps?

Effective Superintendent Participation

How superintendents show up in the monitoring conversation has a huge impact on the conversation’s effectiveness. A few guidelines include:

Don’t Hide the Data: The student performance data being presented during the monitoring conversation should be easy for most parents to understand. As such, monitoring reports should be only 1-5 pages at most, and should be written at no more than an 6th grade reading level.

Don’t Sugar Coat the Data: The data is the data. Whatever it says is what it says – good, bad, or ugly. Never suggest that the data is saying anything other than what you believe it to be saying. If the school system is off track, say that; don’t talk around that. Sugarcoating loses trust.

Align Monitoring with Managerial Action: Data in monitoring report should reflect what staff are looking at to gauge the district’s effectiveness. There should be no need to create data for a monitoring session that isn’t otherwise being considered by the superintendent and cabinet.

Be Prepared: Many superintendents rehearse for monitoring conversations by having their teams throw every conceivable question at them before the board meeting. This is a wise practice not only because it helps with the monitoring conversation but because it can help surface managerial issues and solutions that might not otherwise come up.

Don’t Be Defensive: If the student performance data is disappointing, then it’s natural that board members would be disappointed. Unfortunately, not all of them will manage their disappointment in a mature, adult, and effective manner. Even if this happens, don’t get defensive.

Effective Board Member Participation

Goal monitoring, like board governance in general, is not always intuitive. It is easy to inadvertently conduct monitoring in an ineffective manner. Here are a few guidelines to follow to increase the likelihood of effectiveness:

Do Your Homework: Board members should arrive at board meetings having already read the monitoring report, having already shared technical and tactical questions with the superintendent, and having already come up with at least three or four SMART Questions each regarding the monitoring report (see During Goal Monitoring below).

Understanding Reality: The desired result of monitoring is to understand the current reality for your students as compared to the vision you’ve adopted for them (goals). Whether you enjoy the current reality isn’t the point of monitoring; whether or not you fully know the current reality is.

Keep the Conversation Going: If the superintendent presents a monitoring report that is missing the prerequisites (see Before Goal Monitoring above) or that fails to clarify for board members the extent to which reality matches the goals, consider tabling the conversation and giving the superintendent a chance to fix it and re-offer it at a subsequent meeting, instead of choosing not to accept it and ending the discussion.

No Gotcha Governance: Adopt a monitoring calendar that shows which goals will be monitored during which months and that spans the full term of the goals – for five year goals, the calendar should be five years. Then ensure board members adhere to the monitoring conversation rubric below.

Don’t Offer Advice: Monitoring is never an opportunity for board members to provide advice to the superintendent regarding what should/shouldn’t be done about student outcomes. It’s also not about liking/not liking the superintendent’s strategies.\

Goals / Summative Metrics

Definition: Quantifiable outcome measurement of achieving/honoring your school system’s vision or values

Function: Used to track outcomes; whether the school system/chief executive was successful

Guiding Question: What would honoring the vision and values of our community look like for this school system?

Formula: The [measure] for [population/area] will [increase/decrease] from [starting point] on [starting month/year] to [ending point] by [ending month/year]

Examples:

Number of high performing campuses as measured by the School Performance Framework will increase from W% on X [month/year] to Y% by Z [month/year]

Percentage of graduates persisting in their second year post-secondary without needing remedial courses will increase from W% on X [month/year] to Y% by Z [month/year]

Percentage of graduates having completed an associate’s degree and/or been awarded an industry certification by graduation will grow from W% on X [month/year] to Y% by Z [month/year]

Percentage of students reading on grade level according to ABC instrument by the end of 3rd grade will increase from W% during X [month/year] to Y% during Z [month/year]

Interim Goal / Interim Guardrail / Interim Metrics

Definition: Quantifiable output measurements of achievement that indicate progress towards the Summative Metric.

Function: Used to track outputs; quantifiable progress at critical and ongoing junctures during a cycle

Guiding Question: How will I measure whether or not I am on track to reach my Summative Metric?

Formula: The [measure] for [population/area] will [increase/decrease] from [starting point] on [starting month/year] to [ending point] by [ending month/year]

Examples:

The percentage of students on track in reading will increase from W% on X [month/year] date to Y% on Z date [month/year], as measured by district benchmark assessments

The percentage of 9th graders on track to graduate will grow from W% in X [month/year] to Y% [month/year] by Z as measured by the district’s ABC on-track indicator

Goal Milestones / Guardrail Milestones / Formative Metrics

Definition: Quantifiable output measurements that indicate fidelity of implementation of deliverables – and ideally, though not necessarily, progress toward an Interim Metric.

Function: Used to track fidelity of implementation; quantifiable progress at critical and ongoing junctures during a cycle

Guiding Question: How will I measure whether or not I am faithfully implementing the chosen tactics?

Formula: The [measure] for [population/area] will [increase/decrease] from [starting point] on [starting date] to [ending point] by [ending date]

Examples:

The number of AP seats available will increase from W% on X date to Y% on Z date

The percentage of 9th graders with weekly counselor check-ins will grow from W% in X to Y% by Z

Project Management

Definition: Process of tracking the extent to which the agreed upon tactics are being implemented with fidelity and milestones are being met

Practitioners: This should only be conducted by the level of staff responsible for managerial implementation of a tactic and their supervisors; this is a managerial function, not a governance function

Frequency: Project management should be practiced no less than once per month (though it is most commonly a weekly practice) and no more than the frequency with which implementation fidelity data is refreshed.

Function: Used to track formative metrics (inputs).

Guiding Question: How will I know whether or not I am on track to reach my Formative Metric?

Progress Monitoring

Definition: Process of tracking progress regarding the extent to which the agreed upon results are likely to occur

Practitioners: This should be conducted by staff responsible for managerial creation of results and their supervisors; this is both a managerial function (when it occurs between staff and their supervisors), and a governance function (when it occurs between the board and the chief executive)

Frequency: Progress monitoring should be practiced no less than once per quarter (though it is most commonly a monthly practice) and no more than the frequency with which lead measure/interim measure data is refreshed.

Function: Used to track interim metrics (outputs)

Guiding Question: How will I know whether or not I am on track to reach my Summative Metric?