Showing posts with label Value Added. Show all posts
Showing posts with label Value Added. Show all posts

Thursday, November 22, 2012

Data Driven Nonsense from Harvard and the Gates Foundation



A new Gates-funded Harvard study has found that Los Angeles Unified School District (LAUSD) teachers vary substantially in quality (more than in other districts) and that it disproportionately places inexperienced teachers in lower performing classrooms (as in other districts). The study, Human Capital Diagnostic, was done by the Strategic Data Project (SDP), which is connected with Harvard University’s Center for Education Policy Research.

The biggest problem with this study is that it is a bunch of nonsense.

Let’s start with the authors’ most profound claim: The best teachers in LAUSD provide the equivalent of eight additional months of instruction during the school year compared with the district’s worst teachers. Since their research was based entirely on student scores on the California Standards Tests (CSTs), a high-stakes exam used to rank schools, and the top teachers in the study were the ones with the largest student gains on these tests, what they are really saying is that the best teachers provided the equivalent of eight additional months of test prep.

Big wow!

The authors state that there is “no specific cut-off for determining whether an effect size is large or small,” but they assert that a standard deviation of 0.2 is considered large in education research. The study found that the difference between a 25th and a 75th percentile teacher is one-quarter of a standard deviation (0.25). This would be significant if it was based on a meaningful measurement of teacher effectiveness. Unfortunately, all it really says it that some teachers are better than others at squeezing out student gains on an otherwise lousy exam. It does not tell us whether their students are becoming self-motivated, independent learners or competent critical thinkers and problem-solvers. Furthermore, the study provided no explanation for how it determined that 0.25 standard deviations was equivalent to eight months of instruction.

The authors also claim that Teach for America (TFA) and Career Ladder teachers have higher effects on their students than other novice teachers by 0.05 and 0.03 standard deviations and they even attributed a gain of one to two months in additional learning to these relatively small standard deviations. They make similar claims for National Board Certified teachers, whose students test gains were 0.03-0.07 standard deviations higher than those of other teachers. Yet, if a standard deviation of 0.2 is considered large in education, then a standard deviation of 0.03-0.07 ought to be considered small or even insignificant.

While the standard deviation may be insignificant, the fact that this was being researched in the first place is not. TFA provided 13% of new hires to the district over the past six years (according to the study’s authors) and it would be of great interest to the district’s administrators to show that the investment was worthwhile. So let’s assume for the sake of argument that the difference between TFA recruits and other novice teachers was significant. What would this mean? TFA teachers may in fact be more willing than other novice teachers to work long, unpaid overtime hours and substitute quality student-centered instruction for “drill and kill” style teaching, both of which could produce higher test scores without improving the quality of student learning.

Perhaps a bigger problem with this study (like all studies and reforms based on student test data) is that numbers are not the only relevant type of data in education and sometimes not even the best. Ester Quintaro, writing for the Shanker Blog, talks about the “streetlight effect,” from the parable of the drunk who searches for his lost wallet under the streetlight, not because he lost it there, but because the light is better there and it would be easier to find it if it happened to be there. Student test data is easy to access now that it is required of every district in the U.S. under No Child Left Behind (NCLB)—it is under the streetlight.  Yet, at best it is only a proxy or very rough estimate of teacher quality since it only considers a small part of what teachers are expected to do.

Quintara also correctly points out that NCLB has helped to institutionalize what counts as data. “Scientifically-based research” is now limited to standardized test scores, which, as it turns out, are not particularly scientific. Case studies, ethnographies, teacher observations and portfolios, and other qualitative data are considered unacceptable.

One promising finding from the study was that teacher performance after two years was found to be a good predictor of future effectiveness. In other words, the current system of giving tenure to teachers after two years of good evaluations makes sense. Teachers are not getting worse after two years. Novice teachers are not better than veterans and should not have the right to bump them during layoffs and LAUSD is not top heavy with a bunch of cranky veterans who can no longer teach.

Monday, October 22, 2012

LAUSD Lays Down Hammer on Evaluations

Image from Flickr by r8r

Los Angeles Unified School District (LAUSD) has filed a declaration of impasse, according to the Daily News, after failed negotiations with United Teachers of Los Angeles (UTLA) over a new teacher evaluation system based on student test scores. LAUSD is under court order to revise its evaluation system by December 4. However, Superior Court Judge James Chalfant has mandated that the district negotiate with the union over the new system.

UTLA President Warren Fletcher said last week that his 40,000-member union was engaged in "good-faith bargaining with LAUSD officials over developing a fair and effective teacher evaluation system.” This, of course, is typical union mumbo jumbo meant to convince the public that the union was playing by the rules and trying to do right by the students and that any blame for the stalemate lies squarely on the shoulders of LAUSD.

While Fletcher’s quote may sound good to the press, it is patently untrue. If UTLA was really interested in creating a fair and effective teacher evaluation system they would refuse to accept any use of student test data in their evaluations, as such data is unreliable, inconsistent and leads to many false positives and negatives (see here, here and here). This is obviously bad for teachers who could receive bad evaluations despite being good teachers simply because they work in a low income school with the perennially low test scores that are common in lower income schools. However, it is also bad for children in several ways. They could end up losing excellent teachers because of the inaccuracies inherent in this evaluation system. Conversely, bad teachers could easily slip through the cracks and remain in the classroom because they happen to work in higher income schools, which tend to have higher test scores and larger gains on their scores.

If UTLA and LAUSD were truly interested in a fair and effective evaluation plan they would demand that well-trained outside evaluators be brought in to evaluate teachers blindly, using a combination of classroom observations and portfolios. This would eliminate the bias inherent in being evaluated by the boss (i.e., site administrators), who may ding a teacher for not embracing and carrying out his pet projects and reforms with sufficient vigor or for speaking out on children’s or teachers’ behalf at faculty or board meetings. It also would eliminate the problem of site administrators being poorly trained and lacking the time to make sufficient and competent observations and evaluations. And it would eliminate the bias and problems inherent in the use of student test data.

That LAUSD is declaring impasse suggests that they are fed up with UTLA’s position on the matter. Yet UTLA, despite Mr. Fletcher’s criticisms, has embraced student test data to evaluate its teachers. The big stumbling block, at this point, is that LAUSD wants the data to be based on individual classrooms, which can be directly linked to individual teachers, whereas UTLA wants it aggregated school-wide.

The union is also saying that it wants evaluations that provide useful feedback for teachers so they can improve their practice. Yet regardless of how student test data is acquired or aggregated, it fails to provide such data. This is because the test scores are a measure of student test taking ability. They tell us nothing about how students learned the content or developed their test taking skills and their scores are influenced far more by their socioeconomic status than by their schools and teachers.

UTLA, having already accepted the use of student test data, is unlikely to strike over the matter, especially when the district is under court mandate to include student test data in its new evaluation system. Unions have become overwhelmingly averse to challenging court orders and injunctions (e.g., the Chicago Teachers Union, which supposedly struck over student test data being used to evaluate teachers was, in reality, only fighting over the extent to which it would be used, having already  accepted that it was required by Illinois state law). Thus, the question is not whether, but how, student test data will be abused to evaluate teachers.

Sunday, September 16, 2012

Chicago Teachers, Just Say No


The Chicago Teachers Union (CTU) and Chicago Public Schools (CPS) are were very near a deal by the end of last week. Last minute negotiating today, it is hoped, will lead to a compromise that both sides can accept, allowing teachers and students back in the classroom by Monday. The 800 members of the CTU House of Delegates would have to approve the contract and then teachers would have to ratify it.

The details of the proposal have not yet been made public, but speculation by the mainstream press suggest that it will include the use of student test data in teacher evaluations, but these scores will not count as much toward teachers’ overall scores as CPS had hoped. If this is true, then it is a terrible contract and Chicago teachers should just say no to it.

The test scores are an inconsistent and unreliable proxy for teacher quality. They lead to many false positives and false negatives (also see here, for more on the unreliability of Value Added models). Teachers who are doing everything right in the classroom can still have students who do not progress sufficiently in their test scores, which are correlated far more strongly with students’ socioeconomic backgrounds than with their teachers (see herehere and here). Furthermore, any use of test scores to evaluate teachers will encourage teaching to the test, while also giving tacit approval to the state’s wrongheaded mania for testing.

For all of these reasons, teachers unions across the nation must oppose the use of high stakes tests, period. They are virtually worthless as a tool for assessing teachers and they are destructive of the learning process for children.

Wednesday, August 15, 2012

Giant CTA Sellout: Student Test Scores Could Be Used In Teacher Evaluations


Huck/Konopacki Labor Cartoons
Assembly Bill 5, which will revamp how teacher evaluations are done in California, has been revived thanks to recent support by the California Teachers Association, according to John Fensterwald, writing for Ed Source. CTA lobbyist Patricia Rucker said that it “is a clear and good policy document” now that some of CTA’s preferences have been incorporated into the revised bill.

So what does the bill do to evaluations and how has it been amended to appease CTA?

Mostly it codifies the existing standards for the teaching profession and mandates that all districts use them when evaluating their teachers. These are very reasonable and appropriate expectations for teachers such as engaging and supporting all pupils in learning, setting high expectations, creating and maintaining effective learning environments, and knowledge of content standards.

However, the bill adds a new and potentially dangerous standard: “Contributing to pupil academic growth based upon multiple measures, which may include, but are not limited to, classroom work, local and state academic assessments, and pupil grades, classroom participation, presentations and performances, and projects and portfolios.”

The danger here is that these measures assess where students are, not how they got there. Thus, they do not actually measure teacher skill. Since students’ performance and growth in these assessments is significantly influenced by their socioeconomic backgrounds, English language proficiency and special education status, many teachers will be evaluated poorly due to their students’ backgrounds rather than their own teaching ability.

Furthermore, Value Added Measures (VAM) that rate teachers based on student progress on standardized exams, are notoriously unreliable. Studies show that even when used correctly they are only accurate for the very worst and the very best teachers, not for the vast majority of teachers who lie somewhere in the middle.

So why would the state’s largest teachers’ union concede so much?

Fensterwald argues that they might be responding to a recent court ruling. In June, Superior Court Judge James Chalfant’s ruled that Los Angeles Unified School District (LAUSD) was violating the Stull Act, which requires school districts to use student scores on state standardized tests when evaluating teacher effectiveness. Since AB 5 permits (but does not require) districts to use state test scores in teacher evaluations, CTA may believe that the law preempts the judge’s ruling and will buy them some breathing space.

This is a risky and stupid game for the union. Once student test scores become permissible as evidence of teacher quality, districts will try to impose it locally, while politicians and Ed Deformers will push to make it mandatory for all teachers under state law.

CTA ‘s Patricia Rucker has called the issue overblown, since the existing tests are slated to be replaced by new ones aligned to the Common Core Standards (CCS) and it remains to be seen whether those new tests will be “suitable for teacher evaluations.” However, her perspective shows an incredible naïveté. Considering who is profiting from the CCS and the new exams, it is incredibly unlikely that these tests will provide data that is any more reliable or meaningful than the current state tests. Furthermore, all student tests measure students’ performance, not teachers’. And since they are all influenced by students’ socioeconomic backgrounds, a good teacher can still end up with low student test scores.

Even with CTA’s support, AB 5 still faces an uphill battle. Those on the right are criticizing the revised bill because it does not require test scores be used in teacher evaluations. Still, if the bill passes, it could be years before it would be implemented (if ever). Because the Stull Act requires that the state reimburse local governments for mandated programs, AB 5 would cost the state millions that it currently does not have. Furthermore, the bill would be delayed by a minimum of seven years, until after the state has repaid school districts money it owes them due to budget cuts, a sum currently valued at more than $10 billion.

Monday, August 13, 2012

Chicago Teacher Strike Looming?

Huck/Konopacki Labor Cartoons

Students begin classes Monday at the roughly one-third of Chicago Public Schools (CPS) that are on a year-round schedule. Meanwhile, Chicago Teachers Union (CTU) president Karen Lewis is saying that a resolution of contract negotiations before then is virtually impossible, the Chicago Sun Times reported on Friday. Lewis went on to say that they hadn’t even started talks on compensation because they have been focusing on the smaller items during their 41 bargaining sessions. CTU is warning all teachers to prepare for a strike, as contract resolution may still not occur by September, when the rest of the schools are set to open.

A work stoppage on Monday would necessarily be a wildcat strike, however, as the union is required to give a 10-day warning to CPS and they have been forbidden from striking before August 18.

CTU is not going to sanction an illegal strike, especially with an interim deal still on the table. The interim deal would include the much touted longer school day for students, but not for teachers. This would require the hiring of many more teachers and bring many of the recently laid-off teachers back to the classroom.

The deal was a smart move by CPS, which has been demanding numerous concessions from the teachers. The longer work day may have been the most onerous of these concessions, as the district was demanding a 90-minute longer work day from teachers without extra compensation. It was also an easy one to give up (at least for the short-term), as an arbitrator recommended the district give a 20% raise for the longer hours. Thus, the district can spin itself as being reasonable and fair. It is hoping this compromise will convince teachers to accept the other concessions and avoid a strike—not an unreasonable expectation considering how averse to strikes teachers and especially unions have become.

One of these other concessions is CPS’ demand that student test scores be used in teachers’ evaluations and job security, something all teachers should oppose as the scores are unreliable and correlate much more strongly with students’ socioeconomic backgrounds than with teacher skill (see here, here and here).

Wednesday, August 8, 2012

Another UTLA Sellout—Evaluations Can Include Test Scores


Huck/Konopacki Labor Cartoons
The United Teachers of Los Angeles (UTLA) recently agreed with the Los Angeles Unified School District (LAUSD) to allow the use of student test scores as part of performance reviews beginning this fall, the Los Angeles Times reported, though a UTLA attorney later said the commitment was contingent on whether the union and LAUSD could negotiate an agreement on how the scores would be used in the evaluations.

I am calling the agreement a sellout not only because UTLA had recently opposed using such data in teacher evaluations, but because the data are not an accurate measurement of teacher skill or ability.

The accuracy of such Value Added Measures (VAM) is still up for debate. Teacher trainer Grant Wiggins argues that VAM “models accurately predict over a three-year period, performance at the extremes,” which means that IF you average VAM scores over three years, you can identify the really great teachers and the really lousy ones. The vast majority of teachers (who fall somewhere in the middle) would thus be getting inaccurate VAM scores and potentially bad evaluations as a result. Furthermore, because most school districts that use VAM are using them to evaluate teachers on a yearly or biyearly basis, even those falling at the extremes may be getting inaccurate VAM scores since they are not averaging their scores over a three year period.

This, alone, is a compelling argument against using VAM. However, there are a host of other good reasons not to use VAM.

One of the assumptions of VAM is that a good teacher can help low income students improve as much as higher income students. This is not necessarily the case. Wealth does not simply cause students to earn higher test scores, but provides a variety of advantages that benefit affluent students throughout their lifetimes, including better health, greater access to enriching extracurricular activities, and a significantly lower risk of low birth weight, malnutrition and environmentally-induced illnesses. This decreases the chances that an affluent child will develop learning disabilities or impaired cognitive development and may increase how quickly they can learn and how much of the learning is retained. In other words, teachers at affluent schools may see greater gains in student learning because of their students’ socioeconomic backgrounds.

How much a student improves from year to year is also dependent to some extent on their previous teachers. For example, a chemistry student who had a bad math or science teacher the previous year may be lacking so much of the prerequisite knowledge and skills that their growth in chemistry is severely limited.

Despite the apparent acquiescence by the teachers’ union, there was still criticism from an attorney representing parents who sued the district for “violating” the four-decade-old Stull Act, which requires the use of student achievement in teacher evaluations. The attorney accused UTLA of saying one thing in court and then changing their position.

The Stull Act, however, requires the use of student achievement data, not necessarily test scores. It also does not specific exactly how that data should be utilized.

As terrible as it is for their members that UTLA has agreed to allow the use of such data in their members’ evaluations, at least they are trying to maintain some involvement in determining exactly how that data will be used. This can hardly be seen as being duplicitous, as the parents’ attorney has suggested. Rather, it sounds more like an attempt by the union to collaborate with management in the abuse of their members.

Wednesday, May 16, 2012

VAM Bashing From the Right


Jay Mathews, the conservative foil to Valerie Strauss at the Washington Post, admits he likes Value Added Measures (VAM) in theory, but concedes that the reform is misused and abused and likens it to an action film monster that must be destroyed.

The “best” criticisms he has seen came from teacher trainer Grant Wiggins who points out that VAM “models accurately predict over a three-year period, performance at the extremes.”

In other words, if you average VAM scores over three years, you can identify the really great teachers and the really lousy ones.

Assuming this is true, the vast majority of teachers—who fall somewhere in the middle—would be getting inaccurate VAM scores and potentially bad evaluations as a result. Furthermore, because most school districts that use VAM are using them to evaluate teachers on a yearly or biyearly basis, even those falling at the extremes may be getting inaccurate VAM scores. Thus, no one is being accurately assessed by VAM.

While this is a compelling argument against VAM, there are a host of other compelling criticisms.

One of the assumptions of VAM is that a good teacher can help low income students improve as much as higher income students. This is not necessarily the case. Wealth does not simply cause students to earn higher test scores, but provides a variety of advantages that benefit affluent students throughout their lifetimes, including better health, greater access to enriching extracurricular activities, and a significantly lower risk of low birth weight, malnutrition and environmentally-induced illnesses. This decreases the chances that an affluent child will develop learning disabilities or impaired cognitive development and may increase how quickly they can learn and how much of the learning is retained. In other words, teachers at affluent schools may see greater gains in student learning because of their students’ socioeconomic backgrounds.

How much a student improves from year to year is also dependent to some extent on their previous teachers. For example, a chemistry student who had a bad math or science teacher the previous year may be lacking so much of the prerequisite knowledge and skills that their growth in chemistry is severely limited.

What Does it Mean to Be a “Really Good” Teacher?
Most would argue that there are certain easy to identify practices that characterize a “good teacher” like having a strong background in the content, creative and effective lesson design, good classroom management and a positive rapport with students.

While any teacher who has these qualities ought to be considered a “good teacher,” in reality the teachers identified by administrators as “great teachers” are often the ones who come in at 6 or 7 and stay until 6 or 7. They may in fact be excellent teachers, too, or their VAM could be a reflection of how many extra unpaid hours they are putting in.

Some would likely argue that this is a legitimate use of VAM: A teacher who puts in long hours for her students and gets them to perform better deserves a good evaluation, promotion, bonus pay, etc. However, it is not fair or reasonable to evaluate teachers on whether or not she puts in unpaid volunteer time over and beyond that required by her contract. Under this scenario, an excellent teacher who works the contractual hours or, as most of us do, who works more than the contractual hours, might still get a lower VAM score than a martyr who puts in 70-80 hour weeks.

Friday, March 9, 2012

VAM, Blam, Thank You Ma’am


Huck/Konopacki Labor Cartoons
In a rare moment of lucidity, the New York Times published a piece this week pointing out how even the best teachers (or those with the “best” students) can end up with terrible Value Added (VAM) scores and potentially face reprisals or get fired as a result.

How is this possible?

If 85% of a teacher’s students are proficient in reading but 95% were proficient the prior year, she would earn a low VAM score because she is ostensibly doing a worse job than she did last year. Rather than adding “value” she has supposedly “lowered” the quality of education. Never mind that 85% of her students were proficient—a respectable number that should be honored, rather than punished.

Yet every year our students are different and their scores fluctuate for various reasons that have little to do with teaching, including variations in the tests themselves. Social class is the single biggest influence on test scores. So if a teacher winds up with a less affluent student population one year, test scores are likely to decline. Other nonteaching factors may come into play as well, like how the school structures the exams (e.g., all in two days, or spread out over 1-2 weeks; providing brunch for students; having teachers proctor their own students) or traumatizing social disruptions, like a tornado or earthquake.

One of the inherent problems with the current use of high stakes exams (aside from the fact that they tell us virtually nothing about the quality of teaching or what students have learned) is that they are based on moving targets. Rather than testing if students have reached a benchmark (e.g., being able to comprehend a short passage or use the Pythagorean theorem), they compare students to their peers, with some necessarily always being below the average.

Teachers at a low performing school could help their students make large gains from one year to the next, only to find that similar schools made the same progress, thus precluding them from achieving their NCLB progress goals. Likewise, there is only so far high achieving students can go, resulting in low VAM scores for their teachers when they hit this wall.

Tuesday, January 31, 2012

Teachers Offer the Wealthy an Escape from Poverty


The following is from Anthony Cody’s excellent critique of the education portion of Obama’s state of the union speech. In a nutshell, if a teacher really did increase the lifetime income of a classroom by $250,000, so what? Spread out over the 40 years of each student’s career, that would amount to only $250 per person per year, enough for a nice date, Sunday afternoon beers or some practical work clothes, but nowhere near enough to bring poor kids into the middle class. And if he really wanted us to not teach to the test, he wouldn’t tie our evaluations and salaries to students’ test scores.

Last night in President Obama's State of the Union address, he repeated a familiar refrain about the importance of teachers. 

A great teacher can offer an escape from poverty to the child who dreams beyond his circumstance.
But it seems that it is those in power who are actually using teachers to escape from the realities of poverty these days. 

President Obama offered as evidence a citation from a recent Harvard report:
We know a good teacher can increase the lifetime income of a classroom by over $250,000.

He went on to say,
Teachers matter. So instead of bashing them, or defending the status quo, let's offer schools a deal. Give them the resources to keep good teachers on the job, and reward the best ones. In return, grant schools flexibility: To teach with creativity and passion; to stop teaching to the test; and to replace teachers who just aren't helping kids learn.
 
There are several problems with this. As others have pointed out, if you take a classroom of 25 students, and spread $250,000 over their 40 years of earnings, this amount comes to a grand total of $250 a year per student. This is unlikely to represent an escape from poverty. (see more thorough responses to the Chetty report here, and here.)

The second problem is a glaring contradiction, a logical flaw so huge it has been overlooked by almost every journalist apparently too polite to challenge the administration on it. If you do not wish teachers to teach to the test, if you want them to be passionate and creative, then how can you insist that their performance be measured by the use of test scores?

Let us be crystal clear. The Obama administration has made the use of test scores to evaluate principals and teachers a pre-condition for federal aid. Both Race to the Top and the NCLB waivers require that states develop evaluation processes that incorporate this data.
To see the rest of this article, please click here.