Showing posts with label Testing. Show all posts
Showing posts with label Testing. Show all posts

Sunday, December 11, 2016

End-of-Year Testing

(I started this article several years ago. Just getting around to editing and publishing these drafts.)

I love the idea of End-of-Year testing in a purely hypothetical sense. Students should be able to demonstrate what they've learned. The teacher doesn't grade it or write it. The students can't weasel out of it. A group of math people decide what "Algebra I" should entail and write some questions to measure it.

The test is never perfect, as NY teachers will hurry to emphasize, but it is out of the hands of the teacher and that is good. We should not be afraid to let our students measure themselves against a common standard and we should be open to change if the unexpected happens. SATs serve this purpose as well.

Differences between what you expect them to get and what they get are the prime indicator. If your grades are all As and your kids can't succeed in appropriate tests, then you need to review what you are doing. If the majority of kids can't succeed on an EOY exam, then you the teacher needs to make a determination: is it the individuals, the exam, the curriculum, or me? If you are passing kids and the next teacher isn't, someone might need to adjust.

In theory, EOY exams should be perfect for this - it's just too bad that they'll be useless for it.

not linear.
The State of Oklahoma requires these tests and they will *attempt* to write them to be an honest assessment of the skills that should have been acquired in any particular course. Just like Texas, just like New York Regents.

What will happen is that End of Year testing, like WilyECoyote, is going to run full-speed into the Cliffs of Reality. The "passing" score will be set to 60%, then too many kids fail so it's quietly lowered to 45%, or 35%, or lower. If enough people still can't pass the test, the cut-score will be lowered again. Or you wind up with the weird raw score conversion charts of the NY Regents (right).

So ... is it poor preparation and teaching or poor test-making?

The graph that's been misunderstood
by admin everywhere.
Or could it be that the test is trying to apply the "Higher Standards" that everyone is crowing about? You know the trope: "Raising the Bar improves performance."

Unfortunate reality #1: If you raise the standard, more people will fail to reach it.
Unfortunate reality #2: Calculus kids do better on their SATs than Algebra I students. Selection bias. Duh.


What should we do?

Avoiding all EoC testing is silly. Pretending that some "3-week portfolio question is demonstrably superior" is the canard put forth by all those people who have never watched or judged a science fair. Individual teacher-written final exams are suspect because of quality-control issues and because of grading irregularities. Department-wide final exams are probably the best unless your state has Mr Honner holding your Regents exams to account, in which case, go with the Regents.

Unless your school just voted to eliminate all finals in light of the transition to Proficiency-Based Grading, but that's another post for another day.




Wednesday, November 4, 2015

Incorrect Data isn't Useful

The other day I went to the Health Center for a followup checkup. I had been in previously and had gotten some anti-biotics for an insect bite that got infected. Simple, right? As part of the visit, the nurses are instructed to take routine weight and blood pressure measurements.

I know my blood pressure, so I was surprised that her diastolic reading was 20 points lower than it normally is. I remarked on that. Her reply was "Lower score is good, right?" in the tone of voice that conveyed clearly that I shouldn't be questioning her.

I'm thinking, "Sure ... unless it's a bad measurement." I get that BP is inexact, but it's a bit silly to refuse to re-measure it when the patient points it out. 20 points can make all the difference to the doctor's diagnosis of my overall health.

I decided that I would request the printout from the front desk as I left, the one with all of the day's numbers and decisions from the visit. I read the scale ... same weight as two weeks ago. On the printout, though, it was different ... she had obviously transposed digits when entering the data. In two weeks, I had "gained 23 pounds" and then lost it again in the 30 minutes it took to drive home. My blood pressure changed by 20 points, and that's a lot.

Bad data makes for inappropriate diagnosis.

Bad data makes for bad education policy, too.

The state of Vermont is "suffering" through the release of the first round of SBAC scores despite our scores being better than most other states (we're usually top 5).  "Results are much lower" and already my principal is bitching about it, despite declaring at the time, "We don't care what the scores are, we just want to get the process right." (I'm paraphrasing but that was the intent.)

I'm all for improvement, but I hate basing change on the back of bad data.  Our diagnosis is flawed because our data is flawed, and the prescription runs counter to other policies that the State has imposed.

First, the SBAC has measurement errors just like my nurse had.  Many students took that test knowing that scores would not be held against them, that there was absolutely no chance that anyone would see the scores in fewer than six months or act upon them to set courses for this year or college applications. Additionally, the test itself is drastically different in format (and it's all done through the Chromebook) ... a test completed entirely on-line.

There are no multiple choice questions and kids can have scratch paper, but they're not used to doing math that way. There's a lot of "drag the factors to the answer box" and write three paragraphs explaining why you know that this is a straight line .. and few can stretch out an explanation that far.

Second, and just as  important, the SBAC "passing scores" were decided upon after the fact, to make the percent-passing numbers match what the state had decided they should be ...

That's right. Before the kids even took the test, they told us that there would be a state-wide passing rate of 33% on the HS math test. Then they set the cut-score to match.

Third, add in the fact that we are a small school and we pride ourselves on being able to provide a more personalized education that your average public school, including having personalized learning plans that had quite a few students taking Algebra 2 as seniors. I'm sure you can see where this is going: many of our kids were taking a test heavily based on mathematics they hadn't seen yet.

This runs directly counter to another major initiative in the State of Vermont, the Personalized Learning Plan. Sometimes called "Personal Pathway to Graduation", the initiative requires schools to design different course pathways to graduation for each student as appropriate. This includes allowing schools to schedule certain kids into a faster progression for math and others into a more moderately paced path that might not even include algebra 2. It means that "pre-algebra, algebra 1, geometry, algebra 2" might be the most appropriate for a student.


Taking that approach and then complaining that they don't know algebra 2 by March of their junior year is silly.

It's the rhetorical equivalent of reading a graph that says that Pre-calculus students do better on the NAEP and then concluding that we must make sure every student takes Pre-calculus by the time state tests are given in the junior year
 ....

which a previous principal actually said.

So when the bright bulb in the room points out that we teachers should prepare the kids for college and careers, and should have prepared the kids better for this test, and "If you hold the kids to a higher standard, they'll rise to meet that standard," I will calmly channel Dick Cheney and say that "you go to war with the students you have, not the students you wish you had."

Finally, the teachers are not allowed to know what's on the test.  I don't want to teach to the test, but I'd like to know what is included.I'll give the same assessments I would already have planned, but I might change some questions to a similar format, for example.

Also, I'm not willing to just take their word for it that the test is appropriate. We can't check for bad questions that might have tripped up our students and we can't check that the answers they gave were correct or not. We have to take Pearson's word that the scorers actually knew what they were doing.  After reading Todd Farley's book, Making the Grades: My Misadventures in the Standardized Testing Industry and others with similar tales of the realities of corporate test making and scoring, I'm not particularly willing to do that.

The NY Regents is an example of a relatively open and transparent test-making system, but it has many errors. The NY teachers can catch these problems and get them fixed. If we look at Mr. Honner's long-running series reviewing the NY State Regents exams in mathematics, why should we expect that the SBAC tests will somehow be perfect if there is no chance for oversight?

The SBAC is a closed system with no accountability that scores the tests in strange ways, fails to take into account the realities of the students, will not allow anyone to analyze or even examine any of the questions (unlike the SAT which I can see in its entirely within a few weeks), spits out pre-determined results that do not reflect student abilities, and makes everyone wait an unconscionably long time for those results ... much too long for the school to do anything with them.

I can't use the scores because they aren't detailed enough, timely enough or accurate enough.

I guess I'll just teach math and ignore all that bluster.

Monday, June 15, 2015

Playing the Game

Or perhaps the Principal encourages those who will do poorly to "opt out" ... by telling them they don't have to be in school that day with no consequences. Then, when they skip, they're not truant, they're opting out.


It is so easy to game the numbers.

Saturday, November 8, 2014

Testing Paradigm Needs to Change

Testing in the United States is a sick, diseased system. It is a malignant tumor that must be excised if we are to ever use testing results to improve students, teachers, schools,

Testing in the USA is NOT intended to help teachers or their students. It is only done to give a number that can be used or not, at the whim of the reader. Since most testing is for evaluative purposes, testing provides numbers to punish people with.

As a teacher, I get absolutely no useful information from standardized testing.


None.

On our "Local Common Assessment", I get to know a RIT range and a corresponding percentile, and breakdowns in "Algebraic Thinking, Real&Complex Number Systems, Geometry, Statistics and Probability."

Then, consider that we have our kids taking a test and one of the categories is Real & Complex Number systems - Really? These are 9th graders in algebra 1 ... is the score range of 233-245 based on their less-than-complete knowledge of real numbers combined with no questions on complex numbers or is that 75%-ile based on questions that they would have no reasonable knowledge of?

Okay ... I'm ready to adjust my teaching for Algebra 1 ... What changes should I make?

I see none of the questions, none of the individual responses. I have no idea what kinds of things the test-makers considered to be "Algebraic Thinking" nor do I have any sense of what my students might have replied or understood or didn't, except for the kids who told me they just clicked at random just to be finished more quickly.

Okay ... I'm ready to adjust my teaching for Algebra 1 ... What changes are appropriate? Does the kid who scored "LO" really not understand or is she just lazy?


Yeah, that's the breakdown measurement: LO, AV, HI. Useful? No.

And this is a Pre-Algebra class with some 9th and some 10th graders. I would hardly expect them to get anything other than LO. If they could, they wouldn't be in the class.

Okay ... I'm ready to adjust my teaching ... What changes to my pre-algebra curriculum are appropriate here?

But at least I got those few bits of data within a week, because it was a local assessment.

When it comes to SBAC and PARCC, the problems seem to be the same as for NECAP, and before that, the NSRE.  Too few questions, coupled with long wait times for the scores (test in October, scores in April) and very dodgy scoring of the results for the constructed response questions ...

and we're still not allowed to see the questions, see the scoring, see the individual results ... 
And there was no way you could trust those scores because of the manipulation of the raw score conversion tables for "continuity reasons."

Can't have a big improvement year to year because reasons. The first year of every test has to have similar results as the final year of the test we threw away, so yr1 NSRE was first 58% passing, but was re-scored so we only had 30% passing.

If we're getting rid of a test because it isn't working appropriately, why do we insist that the new test's scores match up with the old test's scores?

And about those scores ... I have never understood how the entire public school populations of five New England states can show results in the way they did:

Highly Proficient: 3%
Proficient: 30%
Below Proficient: 40%
No Evidence of Proficiency: 27%

Really? 33% "passed" a test and you're looking at the teachers, not the test? Of all the kids in all the classrooms with all of the teachers (in VT, NH, ME, RI, and somewhere else that's escaping me right now), how is it possible that only 33% of the students passed a test?


At least the SAT is open ... maybe we should use it instead of paying Pearson far more for less information.

If you can so blithely manipulate scores so as to get a result that your statisticians declare appropriate, maybe the problem isn't in your teachers or your students ... your system needs to change.

If you can so blithely assume that the teachers are the only ones who are responsible for scores but shouldn't be allowed to see any of the test papers or any of the questions ... your system needs to change.

If you can so blithely assume that the students are always "participating fully" and that the results on this worthless and pointless (to them) test, then your system definitely needs to change.

Wednesday, March 5, 2014

A Response to the Income Gap Question and Comments.

Bumped to the top from 2009 ...

Let's recap the last few posts, shall we? I've come to the conclusion that the income-scores correlation that shows up in every testing situation is actually rooted in the parents and their attitudes about school and education.

The smarter, motivated, dedicated parent who is demanding of a good education is also likely to have passed similar traits to his child which, while not being a guarantee certainly shows a strong correlation to that child's success. Show me a bored, unmotivated, or stupid child and I'll bet that you've got a bored, unmotivated or stupid parent. We just had open house and all that I saw reinforced my feelings on that - all of the visitors but one was the parent of a hard-working kid.

Yeah, yeah, I know. There are exceptions and everyone can point to the kid who breaks the mold. But averages don't care about your exceptions.

How else to explain that persistent gap between the rich and the poor? There's the racial gap, but that disappears when you control for income. Likewise for the gender gap, or the city-size gap. Examine the increases for repeat test-takers and you can see that even here, the differences aren't as great as that for income.

It's not a simple A-B correlation: More money does not make a kid smarter. The cause-effect isn't there. We need the confounding factor that causes both effects, sort of like the correlation between the amount of "Bling" and the low death rate in cars. Does "Bling" prevent deaths? No, but the ability to afford it means that you can also afford a safer car.
Jonathan asks "If I can paraphrase: Smart adults --> high income, Smart adults --> smart kids. If I have this right (and please correct me if I do not) I think it is bunk."
He then goes on to describe the exact things that do allow a wealthier parent to raise better students (and those traits invariably show up as "being smarter"):
"intellectually enriched environment, ... more expensive towns ... schools maybe not better, but at least better-funded. ... stimulating things around ... babysat by a reader ...more likely to be exposed to culturally enriching stuff. Plays ... the library ...travel ... fewer stresses from neighborhood, hunger, family ... depends on the parents' wealth, not on the parents' smarts."
And I agree. But what affects the wealth? I would maintain that any person can rise in this world, though for many it is harder than it is for others (racism, classism, sexism, and bigotry are alive and well -- the situation is improving but.) Being smart and motivated is a good start. Lazy and stupid are starting well behind the stagger.

To truly improve the students' performance, you can't just give them money. You can feed them and eliminate some stresses and move them into a wealthier town ... and get nowhere because that's not the cause in this relationship. You can't just move into a better neighborhood and suddenly improve your scores.

But the better neighborhood DOES have better schools and better students. What is cause and what the effect? We need to look at the acquisition of wealth. On average, which type of person will acquire wealth, a boorish lout or a diligent, studious worker?

Over the population, which group of people will stay in a company and rise in the corporate ladder, the smart and driven ones or the unmotivated slackers? Which type will job-hop (or be fired) and remain at the lowest levels of the many companies' pay-scales? Which type will recover from a major setback or rise from the slum and make something of himself? Which will read books, learn math, and practice speaking without an accent that labels him as uneducated? Which type of person will motivate his kids and provide a richer environment than that of his neighbors? Which single mothers will rise at 4:30am to tutor their children -- the alcoholic or the mother of a future President?
Pissed Off Teacher asks "How about all the money rich parents spend on SAT prep classes and private tutors? Lower income kids cannot get this extra help."
They don't get that help and that's a shame. I'd like to see ALL schools offer a half-credit SAT prep course, basically a review of math and English. Now that they've had a taste of what they'll need it for, they might be prepared to pay better attention and break out of the bad habits that they started school with.

I should also mention that the most-advertised grade bump that Kaplan Test Prep and others provide is simply the elimination of the common mistakes. Meeting once or twice a week for ten weeks is not enough time for more than test-taking strategies and gimmickry. Their "guarantee" of 100-point gains are backed by an offer to retake the course free. If the offer was for your original money back, then I'd listen.

Darren, who's obviously been hearing waaaaay tooo much tax-raising talk from his Governator said...
I know what let's do! Let's tax the rich, and if that doesn't work, let's tax them some more! Then, when there's no rich left, there will be none of this gap between rich and poor!
We'll have to cut him some slack. His "Republican" governor has been sounding like a tax-and-spend liberal Democrat lately.  That and those deficits would make anyone cranky.

So what do we do now?

I'm not entirely sure, to be honest. I'm not sure that the income-gap problem CAN be solved because it has already happened; what we are seeing in our classrooms is the fallout.

We already offer free-or-reduced lunch and breakfast. We already introduce them to the things that make intelligent, well-rounded, and educated people and try to help them break whatever mold they're in. We already offer extra help in the corridors because Packemin HS doesn't have any extra rooms.

What we really need to do is to stop agonizing over scores. We should look for improvement in each student and try not to worry so much about improvement from year to year and from cohort to cohort. We certainly should ignore the idea that 100% will achieve proficiency in four years.

Spending more money on the schools might or might not be the answer -- it really depends on whether you've been spending the right amount in the first place. If you've been undermining the system for years, cutting back further certainly won't help. If you're at the right level of funding, spending more won't help either.

Teach the ones in front of you. Do what you can. Don't try to save the world.

There, I said it.

SAT scores are linked to Family income. So what?

Bumped to the top from 2009


Just a warning ... I'm going there.

Flypaper, the edExcellence Blog, has commentary on the SAT scores that refuse to go up.
"Gaps widening a bit by race, income, parental education. Indeed, the tidiest relationships and smoothest curves are those that continue—as they have for as long as anyone can remember—to show the steady upward progression of average SAT scores (pdf) as family incomes and parents’ education rise."
Chester then goes on to show some details and to say,
What does this say about 26 years of education reforming since A Nation at Risk? For starters, it says the reform efforts haven’t seriously penetrated our high schools. Then it says that current moves (e.g., the “Common Core” national standards project of the governors and chiefs) to align high-school exit expectations to college and workforce readiness are urgently needed, indeed long overdue.
That's an interesting take. If you can't get results in 26 years of trying reform after reform, let's try another reform! (Insert sound of Bells and Whistles). Then there is the sentiment that "Reforms are obviously urgently needed."

It's true. The only correlated data that ETS has with an R-value greater than 0.1 or so is between scores and family income. The NYTimes has this:
Everyone, of course, dismisses that income-scores correlation as silly, saying "You can't just give the family more money and raise the scores, ha, ha." That's right, but I don't think that money causes good scores, although there is a lot to be said for SAT prep courses, which are really a total review of math and basic grammar and that's good no matter what.

I think the problem is an incorrect correlation, a confounding factor. It's not that A (money) causes B (scores). It's that C causes D which affects B. Simultaneously D causes A.

Parents are the confounding factor.

As I see it, smart parents are likely to have smart children. More importantly, motivated, dedicated, educated and intelligent parents are likely to have M, D, E, and I children. They are also more likely to have money because they work harder for it and are more capable of holding the job and moving up the ladder to the big bucks and higher family income. Simultaneously, their M,D,I, and E children are more likely to have better scores.

Children whose parents were poor because of misfortune or some other external factor don't tend to fit this pattern. They're the Horatio Algers of the world, the kid who worked his way up from the mailroom to CEO. There's no reason to assume the black or Hispanic kid can't be doctor, lawyer, etc. In fact, the kids of those successful (and higher income) blacks and Hispanics are likewise high-scoring and successful. The single mother who instills dedication, motivation, and a healthy respect for education into her kids might not have money but her kids will.

Children whose parents were not motivated, etc., always fit the income-score correlation. Race has rarely been a factor in determining motivation and dedication (simply look at KIPP schools to see that), but it has been an indicator of income due to longstanding segregation and migration patterns. Money doesn't seem to drive scores, but it does correlate.

The constant desire for improvement and reform and reform and change and reform again is doomed to repeat its cycle of failure.

Are reforms necessary? Sure, if you can point to a definitive improvement that will result. I'd appreciate it if you'd define "improvement," first. Improve what and by how much? Student satisfaction, tech-toys, test scores, athletic titles, graduation rates, future wages? Fundamentally, is "success" as defined by "average scores increasing yearly" or "adequate yearly progress" possible? I don't believe so.

I think we should stop looking for the perfect reform because it doesn't exist. We should instead focus our attention on doing the best we can each year with those we have in front of us.

Just sayin'.

Thursday, July 11, 2013

Question about questions, SBAC edition. Graphs.

So, the SBAC is releasing questions so that we teachers can't complain that they've done all this work behind our backs. At least, they hope we don't.

But I digress.

I got a list of some released questions and started looking. My general rule is to make sure that the first thing on any handout is correct. I also have to assume that SBAC has access to a graphing calculator such as Graph 4.4 ... so why do they come out with the following?


The question they ask the students: "The graph of y = x² is shown on the grid. Drag the graph to show y = (x - 4)² + 2"

The question I have for them is, "Why didn't you graph it properly? It was probably more difficult to get it wrong than it would be to get it right. Did they use Microsoft Word? Just seems weird to me, like a circular arc and then two lines. I know it's picky, but sheesh.



I will point out that this is a great example of the over-reliance on gee-whizardry by SBAC ... everything has to be drag and drop, click and move, glitenbullshit. This could be done so many other ways, just as relevant and equally valid.

I'm curious mostly about the granularity of the placement. What is the tolerance? Can a student with a Chromebook and touchpad do this in a timely fashion?  Here's the "answer" ... note that it doesn't actually have a vertex at (4,2).



I think we'd better get a few extra mice for testing day.

Here's another. Drag the factors to make the equation ... I guess you drag (x-2) out twice. God, what a pain in the ass without a mouse. Graph looks wrong again. It's definitely a Bezier curve from MS Word.



Here's the real one for those who care:


Again, why not do it right?

Saturday, June 16, 2012

How to Create A Monopoly

Laying some groundwork: The Common Core Initiative has been gaining steam in the U.S. for some time.  They've been adopted in 45 of the 50 states. The CCI is fine as it now stands. One could argue with some of the standards or with some of the timelines, but overall they are remarkably similar to most states standards.

Phase Two:  Someone has to test these standards. For 26 states, that would be Smarter Balanced, which is pushing hard on computer-adaptive testing. This is testing that is given on computers or tablet, and using a difficulty rating and some math, delivers questions based on how correctly the student answered previous questions ... answer correctly, it gives you a harder question, answer incorrectly and it tones it down a bit.  Add that to a massive databank of questions (all rated and sorted) and you get a perfect assessment of that students and his abilities.

Like this one.
Every student will have to have access to a desktop running a browser and some lock-down software or to a tablet running a specialty app ... and strangely, the tablet are going to be required for the 11th grade math. N.B.: I am not sure on this last but the presenter (Sue Gendron) said that schools would be required to have at least 25% of the students tested by tablet and then mentioned math questions, so I think that's right.

If all goes well. If they finish in time. If they do their work well. If they finish at all. If the questions are rated correctly. If they're testing the right things.

Here's where the market manipulation comes in.

They are NOT DONE YET; in fact, they've barely started and they probably won't be done completely in time for the rollout in 2014.

So which tablet operating system do you think they starting with, Android or iPad?

With a lot of development money donated by Apple in the form of thousands of iPad IIs that are only costing the states $240 each (per Sue Gendron, ex-Commissioner of Education, Maine), is it surprising that Smarter Balanced is not going to get to the Android app until much later?

Slick, right?

Think about how school buy tech.
iPad money tree.
  • Once the IT buys iPads for the first round, will they really want to introduce a second operating system and purchase the more expensive Androids? No.
  • What if the Android makers bring the price down to a competitive level? Too late. The schools and the teachers and the states are all in the Apple pipeline - that's really hard to break.
  • Isn't the iPad priced below market value to the point of anti-competitive pricing? Yup.
  • Isn't that illegal? Yup.
  • Are states pushing this? Yes, this monopolizing is being done at the behest of the Government.
    • Republican or Democrat, they're both doing this.  Don't get started with that. 
Ramifications?
    •  Doesn't iPad have to be linked to iTunes, and come with a ton of EULA restrictions and bullshit that so many people hate about Apple and all of it's products? You Betcha.
    • Every school in all the 26 states, buying iPads ... as soon as the schools get hooked, the price goes up to it's normal point of three times what you're paying now and twice what a normal-priced Android tablet would cost.And since there aren't any schools with Androids now, who wants to bother developing the app for them? 
    But that's okay, isn't that the only ...
    • Don't you want to use those shiny new tablets for eTextbooks? Yes, purchaseable ONLY through iTunes, naturally. 40%, please.
    • Don't you want the kids to read on those shiny new iPads? Yes, but all books are 40% to Apple.
    • If the kids write something good on an iPad, they are required by the EULA to sell it through iTunes and Apple gets a cut.
    Congratulations, you've just witnessed the beginning of a monopoly.

    Updated for those with 21st Century Learning Skills

    Tuesday, June 12, 2012

    Common core and computer adaptive testing.

    The common core update happened today at our school, with the kindly old lady telling us how much the scores are going to rise when we raise the reading level and the difficulty level of all the questions. (Yeah, I know, but let's play along.)

    She described the computer adaptive questions and how the program would be able to deliver a different next question based on whether the student got the previous one right.  I'm good with that concept, actually.  With a suitable scoring system (like the one for Olympic divers), it should work out quite well. 

    Then we got to the "performance tasks" which were much more extensive, taking 2 hours each for high school students to complete. The students will be responding with text, handwritten (onto a tablet - the goal is for 25% of tests to be taken on tablet), voice over documents, video, and anything else the test creators could dream up.

    Here's the kicker.

    None of this is past alpha development stage. The iPad apps are still in development. The Android apps, laptop programs, desktop programs, Mac, PC, ... all were still in development. Everything is computer-based but they can't figure out how to block the student from getting on Google. They don't know how the schools will supply themselves with all the tech ... "When I was Commissioner in Maine, I just put it in the budget. We're talking to your state about it."

    The math framework and examples will be ready ... in a few months.  The digital clearinghouse ... maybe by June 2013. The funding for technology ... "we're working on that".  The assessment engine ... that'll be ready soon.  The suitable scoring system I mentioned earlier ... "I'm not sure exactly how that'll work".

    I point out that the questions they're displaying are problematic, but her response was to remind me that the exemplars had been posted for some time and teachers had been able to make comments. Oh, my bad.

    Not inspiring confidence, here.

    She put up a reading example about a science topic that was similarly flawed - if you knew the science already, you'd have a huge advantage and your "reading score" would be much higher. Fair enough, you might say, but the elementary teachers in the auditorium didn't know most of the words either.  "Turbidity", for example.

    8th grade - they'll guess and check.
    I'm not sure how that's supposed to
    measure algebra, though.
    Scoring will be a pain, too.  Questions like this one will be scored by ... someone. This will be a source for error right there.  Portfolios had this problem, too.  You'd think it would be easy to get all the scorers to come up with a consistent grade since they all had a page-long rubric to follow.  You might think that, but you'd be wrong.

    I love education.

    Friday, February 24, 2012

    NY's VAM is a Club that Gates Disapproves of.

    Developing a systematic way to help teachers get better is the most powerful idea in education today. The surest way to weaken it is to twist it into a capricious exercise in public shaming. - Bill Gates in the NYT
    It's even worse when the system in use is riddled with errors, has random fluctuations in scores that can make or break a teacher's reputation, doesn't have the support of those being evaluated and makes tenuous  associations between scores and those responsible for them.

    How the scores a student receives on a meaningless test can be much of an indication of the worth of a teacher who didn't take the test, didn't have that kid in class except for part of the current year, and rarely has much control over that student and his personal and academic life, is a mystery to me and a source of much bemusement.

    My state doesn't have VAM yet but it does publish the NECAP scores and tries to shame schools into improvement. Our difficulty up here lies in the fact that these tests are taken in the 7th and 8th grade and then in the beginning of the 11th grade. That's it. This year's juniors took the test a month into the year and I had never had them in class before - what kind of measurement system is that and who is being measured, really?

    It's interesting that the public sees numbers such as "64% of the students are not proficient" and still votes for our budgets every year at town meeting.

    A quick note for all of you big city folk used to 3.5 million-student city-wide districts and mayoral control: Vermont districts are generally a few small towns banding together with schools ranging in size from 200 to 2000 students, heavily skewed towards the small end. Each of these districts has a school board for the district and one for each elementary school. At Town Meeting Day (yes, we still do that), each town votes on it's school budget and town budget, elects school board members, and decides other weighty issues including whether to pay $460 to the ambulance service.

    Monday, February 20, 2012

    Value-added measures don't measure up for Evaluation


    Value-Added Measures don't make a good foundation for a teacher  evaluation system.

    A comment on a Joanne Jacobs article:
    VAM measures the amount of improvement your students make. There are a number of ways to do this. Some of the early VAM methods were highly unstable. More sophisticated methods do seem to hold up well from year to year and also correlate with positive long term outcomes such as lower teen pregnancy rates and better education and employment as adults.

    And here I thought I was supposed to be teaching math.

    I don’t believe VA is anything on which to base bonus or termination. “Seem to correlate” does not mean “cause” … and that’s for the best measurements.

    What of all the poor ones? “Some of the early VAM methods were highly unstable” ("Unstable" is a charitable term for "Any resemblance to a consistent reality is neither implied nor intended.") 

    It means that the results cannot be trusted for grading the student who took them (that's stated plainly and explicitly in the administrator's notes) and it means that the test are worse at evaluating the teacher who didn't take them.

    There are many issues with any kind of testing. What exactly are we supposed to be teaching and what results do we want out of it? What will we consider to be a success? Do the tests measure what we think they're measuring and does that result resemble the state of the student?

    I am given a curriculum that I am to follow. The test is written for a different curriculum. Don't judge me based on something you tell me not to use.

    Then, there's accuracy and repeatability. Use a ruler and get the same height every time - that's data you can trust. Give a test to students a second time, they would score differently. Give the same essay to ten scorers and you'll get 10 different scores. Read Making the Grade for a nasty dose of testing realism. Since the scoring of these tests is so "unstable", evaluations shouldn't be based on the results.

    This graphic to the right cleverly pretends that measuring a child's height is exactly analogous to measuring his grade level.  Unfortunately, the accuracy possible in the one is not possible in the other. I would note with some amusement that the books he's standing on make even that height measurement into an exercise in systematic error.

    Then, the tests claim to be able to discern between fractions of a grade level but the random error in such a measurement is a full grade level or more. The test to test changes on one of the best-known measurement systems, the SAT, can be as high as 100 points. They don't report scores, they report a range (520-540). The test is 600 points and the variation is 100 points. Now imagine the variations on your typical state test.

    States routinely tell the testing company to instruct the scorers that averages HAD to be in a certain range - any test scoring that ran counter to that pre-determined result was wrong. As Todd Farley describes it, accuracy is a fantasy.

    It doesn't make sense to evaluate me based on a test given to a fifteen year-old kid who has only had me for a short while, who has failed again and again, who has attendance "issues", who's strung out on something ("self-medicated"), using a test that pretends to accuracy but fails miserably at it and rarely is aligned to the same curriculum that I've been required to follow.

    What about Value-Added?
    Just the basic premise that you can differentiate teachers based on VAM is flawed. If I have a group of students that improves a lot this year but a different group that doesn’t do as well next year, are we to assume that I’ve been slacking off and just need a goad, a little taste of the whip to perform better or should we assume that my teaching is so variable that I can be bad, then great, then merely good?
    If my students improve from a Level Equivalency of grade 4 to grade 8 in one year (even though no test can honestly make that claim in any accurate way) and my colleague raises his students from grade 10.2 to 11.3, which of us has done a better job? I may have convinced them to work harder at the end of the year but not actually done much teaching.

    If I have a class with “issues” and they only improve from 9.5 to 9.8, that might be a tremendous leap for them but it wouldn’t show that way to the outside observer.

    I find it troubling that we have this blind trust in a standardized testing program.

    What are they good for?

    VA measures are useful to me in a classroom, provided I get them in a reasonable amount of time, disaggregated so I know detail instead of a vague "You Suck" or "You're Great", and the high-stakes are left off it.

    Selling newspapers is not a good use.


    Income gap and Academic Gap is Linked. Well, duh.

    Scott McLeod's Mind Dump: " The test score gap between the richest 10 percent and poorest 10 percent of students has grown by about 40 percent since the 1960s, according to a study by Stanford University sociologist Sean F. Reardon. That's twice the testing gap between blacks and whites, which shrunk significantly in all income levels, he said."

    Which makes sense because the income gap between the richest 10 percent and poorest 10 percent has grown since then, too.

    If only someone could figure out why. I have my thoughts, and thoughts, and thoughts, but there's no hard evidence for the mechanism. We just know that income correlates to scores really well, a direct correlation of (.95).

    Update: Dan Pink coincidentally chimes in here, too: How to Predict a Child's SAT Scores. Look at the parents tax return.

    Monday, February 13, 2012

    Raising the Stakes Causes Inflation

    Joanne reports:

    New York is looking into charges that credit recovery programs make it too easy for students to blow off schoolwork, earn credits for doing very little and pick up a diploma. Principals are evaluated based on graduation rates, providing an incentive to lower standards. (Students can earn P.E. credits online.) Read teachers’ comments on Gotham Schools. It’s not just a New York City thing. Teachers all over the country have been complaining about credit recovery.
    It's a simple thing: when you threaten someone's job, they will react defensively. Sometimes you like the result and sometimes you won't.

    I can't say I'm surprised to read that "They described how principals used credit recovery to boost their schools’ statistics and how students opted for it as an easier way to collect credits." What's a Highly Ineffective Principal to do?

    "Principals are evaluated based on graduation rates, providing an incentive to lower standards."

    I shouldn't even toss this into the HIPster line of posts because this is too depressing and I can't really laugh about it. 
    Every time we raise the testing stakes, more cheating will result. While I expect teachers to be more honorable than bankers and hedge-fund operators, I assume they are subject to similar inducements and pressures. Just because teachers are not too big to fail doesn't make them immune from self-interest. Especially as no student suffers from getting a better score than he otherwise might have gotten. - Deborah Meier
    As you raise the stakes ...

    Sunday, February 12, 2012

    Raising standards will do what, again?

    Here's the bad news:
    This is a four-state test given to all juniors in RI, ME, NH, and Vermont. 6,000 in Vermont. The 4 point scale is
    1. Not proficient
    2. Nearly Proficient
    3. Proficient
    4. Proficient with distinction
    Our 11th grade students do pretty well in English, not so much in math: 36% passed, i.e. got a 3 or 4. If we look into those numbers,though, a weird thing shows up. Only 3% got the top score. Only 33% of the students got more than half of the questions correct for a passing score.

    That's right. 3% got the high score (2/3 correct or better, appr. 70%). (edited 2/13: 800) 180 kids out of 6000, spread out across the state. Three year totals of 600 out of 20,000. The passing score is roughly 50% and only a third of the kids got that.

    Either every single math teacher in Vermont is screwing up and not doing their jobs or we have a test that isn't appropriate. (and after missing that number in the previous paragraph, maybe it's me.)

    If it were just one or two schools, or a single county, you might have a point to make about teachers or demographics but not if the problem is statewide ... and the numbers for Rhode Island, New Hampshire, and Maine are exactly in line with Vermont's.

    If I were to write that test with that kind of passing rate, I'd be excoriated for making it too difficult. Instead, our Commissioner says that we should make the test harder:

    "Commissioner Vilaseca stated: "We are gathering more information about what Math courses all students are required to take, and will carefully consider whether it is time for Vermont to increase our graduation requirements in mathematics."

    Uh, dude, they don't pass the test now. What exactly do you figure raising the cutoff will do?

    Wednesday, February 8, 2012

    Thinking about exams.

    Real Teaching Means Real Learning wants to
    Re-evaluate exam week
    , something that he feels he hasn't thought about deeply enough. After he bemoans the fact that students aren't collaborating on these exams, he throws up a couple of strawman arguments and takes some huge liberties with student motivation.

    "If a kid can fail the exam and still pass, should he take the exam?" Apparently not, which seems like a crazy idea to me. It's been my experience that kids often don't put their knowledge all together in one coherent package until they sit for the exam. "I don't know what I am doing" becomes "Oh, that was easier than I thought. Now I get it." Additionally, how many kids want to only get a 60% passing grade if they could get a 80% ?

    Further, he seems astonished that a kid with an average of 28% can't pass the class and is required to sit the exam anyway. I look at it this way: most kids who have a mark in the twenties are sabotaging themselves; they're rarely stupid. If they take the exam and score well, they have that to build from the next time they take the course - not everyone can do it in the same allotted time.

    You can even make the case that scoring well on the exam is the ONLY criteria for passing a course -- if the exam is comprehensive. It takes a fairly complete knowledge to bring together all of that knowledge in one place, for one exam, during two hours.

    I have never used one.
    Seems like more trouble than it's worth.
    Up pops strawman #2. "As most exams are multiple choice" the student can guess and get a 25%. It's been a while since I gave a MC exam. Tests and quizzes? Yes, there's almost always a MC section. Final exam? No.

    RTMRL wants to take some of the kids out of exams and give them intensive one-on-one tutoring instead. He's ignoring the reality that they weren't working for the previous six months; expecting them to suddenly catch up on all of that during the time the rest are taking exams is a little foolish.

    "Do we take these students, who are obviously struggling with the course, and test them again or do we teach them? Imagine the learning which could occur with 1-1 help in a 3 hour block with a weak student?" Um, probably not much.

    Up pops strawman #3, "If the answer is the latter, then I suggest you be upfront with your stakeholders and put a sign outside your school saying “For an entire month, two weeks during each semester, your child will not learn at this school”.

    He is pretending that the kids are losing 2 weeks at a time for exams. I have been working in education, public and private, for nearly 30 years. Not one school has ever taken this much time for exams. Not only that, there is the idea that learning can't happen at exam time.

    Enough of that. Are exams worth it? I feel they are, for all students.

    It is a summary evaluation. For the entire course, we've broken things down and scaffolded them up. We've looked at pieces of the course in isolation. Projects have been deliberately limited in scope and the students are analyzing and committing to long-term memory the nuts and bolts as well as the grand themes.

    The final exam is the time when you can make problems that bring the whole course together, when the student needs to demonstrate ability and knowledge of a whole vast subject, when the notebook is useless and memory is critical, when a single question requires a deep understanding of parts of 10 or more months of work.

    Some kids crash and burn and still learn that lesson, and come out of the fire tested and ready to do better the next time. Some students realize how much they really do know (and surprise the hell out of themselves). The rest consolidate their knowledge and prove to themselves how much they understand -- or don't.

    You can't fake that.

    Monday, September 5, 2011

    Finland, Teachers and the Common Core


    Joanne quotes A case against standards
    by P. L. Thomas (originally posted in The Answer Sheet, WaPo)
    Since we often choose to demonize U.S. education by international comparison, I suggest we consider the attitude toward the professionalism of teachers in Finland from Henna Virkkunen, Finland’s minister of education:

    “Teachers in Finland can choose their own teaching methods and materials. They are experts of their own work, and they test their own pupils.
    Hold on to your horses, PL., you're mixing apples and orange Volkswagens.

    Teachers in the United States can take four years of college drinking, masquerading as a student of history, attend a 6 week TFA summer course, and be considered the Savior of the Universe -- and some even become decent first-year teachers. Or you can take a couple education classes, student teach with a "master teacher" who ignores you for a couple months and be certified. You can fail out of every other major in college and slide down the slippery slope to elementary education - at which point, you can get As. Or you can participate in Troops to Teachers, which assumes that military training and the ability to order people around is an automatic guarantee of the ability to teach.

    None of these methods is a guaranteed loser - some sergeants are really good at teaching twelve-year-olds, some elementary ed students came from the top of their class, and some TFAs last more than a year and really can teach.

    The problem is that most of these teachers think of themselves as great but are the last people you want to be "autonomous" and they're certainly not masters of their own work. (NSFS - Not Safe for Sanity)

    Finnish teachers, on the other hand, come through a highly selective process, spend years as a teacher-in-training under the direct observation of a master teacher, and generally are all from the top tier of graduating college students. They are all using roughly the same methods because they've been through a lot more training than US teachers and seen what works.

    Many US teachers do get to toss cards in the air and arbitrarily decide that they'll teach in a constructivist style with no books and little direction or lecture. Finnish teachers, by and large, stick to what we denigrate as "traditional" methods. Any that deviate from those "traditional" methods have been around for a long time, actually are masters of their own field, and can be trusted to test their students.

    US teachers are under tremendous pressure to raise grades on tests (to the extent of 25% pay raise or cut, in some places) and are constantly being second-guessed by everyone and anyone while Finns are not. This over-whelming desire to improve and change means that US teachers are constantly swinging between great and lousy at the whims of the most recent fad.

    Fad or Innovation? Wait until your pet innovation fails because the students don't like the "inverted" classroom because of the extra work. (Just to pick one current idea out of the fad-hat.) See how long you are able to continue.

    Ideas fail with students all the time but might work with a different group or a different teacher.  Just because ThinkThankThunk can pull off a non-traditional format doesn't mean that everyone else can. Similarly, dy/dan has great ideas, most well worth stealing, but you can't just blindly copy a few WCYDWTs and hope it's a course - there is the other 90% of one of his algebra courses still out there and the always pesky issue of making his curriculum fit your kids instead of his. Robert Talbot successfully "inverted" his college classroom but had mixed reviews for linear algebra and loved it for programming.
    K-12 teachers and teacher educators must be afforded the autonomy of professionals--not further bureaucracy and invalid accountability.
    Autonomy? Not until you've proven yourself. I don't trust the first-year teacher and I certainly don't trust those who've already demonstrated their incompetence. "Professional" doesn't apply to everyone.

    Thursday, July 7, 2011

    Don't teach to the Tests

    For any new teachers out there, here's a few words of advice.

    • Teach your students math. The tests will take care of themselves.
    • Use released test questions in your own tests occasionally. Some of them are quite good.
    • The book is not your enemy. It's not your script either.
    • Remember the way you learned. It seemed to have worked.
    • "Don't smile before Christmas" is stupid.
    • "Don't threaten what you won't enforce." is critical.
    and for Christ's sake, don't cheat on the standardized testing. It is so easy to detect and you'll accomplish nothing.

    Sunday, February 6, 2011

    Testing, Scoring and Trusting the Data

    A fascinating case study is the NY Regents, scoring, and the unintended (or purposeful) consequences of a line in the instructions to the scorers.

    First, here's the graph of the number of kids getting each score (from WSJ).  The issue can be seen quite clearly. The graph overall is a typical left-skewed distribution, as you'd expect from this type of test. The trouble comes when you explore that jump in the middle at the passing mark.

    Of course, the nattering class is all up in arms over this, claiming fraud and misconduct.  You can almost hear jeers of "Union Bastards trying to save their jobs by lying on the tests."  Unfortunately for those people, the reason comes down to one sentence:
    "[State officials] note that the state actually requires teachers to regrade certain Regents tests where the student barely fails in order to check for grading errors."
    When you are singling out tests for special consideration, and the stated focus is to look for "scoring errors" on "barely failing tests", the only way for the scores to change is up.  Since it's easiest to give one or two points to a 64 or 63, it's logical that those would be the most effected.


    Looking at the graphs, you can see that some teachers (probably a school at a time) set their cut off at 50 while the majority set the cutoff at 55.

    Why the negative slope in that region? When looking for ambiguous answers which could be scored higher, you have to "rescore" each problem until you get an appropriate number of points. It's easier to find one such than 10 such.

    Similarly, since there was no reason to check more answers than just enough to get the kid to pass, the uptick was only to 65, though it seems that many teachers weren't keeping close track of the extra points and brought the kid up to a 66 or 67.

    Conversely, if the teachers had been instructed to rescore all students within ten points of the cutoff instead of just those who were below it, then you would have seen some students adjusted downward, balancing out much of the upward movement.  If there were a similar uptick at the 65 mark in this hypothetical, then and only then can you claim that teachers are deliberately mis-scoring to jigger their VA measures. (and even then, I'd put the reason as teachers wanting to help students rather than being so coldly self-interested.)

    The place where I found the link to this article had this comment: "Teachers don't want to flunk kids that just barely miss the passing score. Until all responsibility for creating and scoring state exams is given to an independent body with no interest in the results of the tests, the results reported should be viewed skeptically."

    Ummm, no. Scoring tests is really complicated.  Pearson, the biggest company, uses part-time, barely out of college, minimum wage people to do the scoring. Getting the "right" score is more a matter of whether or not the scorer speaks English and actually knows the material.

    Frankly, given the mess that the testing industry is in when it comes to scoring, I have a feeling that the teachers are doing a more conscientious job. If you want a nasty introduction to the follies of testing company scoring sessions, check out Todd Farley's "Making the Grades." It's well-written but damn depressing if counting on accurate test scores because you're stuck in the hell of value-added and merit pay.

    Thursday, January 20, 2011

    PISA scores - another look.

    What do you see? When you look solely at schools with fewer than 10% of students on FRL (i.e., poor), US schools would be better than those of the top of the chart, Korea. When you include the whole spectrum of US schools and their students, the US is much lower.

    The US has a much bigger spread than any other country (the largest standard deviation of wealth in the developed world).
    Our overall scores are unspectacular because we have a high percentage of children living in poverty, over 20%. This is the highest among all industrialized countries. In contrast, child poverty in high-scoring Finland is less than 4%. For Cleveland and for the US as a whole, the major problem is poverty. Before we worry about teacher quality, institute longer school days, and increase testing, we need to make sure that all children are protected from the effects of poverty: This means adequate health care and nutrition, and access to books. When we do this, American test scores will be at the top of the world.-- Stephen Krashen

    Saturday, December 25, 2010

    That'll improve the school

    Think they'll pass muster?
    Forget about replacing the teachers ... replace the students. Bring in the superstars, the superheroes and the supernerds.Then our scores are sure to go up.  New Trier High School, Fairfax High School -- we're looking at you. Let's trade yours for ours!

    Never mind.

    Funny idea, though.