Louis C.K. hates the Common Core standards.
I first saw it here, on the HuffPo, from David Letterman. Later, a friend/colleague (I don't name her only because she is in the business, too, and I don't want to get her in trouble) sent me a message to make sure I saw it--thanks for keeping me in the loop.
Toward the interest of full disclosure, I have to say I’ve always had a little crush on Louis C.K.
Why, you may ask?
You may ask this because you’re thinking of the fictional-but-based-on-real-Louis and the elevator fantasy scene or fictional-but-based-on-real-Louis passed out and surrounded by empty pizza boxes and ice cream containers or fictional-but-based-on-real-Louis being rejected by a woman because she witnessed him shrinking from confrontation with a high school bully.
Why do I have a crush on Louis C.K.?
Oh, let me count the ways: for the pure and sheer humanity and vulnerability that he just goes ahead and expresses, seemingly without any filter, and most of all, for the breathtaking courage it must take for him to expose his humanity and vulnerability to the world. He’s willing to be naked, figuratively and literally, when most of us are frantically swaddling ourselves with ego padding, trying to keep our humanity and vulnerability zipped up, buckled tight, under wraps, armored up. We post only flattering glamorous pictures online, nothing that makes us look dumpy or frumpy or dorky, even though surely all of us spend more time being dumpy, frumpy, or dorky than we do being smooth and suave and glamorous and elegant and unruffled.
Unless we’re, you know, Kimye.
(That picture on my blog? Taken three years ago. I've aged. I hate having pictures taken of myself and probably won't update it until I'm seventy.)
Besides, I hail from the working class, as does Louis C.K.,and so I applaud and cheer preach it, brother! whenever he criticizes entitlement or laziness or ingratitude.
All right, so we have this comedian who is a father--a good father, if by “good,” we mean someone who engages in thought about parenting and participates in his kids’ lives, which is all fantastic, and those of us who didn’t have fathers like that think he is really amazing for being that kind of father, and probably those of us who did have fathers like that feel a bit of fondness for him because this is familiar territory--who pays attention to his daughters and their inner lives and who worries when his daughters suffer, and so when his daughters, upon encountering mandatory statewide standardized testing, feel anxious, Louis C.K. has something to say about it. Something really not flattering to the people who write the tests. Something really not flattering to me.
In the spirit of respectful discourse and intellectual debate, I’d like to address these points:
1. Who writes these tests?
I do. Not the bad ones--unless instructed by a client to write badly, and sometimes that do happen, much to my chagrin--and only English language arts. Someone else is to blame for math, science, and social studies. Not my areas.
You could talk to my friend Scott about math or my colleague Jim about science, I guess, but they don't write bad tests, either.
There are lots of bad tests, yes. It's a systemic problem. More on that.
2. Why do I do this horrible, horrible thing?
To earn a living and support my two children.
3. What are my qualifications?
I have a bachelor’s degree in literature from the University of California at Santa Barbara, a master’s degree in English with an emphasis on writing from Sonoma State University, almost two years community college teaching experience, twenty years experience in educational publishing, and five years elementary classroom volunteer experience, as well as other miscellaneous tutoring experience from college and grad school. Also I was a TA
in grad school, for the creative writing class, oh God, was that awful, all those stories about cats and sexual abuse and suicide mixed with the occasional fantasy of being a wealthy celebrity writer driving a red Corvette, clearly no
student in that class had ever read a word written by any writer other than their favorite writers: themselves. In addition, I’ve put in many, many hours of study--in education, in reading and language acquisition and of literature and literary criticism, and especially in assessment, and even more especially in the writing of test
questions.
4. What’s with the Common Core?
It’s a good idea to have national standards. Other countries do, and that’s how they make sure that all the kids in the country are learning the same things at the same pacing. It’s a good idea to consider career- and college readiness, and how to make that happen, particularly when kids in the United States are undereducated to a degree that must make us the laughingstock of industrialized nations. Finland and South Korea especially must snigger at our national ignorance and celebration thereof--is there any country in the world that makes a point of being so dang proud of being stupid? I ask you.
5. Why do people hate the Common Core so very much?
People fear change. People hate what they fear.
No one understands what the Common Core standards are, or what the shift means, or that it's really a good thing that kids in Alabama learn the same things as kids in Connecticut. There’s too much hype and not enough real education about the standards and their purpose. Teachers are scared because tests are being used for wrongful purposes (never a good idea to link teacher pay to test scores), and scared teachers are scaring the kids.
6. What is Louis C.K. really upset about?
Like any caring parent, he’s upset that his daughters are upset.
He doesn’t know enough about the Common Core to be upset about them.
That’s not his fault; it’s the fault of the top-secret test publishing industry that keeps all information under lock and key, supposedly to preserve confidentiality, but, really? Wouldn’t it be smarter to explain what’s happening and why? No. Because then they would have to explain everything else, like the billions of dollars spent on testing and how little of it changes anything really, and also how little of it trickles down to the people who are doing the
actual work which means that the majority of content developers (not me, I'm the exception, this is my career) are inexperienced hobbyists or inexperienced part-time teachers or hustlers who think they're getting away with something by getting paid to do something they don't know anything about and how much of the big money in testing gets bottlenecked up at the executive and shareholder level.
7. What should Louis C.K. really be upset about?
Capitalism. The one percent. The war on poor people instead of a war on poverty. The state of education in the United States. Dogs that need rescue at animal shelters.
What’s so unfortunate here is that Louis C.K. is someone who’s got a public forum--people (including me) listen to him, laugh at his jokes, care about his opinions. He has an opportunity to make people think (at least a little) and that would be a really great thing if--and I don’t at all intend this as a snarky sarcastic dig--he knew what he was talking about. I'm sure there are a gazillion things he knows plenty about, but the Common Core standards are not on that list. Really, do you think he has even read them? I mean no offense, but I would be surprised by an affirmative.
There are so many things that are terribly wrong in education in general and in educational
assessment in particular, but from my perspective--as someone who does know what she’s talking about here--the Common Core is a paper dragon. Let’s talk instead about the corporatization of education.
How about that fewer than half a dozen test publishing companies rule assessment, and the king is Pearson? (Which company is now involved in a controversy over the award of the PARCC assessments as the result of a lawsuit filed by AIR.)
How about that the people who are actually doing the work are paid woefully inadequately (Hello? My yearly income today is the same as it was twelve years ago when I started my business, but guess what, inflation--can you see why this is a problem?) while the companies continue to earn profits that are obscene in comparison?
Let’s take me, partly because I am monumentally self-absorbed, but also because my experience is what I know. I do know many other people in this line of work, but very few have the depth and breadth of experience that I have in educational assessment: I’ve worked in hand-scoring, program management, content development from the ground up (item writer to editor to supervisor to manager to director and back to item writer and editor). I’ve worked directly with state department of education officials. For seven years, I had an annual contract with Miami-Dade County Public Schools, one of the largest districts in the country, a district that has more students than some states.
For the last twelve years, I have concentrated mainly (with some side jobs involving higher level consulting of test design and product research) on the hands-on work of content development: writing and editing material (reading passages and questions) for tests. That is unheard of. In this industry, as soon as anyone shows a spark of initiative, and especially if that initiative is accompanied by a pebble of intelligence, that person gets promoted. Anyone else with twenty years' experience has been in management for at least ten of those years, and management is not the same as actually doing the work, as any line cook at KFC could tell you.
I mean no arrogance when I say that I’m the perfect person to write tests, considering the combination of education, experience, and
dedication--because I care about what I do, quality matters to me, the kids matter to me--when I write reading passages and test questions, I’m thinking about the experience of the kids who are going to take the test just as much as I’m thinking about my paycheck. Maybe more.
Not that I don’t think about my paycheck. I do. I have to. I’m a single mother with two kids.
Thinking about those kids who take these tests breaks my heart. Not so much kids like my daughters and Louis C.K.’s daughters--these girls are all going to be fine. They have parents who love them, ready access to books, music, art, libraries, documentaries on penguins and whales and volcanoes and subscriptions to the National Geographic and visits to the Smithsonian. Their parents talk to them all the time (maybe too much, in my case; Louis C.K. is probably a lot more interesting and a lot less pedantic when he talks to his daughters) and are willing to listen and answer questions and explain all about why everything in the world is the way it is. We the parents will support our daughters, consider their happiness, find ways to challenge them, look for opportunities to help them navigate the complexities of relationships, communication, education, and, eventually, careers.
And Louis C.K.’s kids? They’re especially going to be fine. They’re rich. They’ll have their pick of colleges, go wherever they want, do whatever
they want from now until they die and leave their truckloads of dollar bills (remember investment income is taxed at about half the rate of labor income, so their money is constantly making money, they'll have more money than they could ever spend) to their kids and their kids' kids.
That is awesome for them, and while I envy their good fortune (which I acknowledge comes from the hard work and talent of their father), I don’t begrudge it them. If they get a little upset about a
test, I understand and I sympathize and it's nice that their dad sympathizes, too, but really, there are a lot worse things in the world to happen when you’re a kid, and a lot worse things do happen to many of the kids in the world. Maybe some of that righteous indignation could go to someone else’s kids, kids who really don’t ever get a chance.
Not that I mean to be all sassy to Louis C.K.
Note: A little crush. Not a stalker crush. Have I ever written to or tried to contact him in any way? No. Geez. Of course not. Would I ever? No. Oh, God, no. What do you take me for? Have I watched his show and stand-up routines? Yes. Do I laugh at his jokes? Most. Some of the humor is a bit past my endurance, but I celebrate his right to express himself.
What I'm thinking about what I do, which includes but is not limited to all matters related to education, reading, writing, editing, content development, and curriculum and assessment design and implementation
Showing posts with label test publishing. Show all posts
Showing posts with label test publishing. Show all posts
Thursday, May 8, 2014
Friday, July 26, 2013
A Great Deal Done Imperfectly
Better a little which is well done, than a great deal imperfectly.--Plato
In my last post, I may have given the impression that editors have more power than they do, and that perhaps that they have time to consider the ramifications of failing to provide all the necessities, or that they are willfully negligent. Salaried as they may be, editors often find themselves in an unenviable pickle, with compressed development cycles and few resources. The industry's reliance on freelance personnel increases the workload of front-line staff, who may now have to manage groups of writers in addition to performing other duties. Each writer must add about an hour a week in emails, phone calls, and admin tasks--and that's if the writer is low-maintenance.
It's also likely that editors want to provide all the necessities, but those necessities don't exist and the schedule doesn't allow time for editors to develop them. (Some of the most experienced item writers are able to work around the deficiencies, but the work of the less experienced will be affected.)
No one--except the one at the top of the pyramid, I imagine-- is resting on a velvet cushion.
I may have left another inaccurate impression: that it's all about the money. It's not. How can it be? This is not a high rolling game. What I mean to say is that when writers don't have what they need to do their best work, everyone loses.
The industry continues to become less hospitable to the people actually doing the work of creating the tests--or, more accurately, the people writing the passages and questions from which the tests are assembled--which results in a great deal done imperfectly.
Writers lose time and money; they also lose the best of all rewards, the satisfaction of a job well done, simply because how can you do a task perfectly when the task hasn't been clearly defined, and when you ask for clarification, you're directed to figure it out?
The companies lose much, much more. The lower pay and the more pain (inconvenience? Call it what you will. I mean all of those tiny ducks that are pecking us to death) to the writers, the lower quality the work, and the fewer writers willing to undertake that work, those fewer writers being the ones who have no choice: the least proficient, the least experienced. And the most highly skilled writers simply decide they've had enough and they move on to greener (or at least different) pastures.
Most importantly, the children who are taking the tests have already lost when they're faced with low-quality materials that don't provide them with a fair chance to demonstrate what they know and can do.
All right. Let's move on. I'm eager to address the basic rules of item writing (a version of which you can see here, in the Quality Control Checklist published by CCSSO), but I realize I should first define some terms.
An item is a test question. An item may be discrete, or may depend on some external stimulus, such as a reading passage or a chart or a map or something else.
Here is a discrete item:
The above is a multiple-choice question, and contains a stem ("Why does my dog bark at mail carriers?") and four answer choices: one correct response (B, as far as I can tell, but I think maybe C is a possible right answer) and three distractors. Distractors, which used to be known as "foils," are wrong answers. Don't get hung up on the language--the point is never to distract nor entice the test-taker to bubble the wrong answer; the point is to create wrong answers that have a reasonable foundation in common mistakes kids would make with that particular skill or bit of content knowledge. More on this later. But tests should never be tricky.
A multiple-choice item is usually worth one score point, and used to be budgeted for one minute of test-taking time, not including the time it takes to read a passage or examine whatever stimuli is needed to answer the question.
There are other item formats: constructed-response items, which are also known as open-ended items. These require the student to provide a response. The response may be as short as a word or a phrase, or, in the case of extended-constructed-response items, the response may be a complete essay.
Here is a short constructed-response item:
An extended-constructed-response item would look like this:
Bear in mind that these sample items are jokes, and as such, aren't examples of exemplary items, primarily because they require a great deal of prior knowledge, and so the test-taker who is unfamiliar with Sophie and dogs in general will perform less well than the test-taker who is on a first-name basis with Sophie and/or other dogs. There are other, less egregious flaws, but we'll get to those when we get to them.
If you have an item you'd like me to examine, explain, or deconstruct, feel free to post it in the comments. Check the copyright first.
What I'm reading: Forgot to mention I was also finishing up The Claverings by Anthony Trollope. Then it's back to As I Lay Dying. I gave up on the other.
In my last post, I may have given the impression that editors have more power than they do, and that perhaps that they have time to consider the ramifications of failing to provide all the necessities, or that they are willfully negligent. Salaried as they may be, editors often find themselves in an unenviable pickle, with compressed development cycles and few resources. The industry's reliance on freelance personnel increases the workload of front-line staff, who may now have to manage groups of writers in addition to performing other duties. Each writer must add about an hour a week in emails, phone calls, and admin tasks--and that's if the writer is low-maintenance.
It's also likely that editors want to provide all the necessities, but those necessities don't exist and the schedule doesn't allow time for editors to develop them. (Some of the most experienced item writers are able to work around the deficiencies, but the work of the less experienced will be affected.)
No one--except the one at the top of the pyramid, I imagine-- is resting on a velvet cushion.
I may have left another inaccurate impression: that it's all about the money. It's not. How can it be? This is not a high rolling game. What I mean to say is that when writers don't have what they need to do their best work, everyone loses.
The industry continues to become less hospitable to the people actually doing the work of creating the tests--or, more accurately, the people writing the passages and questions from which the tests are assembled--which results in a great deal done imperfectly.
Writers lose time and money; they also lose the best of all rewards, the satisfaction of a job well done, simply because how can you do a task perfectly when the task hasn't been clearly defined, and when you ask for clarification, you're directed to figure it out?
The companies lose much, much more. The lower pay and the more pain (inconvenience? Call it what you will. I mean all of those tiny ducks that are pecking us to death) to the writers, the lower quality the work, and the fewer writers willing to undertake that work, those fewer writers being the ones who have no choice: the least proficient, the least experienced. And the most highly skilled writers simply decide they've had enough and they move on to greener (or at least different) pastures.
Most importantly, the children who are taking the tests have already lost when they're faced with low-quality materials that don't provide them with a fair chance to demonstrate what they know and can do.
All right. Let's move on. I'm eager to address the basic rules of item writing (a version of which you can see here, in the Quality Control Checklist published by CCSSO), but I realize I should first define some terms.
An item is a test question. An item may be discrete, or may depend on some external stimulus, such as a reading passage or a chart or a map or something else.
Here is a discrete item:
Why does my dog Sophie bark at mail carriers?
A She is flat-out crazy.
B She is outraged by uninvited guests.*
C She knows something about them that we don't.
D She wants to register a protest about mail delays.
The above is a multiple-choice question, and contains a stem ("Why does my dog bark at mail carriers?") and four answer choices: one correct response (B, as far as I can tell, but I think maybe C is a possible right answer) and three distractors. Distractors, which used to be known as "foils," are wrong answers. Don't get hung up on the language--the point is never to distract nor entice the test-taker to bubble the wrong answer; the point is to create wrong answers that have a reasonable foundation in common mistakes kids would make with that particular skill or bit of content knowledge. More on this later. But tests should never be tricky.
A multiple-choice item is usually worth one score point, and used to be budgeted for one minute of test-taking time, not including the time it takes to read a passage or examine whatever stimuli is needed to answer the question.
There are other item formats: constructed-response items, which are also known as open-ended items. These require the student to provide a response. The response may be as short as a word or a phrase, or, in the case of extended-constructed-response items, the response may be a complete essay.
Here is a short constructed-response item:
Write two words to describe my dog Sophie. Use details to support your answer.And here is the scoring rubric:
2 points: The response includes two accurate describing words, and is supported by relevant evidence.
1 point: The response includes one accurate describing word, and is supported by relevant evidence, OR the response includes two accurate describing words with no supporting evidence.
0 points: The response is blank, illegible, off-topic, or otherwise impossible to score.A short constructed-response item would usually have a score point range of 0-2 or 0-3, and would be budgeted for 5-10 minutes. More than that is usually reserved for an ECR, which could take as few as 15 minutes, or as long as an hour or more for a full essay.
An extended-constructed-response item would look like this:
Considering Sophie's protective nature, do you think it is wise for strangers to approach her? Why or why not? Write an essay in which you discuss the wisdom of approaching a dog with whom you are personally unacquainted.I don't provide a writing rubric because they are complex creations, but you may see some examples here and here. The score point ranges for ECR items vary, depending on the traits of writing and number of domains. That is, an essay might be scored for organization, style, and conventions. If the question depends on the student's comprehension of a passage, the essay might be scored for both reading and writing.
Bear in mind that these sample items are jokes, and as such, aren't examples of exemplary items, primarily because they require a great deal of prior knowledge, and so the test-taker who is unfamiliar with Sophie and dogs in general will perform less well than the test-taker who is on a first-name basis with Sophie and/or other dogs. There are other, less egregious flaws, but we'll get to those when we get to them.
If you have an item you'd like me to examine, explain, or deconstruct, feel free to post it in the comments. Check the copyright first.
What I'm reading: Forgot to mention I was also finishing up The Claverings by Anthony Trollope. Then it's back to As I Lay Dying. I gave up on the other.
Saturday, April 28, 2012
Pass the Pineapple
This, from Jo Perry, the beginning of a discussion about the larger context for the sleeveless talking pineapple:
An American child could go to a public school run by Pearson, studying from books produced by Pearson, while his or her progress is evaluated by Pearson standardized tests. The only public participant in the show would be the taxpayer.
If all else fails, the kid could always drop out and try to get a diploma via the good old G.E.D. The General Educational Development test program used to be operated by the nonprofit American Council on Education, but last year the Council and Pearson announced that they were going into a partnership to redevelop the G.E.D. — a nationally used near-monopoly — as a profit-making enterprise.
I'm very interested in this conversation. I'll say upfront that although I find it disagreeable to point at problems without offering possible solutions, this one's got me baffled.
There are not-for-profits that publish curriculum and assessment materials. From what I've seen, many operate just as corporations do, but perhaps more cheerfully, said operations being subsidized by what I imagine are tax breaks that lend some comfort to the proceedings.
Twice I've been recruited by not-for-profit agencies that publish test materials. Nothing seemed any different than any corporation. During the come-work-with-us talk, the VP assured me that just because their agency was a not-for-profit, this did not mean they didn't make a profit. He told me they liked to think of themselves as a meritocracy, and then he wrote a salary figure on a piece of paper and slipped it across the table.
From what I've seen of public education--and I've spent a tremendous time in classrooms at every grade, in review meetings with teachers, administrators, and other education professionals, and in state DOE conference rooms--I can't say that the public sector manages anything better than businesses or not-for-profits do.
(The elephantine factor is one problem. The larger a system is, the more difficult its management.)
In the immortal words of Tolstoy, "Everyone thinks of changing the world, but no one thinks of changing himself."*
Every time I've emerged from a classroom or a conference room (or even a presentation at an industry conference) feeling optimistic, it's been because of one person. A person who cares and whose work and words show that she cares. (I use "she" out of habit, not to be exclusionary.) There are brilliant and dedicated teachers in our schools. There are brilliant and dedicated leaders in education. (Some of these work with the corporations, by the way--there are certain names that always reassure me even before I read the recommendations based on their research.) There are people working in the corporations who are deeply and sincerely dedicated to serving students in their work.
There are many who aren't.
My feeling is that whatever your work, if you're just in it for the paycheck, you're doing yourself, your employer, and the end-user a tremendous disservice. We humans need to find and serve a higher meaning.
It shows when we don't. It shows, whether we work behind the counter at Starbucks or with a bunch of tiny little savages kindergarteners in an elementary school, or in a partitioned cubicle in a big corporation.
* I hope you understand I don't mean Gail Collins when I say this. She is calling our attention to a matter worthy of discussion.
Monday, February 13, 2012
In Defense of Quality
If you cannot learn to love real art, at least learn to hate sham art.
This, from William Morris.
By “real art,” let’s say James Lesesne Wells, for example, and by “sham art,” let’s say Thomas Kinkade.
I’d like to apply this sentiment to the work of writing, and more specifically, to the work of test content development. Quite frankly, I’m mystified (a phrase I borrow from an assistant district attorney with whom I used to work when I was on a very different career path, and who was in the habit of using this phrase to sharpen his tongue as he prepared to slice me up for having done something with which he disagreed) by not only the deep and devastatingly obvious diminution of quality in test content in the past few years, but also by the failure of people in this silo of the industry to recognize this trend.
Quite frankly, it breaks my heart. As silly as it sounds. But when you love, you expose yourself to the risk of heartbreak. Again, I turn to William Morris, who said, “Give me love and work – these two only.” I’ve been doing this work for 19 years now; though I got into it thinking it was a temporary rope to keep me out of the quicksand until I found my magic circle niche place on solid ground, I think we can all agree it’s become a long-term relationship.
If this lack of quality trend were limited to newcomers to the business, we could propose that them entry-level young’uns [*sigh*] are poorly educated and ill equipped to express themselves except via texting, which you can certainly see in their editorial comments (which are lamentably rich in acronyms, emoticons, and which betray an unfortunate juvenile fondness for excess punctuation and using all caps in directions, which cannot help but set one’s back up, however patient one might be, and anyone who knows me knows that overly patient I be not).
But no, all we content dev folk – ELA, math, science, and social studies, not one is immune, no, not one – have noticed, and we do talk about it, and the conversation and all the various repetitions and iterations of the same conversation bore and horrify us so that we are reduced to shaking our heads and turning our attention to some vision of an oasis, such as the cocktail that awaits the end of the day.
Back in the day, when I worked at Great Big Huge Test Publishing Company, I went to a mandatory training on root cause analysis. We used the fishbone chart. As trainings go, it was all right. Certainly better than the one at which I was accused of not doing my work and letting my teammates pick up my slack because I failed to participate in the assembling of a puzzle, which failure actually had a lot to do with my abysmally poor spatial intelligence and equally poor vision (since corrected through the wonders of Lasik surgery) and little to do with my work ethic, which, as it happens, is about as Puritan as a work ethic might be. You can take the girl out of the working class, but you can’t take the working class out of the girl. But I digress.
If we performed a root cause analysis on the wreckage of Good Ship Quality, what would we find?
To answer that we’d have to go back to the beginning. When I started as a content editor, I was dedicated to one project. That project was my one, my only, my all in all. It was the same for my co-workers. That was the early 90s. In most states, large-scale tests were restricted to reading, writing, and math, and were administered at three or four grades (usually something like 4, 6, 8, and 10, or 5, 7, and 11).
Five years later, it was a whole different and bigger but not necessarily better ballgame. More states were testing more grades, and NCLB loomed on the horizon. As a supervisor, I was responsible for five projects. No one on my team was solely dedicated to any one project; each person, from editor to supervisor, worked on several.
I had a meeting with my manager that went like this:
Manager: [peering at her clipboard] All right, so you have State V, State W, State X, and State Y.
Me: And State Z.
Manager: Oh, I forgot about Z. Right. State Z. [scribbles a note on her clipboard]
Me: What is the order of priority?
Manager: [pause] They’re all priorities.
Me: With five states, mistakes are going to be made. It’s impossible to supervise five projects of this scale. Which state is going to be the mistake state?
Test publishing companies couldn’t handle the workload. Companies that had never done any testing smelled the money and jumped into the fray. All companies got hiring fever. By then, I was a content development manager hiring entry-level candidates at more than twice my starting salary as an editor (and did that ever sting, I tell you what).
But the equation for meeting a production deadline is
TIME + WORKERS = MEETING PRODUCTION DEADLINE
If you have less time, you need more workers. Fewer workers, you need more time. I am no math expert, but this equation I know.
Deadlines got more and more compressed, development cycles shrank, and everyone starting skipping steps. Real training gave way to on-the-job training, which really means sink-or-swim training. New-hires were handed the comprehensive binder containing lists of processes and procedures, which binders were relegated to shelves in cubicles because no one had time to read them. Early field tests were cast aside. Sometimes all field tests. There were fewer internal reviews. The few remaining reviews were performed by overworked and/or underexperienced staff—and you can actually determine which is which (and which is both) when you see the editorial feedback coming out of such reviews.[1]
Deadlines got more and more compressed, development cycles shrank, and everyone starting skipping steps. Real training gave way to on-the-job training, which really means sink-or-swim training. New-hires were handed the comprehensive binder containing lists of processes and procedures, which binders were relegated to shelves in cubicles because no one had time to read them. Early field tests were cast aside. Sometimes all field tests. There were fewer internal reviews. The few remaining reviews were performed by overworked and/or underexperienced staff—and you can actually determine which is which (and which is both) when you see the editorial feedback coming out of such reviews.[1]
Another significant factor may be a corollary to the Peter Principle. The most highly skilled, knowledgeable, and experienced line staff keep getting promoted to management, where they may be doing a fantastic job, but their spots are filled either by new hires or old hands who are left behind (how can I say this delicately? Their remaining behind may not always be by their own choice). Combine this with the absence of training, and it’s a chaos cocktail.
Not to mention the dependence on freelancers. Companies started laying people off and then rehiring them as subcontractors. For some, it’s a win-win—the company don’t have to pay your benefits, and you get to work at home in your pajamas—but it do mean there are a heck of a lot of people at their keyboards writing test questions who neither have experience in education nor in publishing, let alone assessment, which some of us choose to believe is both an art and a science.
There is value in enduring years of slogging through the entire publishing cycle from first draft through bluelines over and over and over again. There is value in having logged many hours in the company of small children struggling to read. There is value in meeting with what the industry calls the stakeholders—teachers, administrators, community leaders, DOE officials. There is value in stretching to accommodate the demands of the stakeholders. There is value in educating oneself about the history and practices of one’s profession. Those learn-to-play-the-piano-in-10-minutes books aside, there is no shortcut to attaining mastery in anything.
There are so many facets to what we do in assessment content development, and when one’s experience is restricted to one tiny mirrored triangle of the great big disco ball, well, that creates a problem because one hasn’t constructed a greater context which allows for greater meaning to inform and guide the work. When the work is simply writing questions for a paycheck and meaning goes out the window, the questions get lamer and lamer, by which I mean trivial, superficial, and plagued by error.
However, the purpose of identifying a problem is not to castigate wrong-doers, nor to enjoy that most basic human pleasure of being right, but to use such identification to find a solution.
The answers are probably as clear to you as they are to me:
- 1. Only hire content developers (freelance or in-house, I have no axe to grind here) who either have a proven track record of providing high-quality work or who have the capability (combination of education, writing skills, content area expertise, intelligence, creativity, and persnicketiness) to learn how to do the work well
- 2. Provide not only adequate but excellent training
- 3. Employ senior content development personnel [*ahem* not naming any names] to review items and provide specific instructional feedback to writers
- 4. Budget sufficient time and money for the given project
Well, there you go. There’s nothing new under the sun, in the immortal words of King Solomon.
2/23/12 UPDATE: This just landed in my inbox, posted by Erik Robelen at Curriculum Matters at EdWeek:
You will want to read Annie Keeghan's original blog post here at Open Salon.
2/23/12 UPDATE: This just landed in my inbox, posted by Erik Robelen at Curriculum Matters at EdWeek:
The “new normal” in educational publishing is “a severe lack of oversight in the quality of curriculum being produced” and a “frightening apathy” to do anything about it. Keeghan’s piece, “Afraid of Your Child’s Math Textbook? You Should Be” is a jeremiad. It does for textbook publishing what The Jungle did for the meatpacking industry.
Keeghan paints a bleak and dispiriting picture of a business gutted by mergers, competition for fewer available dollars, and an increased focus on sales and marketing at the expense of producing quality products. Materials rushed to market at breakneck speed are “inherently, tragically flawed.” Plus the pool of qualified writers and editors is drying up, and those doing the work “often don’t have the necessary skills or experience to produce a text worthy of the publisher’s marketing claims,” she writes.
You will want to read Annie Keeghan's original blog post here at Open Salon.
[1] Overworked but highly experienced people skim text, which forces their brain to fill in the gaps. Which means erroneous assumptions and conclusions drawn from limited evidence and resulting in unnecessary, ill-advised edits. Subtleties or fine distinctions are impossible to detect when skimming. Underexperienced people often restrict their scope of what’s acceptable to their own narrow band of direct experience, and then reject what lay in the outer darkness of their ignorance. This is bad enough on its own, but they will then assume a pedantic tone and lecture the writer for having written items that exceed the demands of the specifications.
Wednesday, October 7, 2009
An Angry Little Toot from a Lone Brave Whistle-Blower
The books I am reading right now are: Love and Will by Rollo May, Aretha Franklin's autobiography Aretha: From These Roots, and Daniel Goleman's Vital Lies, Simple Truths: The Psychology of Self-Deception. And I just finished Making the Grades: My Misadventures in the Standardized Testing Industry.
There may not, at first glance, appear to be any common ground (in the immortal words of Rev. Jesse Jackson) among these, but I will argue that as humans, we take ourselves wherever we go, and in so doing, we drag along the burdens of either consciousness or self-deception. As we choose. Sometimes one, sometimes the other, as best meets our needs and best fits our capacities at the time, the question of existence often being the same as that faced by Oedipus, (as Rollo May says): how much self-awareness can a human being bear? One hopes--I hope--that one may travel an upward trajectory in which one increases one's awareness of oneself and of the world, a trajectory that leads to some higher plane in which we can learn to live with truth.
We take ourselves to work, where we sometimes must walk a tightrope between our values and the need to earn a living. Most of us must make some compromises, must sell ourselves in some way. Some compromises are small and meaningless, but others may put our very integrity at stake. This is the real story of Farley's book.
Though the book is purportedly about the testing and test publishing industries, it is just as much about Farley, who presents himself as a whistle-blower, though one might be forgiven for the mean-spirited thought that Farley sure did take his time in finding the whistle, being as he kept on collecting a paycheck from The Great Satans for many years. And that perhaps this delay lends a bit of tarnish to his credibility.
It must also be said that Farley does not appear to best advantage when he writes about how he copied other workers' scores to get out of doing the work himself, or about spending most of his workday surfing the Internet, or about his hand-rubbing glee in charging exorbitant fees as a consultant (though maybe my pointing out the latter is evidence of envy on my part, as I cannot help but wonder how he managed to pull this off, as I have been a consultant in this industry for 8 or 9 years, and though I do support my little family, our style of living cannot be described as high off even a tiny hog). (Not to mention the subtle sexism in Farley's thinking that is revealed in his writing. Look at how his view of women is first invariably filtered through the lens of whether or not he finds them attractive. In his mind, women--no matter how accomplished or intelligent--are reduced to decorative objects because of course, a woman's main value has to do with whether you enjoy gaping at her. Then consider the adjectives and nouns he uses when writing about women. He says he traveled with "a gaggle" of women. Oh, please. Yes, we of the feminine persuasion are all just clacking geese, you know how ladies love to gab. Sigh.)
Reading Farley's book raises as many questions about him as it does about the industry he intends to expose. If it were so chock-full of despicable practices, why did he remain there for 15 years? How did making such a sacrifice of his own values and beliefs affect him? What exactly was going on in his mind as he participated in these ethics violations? While working in that industry, what efforts did he make for reform?
It seems that Farley wants to rail against this industry-wide malfeasance without taking any responsibility for his own role in it, but as it do say in the Bible, one cannot touch pitch and not be defiled.
This is a dilemma in which you might say I have a deep and abiding interest. The spirit of full disclosure compels me to state that I worked at CTB McGraw-Hill from 1993 to 2001, since which time I have been a content development consultant for a variety of test publishers, school districts, and one state department of education. Having worked with most of the major test publishing companies, I can say that I have seen a good share of corporate culture, and the more I see, the more I am glad I work for myself. Anyone who has ever worked in a corporate setting is probably familiar with at least some of the horrors Farley describes. People who are like the walking dead, who are so eccentric and odd-mannered as to seem unemployable, catch-22 mandates handed down from upper-level management, meetings that seem designed to showcase pomposity and vanity and futility.
But the shenanigans of wrong-doing in hand-scoring that Farley reveals, the behind-the-scenes falsification of scores, the pressure to score in one direction or another as the wind from on high changes, demands from psychometricians to increase the number of scores at a given grade point level, hand-scorers who were the dregs of society--these I did not see. Pressure to work faster, for higher productivity, yes. Unreasonable demands, yes. Obsequious sucking up to district or state officials, yes, yes, yes. Lots of co-workers with their little quirks, oh, yes.
As there is in the world, there is much that could and should be improved in this industry. On all sides, and probably in every department in every test publishing company. If Farley says that what he describes in the book was his experience, then that was his experience, I will not dispute that, though my own experience has been different. I agree completely that tests today are being used for purposes for which they should not be used. I agree completely that there is more to learning than can be measured by a paper and pencil test (or a keyboard test). I agree completely that one of the unintended consequences of the No Child Left Behind legislation has been the unleashing of unprincipled money-sniffing dogs into the industry (no offense to literal dogs), and the muck they try to pass off as genuine content--well! I have seen some awful terrible bad no-good things, is what I am saying.
And yet, testing is never going to disappear. Nor should it. The example I always give when a stranger tries to hold my feet to the fire is whether you would want to undergo an invasive medical procedure at the hands of a surgeon who had never submitted to (let alone achieved a passing grade from) any kind of examination as to his or her knowledge and skill. Let's face it, we don't even want to take our cars to mechanics who are not certified in some manner.
See, we can construct this evil villain testing empire, we can make paper dragon cut-outs all we want, but how does that effect real change? What about starting where we are? For myself, I find that much of my work has been coming more from curriculum the last few years, partly because these particular clients are just plain charming and nice to work with, but partly because, given the choice, I would rather be working on the side of remediation and intervention. Not that I have stopped my assessment work. However, if I ever do feel about it the way Farley did--if it stole from me my integrity--I hope that I would not hesitate to leave it behind. I can't really know what I would do unless I find myself in that situation. Some sleepless nights would result, I am sure. Rollo May says that fate plus guilt equals no rest for the wicked. (My paraphrase.)
There may not, at first glance, appear to be any common ground (in the immortal words of Rev. Jesse Jackson) among these, but I will argue that as humans, we take ourselves wherever we go, and in so doing, we drag along the burdens of either consciousness or self-deception. As we choose. Sometimes one, sometimes the other, as best meets our needs and best fits our capacities at the time, the question of existence often being the same as that faced by Oedipus, (as Rollo May says): how much self-awareness can a human being bear? One hopes--I hope--that one may travel an upward trajectory in which one increases one's awareness of oneself and of the world, a trajectory that leads to some higher plane in which we can learn to live with truth.
We take ourselves to work, where we sometimes must walk a tightrope between our values and the need to earn a living. Most of us must make some compromises, must sell ourselves in some way. Some compromises are small and meaningless, but others may put our very integrity at stake. This is the real story of Farley's book.
Though the book is purportedly about the testing and test publishing industries, it is just as much about Farley, who presents himself as a whistle-blower, though one might be forgiven for the mean-spirited thought that Farley sure did take his time in finding the whistle, being as he kept on collecting a paycheck from The Great Satans for many years. And that perhaps this delay lends a bit of tarnish to his credibility.
It must also be said that Farley does not appear to best advantage when he writes about how he copied other workers' scores to get out of doing the work himself, or about spending most of his workday surfing the Internet, or about his hand-rubbing glee in charging exorbitant fees as a consultant (though maybe my pointing out the latter is evidence of envy on my part, as I cannot help but wonder how he managed to pull this off, as I have been a consultant in this industry for 8 or 9 years, and though I do support my little family, our style of living cannot be described as high off even a tiny hog). (Not to mention the subtle sexism in Farley's thinking that is revealed in his writing. Look at how his view of women is first invariably filtered through the lens of whether or not he finds them attractive. In his mind, women--no matter how accomplished or intelligent--are reduced to decorative objects because of course, a woman's main value has to do with whether you enjoy gaping at her. Then consider the adjectives and nouns he uses when writing about women. He says he traveled with "a gaggle" of women. Oh, please. Yes, we of the feminine persuasion are all just clacking geese, you know how ladies love to gab. Sigh.)
Reading Farley's book raises as many questions about him as it does about the industry he intends to expose. If it were so chock-full of despicable practices, why did he remain there for 15 years? How did making such a sacrifice of his own values and beliefs affect him? What exactly was going on in his mind as he participated in these ethics violations? While working in that industry, what efforts did he make for reform?
It seems that Farley wants to rail against this industry-wide malfeasance without taking any responsibility for his own role in it, but as it do say in the Bible, one cannot touch pitch and not be defiled.
This is a dilemma in which you might say I have a deep and abiding interest. The spirit of full disclosure compels me to state that I worked at CTB McGraw-Hill from 1993 to 2001, since which time I have been a content development consultant for a variety of test publishers, school districts, and one state department of education. Having worked with most of the major test publishing companies, I can say that I have seen a good share of corporate culture, and the more I see, the more I am glad I work for myself. Anyone who has ever worked in a corporate setting is probably familiar with at least some of the horrors Farley describes. People who are like the walking dead, who are so eccentric and odd-mannered as to seem unemployable, catch-22 mandates handed down from upper-level management, meetings that seem designed to showcase pomposity and vanity and futility.
But the shenanigans of wrong-doing in hand-scoring that Farley reveals, the behind-the-scenes falsification of scores, the pressure to score in one direction or another as the wind from on high changes, demands from psychometricians to increase the number of scores at a given grade point level, hand-scorers who were the dregs of society--these I did not see. Pressure to work faster, for higher productivity, yes. Unreasonable demands, yes. Obsequious sucking up to district or state officials, yes, yes, yes. Lots of co-workers with their little quirks, oh, yes.
As there is in the world, there is much that could and should be improved in this industry. On all sides, and probably in every department in every test publishing company. If Farley says that what he describes in the book was his experience, then that was his experience, I will not dispute that, though my own experience has been different. I agree completely that tests today are being used for purposes for which they should not be used. I agree completely that there is more to learning than can be measured by a paper and pencil test (or a keyboard test). I agree completely that one of the unintended consequences of the No Child Left Behind legislation has been the unleashing of unprincipled money-sniffing dogs into the industry (no offense to literal dogs), and the muck they try to pass off as genuine content--well! I have seen some awful terrible bad no-good things, is what I am saying.
And yet, testing is never going to disappear. Nor should it. The example I always give when a stranger tries to hold my feet to the fire is whether you would want to undergo an invasive medical procedure at the hands of a surgeon who had never submitted to (let alone achieved a passing grade from) any kind of examination as to his or her knowledge and skill. Let's face it, we don't even want to take our cars to mechanics who are not certified in some manner.
See, we can construct this evil villain testing empire, we can make paper dragon cut-outs all we want, but how does that effect real change? What about starting where we are? For myself, I find that much of my work has been coming more from curriculum the last few years, partly because these particular clients are just plain charming and nice to work with, but partly because, given the choice, I would rather be working on the side of remediation and intervention. Not that I have stopped my assessment work. However, if I ever do feel about it the way Farley did--if it stole from me my integrity--I hope that I would not hesitate to leave it behind. I can't really know what I would do unless I find myself in that situation. Some sleepless nights would result, I am sure. Rollo May says that fate plus guilt equals no rest for the wicked. (My paraphrase.)
Poppycock, Folderal, Nonsense
. . . in the immortal words of Todd Farley.
About a week ago, someone sent me a link to an Op-Ed piece in the New York Times by Todd Farley, author of Making the Grades: My Misadventures in the Standardized Testing Industry.
Farley's experiences aren't unique. Like Farley, I am a writer who sort of fell into the test publishing industry by accident. Like Farley, I stayed in the industry long after I thought I would have gone on to what I thought would be my real career of writing novels or screenplays or something, anything.
Both of us started our careers in hand-scoring, so hand-scoring is what I will talk about, specifically the hand-scoring of open-ended test questions. Multiple-choice questions are simple to score, because there is only one correct answer. All multiple-choice test questions are machine-scored. The answer sheets or test booklets are scanned, the answer choices verified by machine, and the scores are then computer-generated. Sometimes there are mistakes in the programming that must be corrected, for example, the correct answer to a given question was actually C but was identified somewhere along the line as A. Sometimes there are mistakes with a student's name or identification number that lead to a mistaken score. Sometimes--and this happened with my daughter's third-grade California STAR testbook--the testbook or answer sheet has juice spilled all over it, and so a false score may be generated. Where humans are involved, there will be some error somewhere, it is unavoidable, let us simply endeavor to put checks in place to catch the errors and processes to correct them.
The scoring of open-ended questions is a horse of a whole nother color. By its nature, there must be some subjectivity. In support of standardization are an array of tools that include a scoring guide or rubric, sample student responses at each score-point-level, and anchor papers and rangefinders. A rubric lists the characteristics of the response at each score point, a sample response gives an example of what kind of response is expected, an anchor is a student response that embodies the score point level, and rangefinders show what may be expected at the high, middle, and low ends of the spectrum within a score point level.
It sounds like a complicated process, and it is. And it's not without its ridiculous moments. And I have to say that though I found much about handscoring interesting, the work itself was tedious and the routine unbearable. But it's not the Orwellian circus of nonsense Farley describes. Or maybe it is at the company where Farley worked; it wasn't at CTB McGraw-Hill when I worked in hand-scoring there.
I am only about a quarter into the book, so maybe there will be some sort of Aristotelian discovery on Farley's part. At this point, he sounds like one of the disgruntled hand-scorers, and there were some of those, people who just never got it, never were able to internalize the scoring criteria and constraints, the ones whose scores had to be checked and re-checked so often that eventually they were let go. He says that he failed to qualify as a scorer for a writing test, which does make one wonder whether this type of work simply was not a good fit for him. Not that I can vouch for what happened at Pearson, as I've not worked there.
I will also say that--although I do not at all see myself as a flag-waver for the test publishing industry, and that I have my own strong feelings about the mis-use of tests and what seems to me to be an abuse of tests and how they are used and what they symbolize and how the data are manipulated--sitting in the mocking judgment seat is generally easy to do. I have plenty of ridiculous stories of my own. We humans are ridiculous, it's in our nature, and thank God that we are, it makes the world so much more entertaining.
And this book is just that--entertainment, a joke that is masked as an indictment of the industry. For myself, I'd be a lot more interested in a thoughtful exploration of the subject, one that takes into account the need for measurement in teaching, and the demand for standardization (because that seems to be the only way to ensure any kind of fairness or equity), and how we could possibly balance these kinds of standardized measurements with classroom performance and evaluations from teachers.
CORRECTION: I mean "Folderol." Geez. And to think I won first place in the 8th grade spelling bee. What did I tell you? Human error.
About a week ago, someone sent me a link to an Op-Ed piece in the New York Times by Todd Farley, author of Making the Grades: My Misadventures in the Standardized Testing Industry.
Farley's experiences aren't unique. Like Farley, I am a writer who sort of fell into the test publishing industry by accident. Like Farley, I stayed in the industry long after I thought I would have gone on to what I thought would be my real career of writing novels or screenplays or something, anything.
Both of us started our careers in hand-scoring, so hand-scoring is what I will talk about, specifically the hand-scoring of open-ended test questions. Multiple-choice questions are simple to score, because there is only one correct answer. All multiple-choice test questions are machine-scored. The answer sheets or test booklets are scanned, the answer choices verified by machine, and the scores are then computer-generated. Sometimes there are mistakes in the programming that must be corrected, for example, the correct answer to a given question was actually C but was identified somewhere along the line as A. Sometimes there are mistakes with a student's name or identification number that lead to a mistaken score. Sometimes--and this happened with my daughter's third-grade California STAR testbook--the testbook or answer sheet has juice spilled all over it, and so a false score may be generated. Where humans are involved, there will be some error somewhere, it is unavoidable, let us simply endeavor to put checks in place to catch the errors and processes to correct them.
The scoring of open-ended questions is a horse of a whole nother color. By its nature, there must be some subjectivity. In support of standardization are an array of tools that include a scoring guide or rubric, sample student responses at each score-point-level, and anchor papers and rangefinders. A rubric lists the characteristics of the response at each score point, a sample response gives an example of what kind of response is expected, an anchor is a student response that embodies the score point level, and rangefinders show what may be expected at the high, middle, and low ends of the spectrum within a score point level.
It sounds like a complicated process, and it is. And it's not without its ridiculous moments. And I have to say that though I found much about handscoring interesting, the work itself was tedious and the routine unbearable. But it's not the Orwellian circus of nonsense Farley describes. Or maybe it is at the company where Farley worked; it wasn't at CTB McGraw-Hill when I worked in hand-scoring there.
I am only about a quarter into the book, so maybe there will be some sort of Aristotelian discovery on Farley's part. At this point, he sounds like one of the disgruntled hand-scorers, and there were some of those, people who just never got it, never were able to internalize the scoring criteria and constraints, the ones whose scores had to be checked and re-checked so often that eventually they were let go. He says that he failed to qualify as a scorer for a writing test, which does make one wonder whether this type of work simply was not a good fit for him. Not that I can vouch for what happened at Pearson, as I've not worked there.
I will also say that--although I do not at all see myself as a flag-waver for the test publishing industry, and that I have my own strong feelings about the mis-use of tests and what seems to me to be an abuse of tests and how they are used and what they symbolize and how the data are manipulated--sitting in the mocking judgment seat is generally easy to do. I have plenty of ridiculous stories of my own. We humans are ridiculous, it's in our nature, and thank God that we are, it makes the world so much more entertaining.
And this book is just that--entertainment, a joke that is masked as an indictment of the industry. For myself, I'd be a lot more interested in a thoughtful exploration of the subject, one that takes into account the need for measurement in teaching, and the demand for standardization (because that seems to be the only way to ensure any kind of fairness or equity), and how we could possibly balance these kinds of standardized measurements with classroom performance and evaluations from teachers.
CORRECTION: I mean "Folderol." Geez. And to think I won first place in the 8th grade spelling bee. What did I tell you? Human error.
Subscribe to:
Posts (Atom)