Showing posts with label content development. Show all posts
Showing posts with label content development. Show all posts

Tuesday, April 3, 2012

Gap

As obvious as the link is between the quality of content (we could also say validity) and item writing training, it remains a mystery to me how sound practices in item writing training have nearly become obsolete.


Steven M. Downing discusses this link in "Twelve Steps for Effective Test Development" from the Handbook of Test Development:
Yet knowing the principles of effective item writing is no guarantee of an item writer's ability to actually produce effective test questions. Knowing is not necessarily doing. Thus, one of the more important validity issues associated with test development concerns the selection and training of item writers . . . . The most essential characteristic of an effective item writer is content expertise. Writing ability is also a trait closely associated with the best and most creative item writers.
. . .
Effective item writers are trained, not born. Training of item writers is an important validity issue for test development. Without specific training, most novice item writers tend to create poor-quality, flawed, low-cognitive-level test questions that test unimportant or trivial content. Although item writers must be expert in their own disciplines, there is not reason to believe that their subject matter expertise generalizes to effective item writing expertise. Effective item writing is a unique skill and must be learned and practiced. For new item writers, it is often helpful and important to provide specific instruction using an item writer's guide, paired with a hands-on training workshop (Haladyna, 2004). As with all skill learning, feedback from expert item writers and peers is required. The instruction-practice-feedback-reinforcement loop is important for the effective development and maintenance of solid item writing skills. . . .
The best and most effective training, then, is to teach item-writing skills to content area experts, give them a guide for reference, let them practice, review and comment on their work, have them make revisions, and review and comment again until the items are satisfactory. Add a peer review step -- but only if the peers have a thorough understanding of the principles for developing sound items. Repeat as necessary.


In other news, April is National Poetry Month. I like the idea of carrying a poem in your pocket.

Thursday, March 15, 2012

It's Like That

. . . in the immortal words of Run DMC


Recently, as part of my quality reform campaign, I took several protégées under my wing. It was a responsibility I'd been ducking for years. I've always thought that I more than met my obligations to my industry by simply doing the best work that I could do.


But how much of a difference can I make by myself? Yesterday I was talking with a friend who mentioned that 22,000 items will be coming through the agency where she wields her magic works. 22,000! How is it even possible to glance at 22,000 items, let alone perform a thorough content review?


Hence the protégées, whom I've been training in the manner in which I was trained, but more so. That is, much of the training I received was on-the-job, and even though it was much better, much more comprehensive than what is passing (or not passing) for training in most places nowadays, there certainly could have been more of it. (More did come later, when the director of our department launched the development of a processes and procedures manual.) Which I recognized at the time, and which inspired me to educate myself in the work of the industry.


I spent a great deal of time talking to colleagues in various departments of the Great Big Huge Company where I was employed. Like many companies, Great Big Huge Company had a silo culture. People mostly stuck with their own. The programmers ate lunch with the programmers. The style editors took walks together. The content people went out for coffee together. The psychometricians mostly remained in their offices, except for when they appeared at meetings. The upstairs people stayed upstairs, while the people in operations kept everything in motion downstairs. I went everywhere (I got lost a few times in that first few months, once in the warehouse, a cavern with concrete floors on which were stacked towers of testbooks on wooden pallets) and talked to everyone, from those in shipping and operations to manufacturing and finance. I wanted to know what everyone did and how everyone's work fit together to create the bigger picture. (With much envy, I listened to stories of the glory days of Great Big Huge Company, the days of raucous St. Patrick's Day parties and of flights in some executive's small plane and of leisurely, sociable Friday lunches. It was the '80s.)


I started reading. A lot of what I read made no sense to me. I didn't speak the language yet. It was discouraging to understand so little, but I approached it as if it were grad school and kept reading, kept asking questions, and always, kept doing the work, and then eventually I knew what I was doing.


This sense of mastery is what I hope for the protégées--this group, the group to come, maybe many groups in the future. Maybe these will go on to train others in best practices. Now that I've begun, I see that it's a worthwhile endeavor. Two or three or four protégées who may go on to become experts and provide work of the highest quality may not sound like a lot. But these two or three more people doing excellent work will leave less room for two or three bumblers. And then when clients get into the habit of receiving excellent work, they'll have less patience for the substandard. That's my fond hope, anyway.


Assessment content development is a funny little world. There's not much published information about what it is and how to do it right. Content developers don't often meet to share information, to talk about challenges and discuss solutions. When people ask me for sources, I point them in the direction of Thomas Haladyna. James Popham. Grant Wiggins. Jay McTighe.


Even so, one won't always agree, and when one does agree, there will be gaping chasms gaps between theory and execution.


[Thanks to Bob DeBris, who introduced me to the video posted above.]

Monday, March 5, 2012

First Things First

Years ago, I attended a writing workshop with Tom Jenks and Carol Edgarian. Now I don't even remember how I decided to go to this workshop, but I knew it was the right place for me when I learned that Anna Karenina (this translation) and Aristotle's Poetics were among the required reading for the week.


Now I'm reading Anna Karenina (again, I don't know how many times I have read it, though I don't think one could ever tire of it), and so I was thinking about something Tom and Carol said during the workshop, about how one's writing needs to have a big idea, that everything--everything!--in the writing should be developing and supporting this big idea.


As modest an endeavor as writing a test question is, it is still writing. Each question must have a big idea. And the item writer must know what that big idea is, and then marshal all of the everything of the item in support of that big idea.


Last November I went to that heaven on earth, home of all manner of delights, such as Frank Lloyd Wright houses and the Art Institute and French-fried green beans. I love Chicago. I love the people there--Stormy, husky, brawling/City of the Big Shoulders. (As the doorman of my hotel, whom I had come to know fairly well in the course of a few days of comings and goings, put me in a cab to travel to Oak Park late on a Friday evening, he said to the driver, "This here is my sister. You take care of her like she is my sister. You understand?" The funny part is that, indeed, we were of an age and did resemble each other closely enough to be siblings. Siblings raised apart, as we didn't share that distinctive accent.)


The reason for my visit was not just to have the best time in the world with my friend Carrie; it was to give a presentation on an assessment content development-related topic of my choosing. I rummaged around in my brain for a while before I came up with intention.


I thought of intention because of something that Kate Nash, my friend and genius dance teacher once said, probably in response to all of our awkwardly flailing limbs: that if you focus your intention, your bones will organize themselves. It reminded me of how, when I learned to ski, everyone told me not to look down, because I would fall. You cannot help but follow the direction of your gaze.


Then I found there was what seemed like an orienting magic in setting an intention. Having a laserlike focus burns off what is inessential. It builds a solid foundation. I tried intention in areas of my life other than dance. There was something about doing this simple step first that cleared away debris distractions.


Around that time, I was asked to perform triage ride in like the cavalry edit a set of questions that had been rejected by a customer. It soon became apparent that the main problem with the set was that the item writer had no firm intention. Not only was there no unifying purpose for the questions as a group, but each question could have been measuring two or more skills. Which meant that none of the questions could measure any skill accurately.


It takes time to do this, but more than time, it takes a reflective pause. Which may not sound like much--but when folks are busy and feeling overwhelmed, the reflective pause gets thrown overboard.


In the immortal words of Epictetus, First say to yourself what you would be, and then do what you have to do.


(When I returned home, my daughters were so jealous of all of my Chicago adventures, especially the French-fried green beans. I found a recipe and made a big batch and we ate a million each. It is safe to say neither of them will ever look at a French-fried green bean again. In fact, there is a ban on saying the words aloud in their presence.)


UPDATE: Identified supernova Kate Nash by name.



Monday, February 13, 2012

In Defense of Quality

If you cannot learn to love real art, at least learn to hate sham art.

This, from William Morris.

By “real art,” let’s say James Lesesne Wells, for example, and by “sham art,” let’s say Thomas Kinkade

I’d like to apply this sentiment to the work of writing, and more specifically, to the work of test content development. Quite frankly, I’m mystified (a phrase I borrow from an assistant district attorney with whom I used to work when I was on a very different career path, and who was in the habit of using this phrase to sharpen his tongue as he prepared to slice me up for having done something with which he disagreed) by not only the deep and devastatingly obvious diminution of quality in test content in the past few years, but also by the failure of people in this silo of the industry to recognize this trend.

Quite frankly, it breaks my heart.  As silly as it sounds. But when you love, you expose yourself to the risk of heartbreak. Again, I turn to William Morris, who said, “Give me love and work – these two only.” I’ve been doing this work for 19 years now; though I got into it thinking it was a temporary rope to keep me out of the quicksand until I found my magic circle niche place on solid ground, I think we can all agree it’s become a long-term relationship.

If this lack of quality trend were limited to newcomers to the business, we could propose that them entry-level young’uns [*sigh*] are poorly educated and ill equipped to express themselves except via texting, which you can certainly see in their editorial comments (which are lamentably rich in acronyms, emoticons, and which betray an unfortunate juvenile fondness for excess punctuation and using all caps in directions, which cannot help but set one’s back up, however patient one might be, and anyone who knows me knows that overly patient I be not).

But no, all we content dev folk – ELA, math, science, and social studies, not one is immune, no, not one – have noticed, and we do talk about it, and the conversation and all the various repetitions and iterations of the same conversation bore and horrify us so that we are reduced to shaking our heads and turning our attention to some vision of an oasis, such as the cocktail that awaits the end of the day.

Back in the day, when I worked at Great Big Huge Test Publishing Company, I went to a mandatory training on root cause analysis. We used the fishbone chart. As trainings go, it was all right. Certainly better than the one at which I was accused of not doing my work and letting my teammates pick up my slack because I failed to participate in the assembling of a puzzle, which failure actually had a lot to do with my abysmally poor spatial intelligence and equally poor vision (since corrected through the wonders of Lasik surgery) and little to do with my work ethic, which, as it happens, is about as Puritan as a work ethic might be. You can take the girl out of the working class, but you can’t take the working class out of the girl. But I digress.

If we performed a root cause analysis on the wreckage of Good Ship Quality, what would we find?

To answer that we’d have to go back to the beginning. When I started as a content editor, I was dedicated to one project. That project was my one, my only, my all in all. It was the same for my co-workers. That was the early 90s. In most states, large-scale tests were restricted to reading, writing, and math, and were administered at three or four grades (usually something like 4, 6, 8, and 10, or 5, 7, and 11).

Five years later, it was a whole different and bigger but not necessarily better ballgame. More states were testing more grades, and NCLB loomed on the horizon. As a supervisor, I was responsible for five projects. No one on my team was solely dedicated to any one project; each person, from editor to supervisor, worked on several.

I had a meeting with my manager that went like this:

Manager: [peering at her clipboard] All right, so you have State V, State W, State X, and State Y.
Me: And State Z.
Manager: Oh, I forgot about Z. Right. State Z. [scribbles a note on her clipboard]
Me: What is the order of priority?
Manager: [pause] They’re all priorities.
Me: With five states, mistakes are going to be made. It’s impossible to supervise five projects of this scale. Which state is going to be the mistake state?


Test publishing companies couldn’t handle the workload. Companies that had never done any testing smelled the money and jumped into the fray. All companies got hiring fever. By then, I was a content development manager hiring entry-level candidates at more than twice my starting salary as an editor (and did that ever sting, I tell you what).

But the equation for meeting a production deadline is

TIME + WORKERS = MEETING PRODUCTION DEADLINE

If you have less time, you need more workers. Fewer workers, you need more time. I am no math expert, but this equation I know.

Deadlines got more and more compressed, development cycles shrank, and everyone starting skipping steps. Real training gave way to on-the-job training, which really means sink-or-swim training. New-hires were handed the comprehensive binder containing lists of processes and procedures, which binders were relegated to shelves in cubicles because no one had time to read them. Early field tests were cast aside. Sometimes all field tests. There were fewer internal reviews. The few remaining reviews were performed by overworked and/or underexperienced staff—and you can actually determine which is which (and which is both) when you see the editorial feedback coming out of such reviews.[1]

Another significant factor may be a corollary to the Peter Principle. The most highly skilled, knowledgeable, and experienced line staff keep getting promoted to management, where they may be doing a fantastic job, but their spots are filled either by new hires or old hands who are left behind (how can I say this delicately? Their remaining behind may not always be by their own choice). Combine this with the absence of training, and it’s a chaos cocktail.

Not to mention the dependence on freelancers. Companies started laying people off and then rehiring them as subcontractors. For some, it’s a win-win—the company don’t have to pay your benefits, and you get to work at home in your pajamas—but it do mean there are a heck of a lot of people at their keyboards writing test questions who neither have experience in education nor in publishing, let alone assessment, which some of us choose to believe is both an art and a science.

There is value in enduring years of slogging through the entire publishing cycle from first draft through bluelines over and over and over again. There is value in having logged many hours in the company of small children struggling to read. There is value in meeting with what the industry calls the stakeholders—teachers, administrators, community leaders, DOE officials. There is value in stretching to accommodate the demands of the stakeholders. There is value in educating oneself about the history and practices of one’s profession. Those learn-to-play-the-piano-in-10-minutes books aside, there is no shortcut to attaining mastery in anything.

There are so many facets to what we do in assessment content development, and when one’s experience is restricted to one tiny mirrored triangle of the great big disco ball, well, that creates a problem because one hasn’t constructed a greater context which allows for greater meaning to inform and guide the work. When the work is simply writing questions for a paycheck and meaning goes out the window, the questions get lamer and lamer, by which I mean trivial, superficial, and plagued by error.

However, the purpose of identifying a problem is not to castigate wrong-doers, nor to enjoy that most basic human pleasure of being right, but to use such identification to find a solution.

The answers are probably as clear to you as they are to me:

  • 1.     Only hire content developers (freelance or in-house, I have no axe to grind here) who either have a proven track record of providing high-quality work or who have the capability (combination of education, writing skills, content area expertise, intelligence, creativity, and persnicketiness) to learn how to do the work well
  • 2.     Provide not only adequate but excellent training
  • 3.     Employ senior content development personnel [*ahem* not naming any names] to review items and provide specific instructional feedback to writers
  • 4.     Budget sufficient time and money for the given project



[1] Overworked but highly experienced people skim text, which forces their brain to fill in the gaps. Which means erroneous assumptions and conclusions drawn from limited evidence and resulting in unnecessary, ill-advised edits. Subtleties or fine distinctions are impossible to detect when skimming. Underexperienced people often restrict their scope of what’s acceptable to their own narrow band of direct experience, and then reject what lay in the outer darkness of their ignorance. This is bad enough on its own, but they will then assume a pedantic tone and lecture the writer for having written items that exceed the demands of the specifications.  

Thursday, March 18, 2010

You Gets What You Pays For

A friend forwarded this ad from a freelance job website:

Looking for 200 multiple choice trivia questions on the subject of Easter. Questions must be divided into ten different topics, each with an easy and hard section. All sections must have the same number of questions. Looking to spend no more than $30 on this, and the questions must be completed within 2 days. All i require is a spreadsheet, with ten different sheets. There should be four possible answers for each question, with the correct one in bold. I am only looking for people who speak fluent English. Please send two samples questions if you are interested. Copyright of all questions will belong to me and you may not reproduce them elsewhere.[Ed. sic, sic, sic.]
Though I often confess that numbers and me, we're not the best of friends, sort of like cranky neighbors who give each other a grudging nod once in a while when we catch sight of each other taking out the trash, let us do the math: 200 test questions in 2 days for $30. Which means, assuming an 8 hour work day (8 hours for work! 8 hours for sleep! And 8 hours for what we will!), that the rate of production will be 4.8 minutes per question at the whopping rate of pay of $1.88 per hour. Don't spend it all in one place.

My friend and I decided that a zero or two must have been left off the $30. But I do have a tiny little voice nagging doubt. Could it be?

Wednesday, November 11, 2009

Fun with Point Biserials

A few weeks ago, I was spending a lot of time analyzing test data, specifically, the p values and pt. biserials of a set of English language arts tests for grades 3 through 11, in order to determine What Went Wrong.

In the preponderance of cases, of course, nothing went wrong. The items performed more or less as expected, with the correlations one might expect--i.e., high-achieving students got the easy questions and most of the hard questions right, while struggling students may have gotten the easy questions right, but pretty much got the hard questions wrong. And it follows that no wrong answer in these solid items attracted more students than the right answer.

But there were items with wacky data, items for which the high-achieving students picked wrong answers, or items with wrong answers that lured a higher number of students, leaving the right answer feeling like an awkward wallflower in a darkened gym at the middle school dance. My task was to review these items and figure out the big why. It is a testament to my thorough and absolute geekiness that I LOVE DOING THIS WORK. I could do it all day long, every day. Oh, my goodness. More fun than a barrel of monkeys. It was such a pleasure that I found it difficult to tear myself away at the end of the day (and honesty compels me to add that I did sneak back to it at night after my daughters were asleep).

Why, you may ask. Oh, for so many reasons! One being that it is fun to play detective, to deconstruct an item by conducting an investigative inquiry which concludes in determining the most likely source of the trouble, whether it be a stem of staggering verbosity, or a fundamental unsoundness in the premise of the item, or a bad practice (e.g., attempting to assess two or more skills with one item). It's like picking up a big tangled knotted mess of yarn and, starting with one end, delicately unraveling the knots and twists and then rolling the yarn back into a nice orderly ball. But with your brain. How often do we really have the opportunity to use our brains in this way, and truly, what can possibly be more satisfying that solving a problem?

Because assessment content development, as its own universe, is governed by its own rules, and--to me, in my ridiculous and relentless geekiness--these rules have a sort of simple elegance, the beauty of rigor and orderliness. (Were I to wax biblical on the matter, I would start talking about the necessity for doing things decently and in order.) So to use the data as a mirror to reflect the soundness (or unsoundness) of the item, diagnose the trouble, and then find a means of repair requires the use of a range of tools that include a knowledge of the rules (and knowing when it is acceptable to bend them) and of how students might interpret an item (which is not always how a test developer intends) and of the content area itself. And it also requires a bit of intuition, that ability to sense your way in the dark that just comes from thoroughly knowing something, in the same manner you could navigate your home with the lights off because you know where the couch is and where stands that big coffee table with the treacherously sharp corners.

Tuesday, March 3, 2009

Stakes

Today I got a shout-out asking about a practice to be employed in a classroom assessment, and it got me thinking about classroom assessments.

In this business, we think of classroom assessments as informal and low- to no-stakes, meaning that there will be no decisions made about student promotion/retention, and that there is no teacher/school/district accountability. There are probably some stakes for the students--they will receive a grade, which will contribute to an overall grade in the class, but the effect of this one grade on this one assessment is fairly minor in the great grand scheme of the universe.

Which is an excellent thing, as so few teachers know much about assessment, what skews an assessment, what practices should be avoided in assessment, which question constructions put up unnecessary obstacles for the test-takers, what kind of wording should be used. How do that happen, I wonder. Not to blame the teachers; the mystery is why teachers aren't generally required to take a class on test content development. Creating a sound assessment that accurately measures the targeted skills or knowledge requires knowing how to do it. You don't just get up from the couch one day and say, Hey, Ima build me a house without doing at least a tiny bit of research. But we don't know what we don't know.

Test content development is not difficult, but one does have to know what are the best practices in order to build a test that has at least some potential of giving a relatively fair and accurate measure of what a student knows and can do. Even if it is just in the classroom, even if the results only affect a tenth of the student's semester grade in one class.

Same thing at the DMV, you know. And those high-stakes tests could not possibly be legally defensible, constructed as they are in such a haphazard way, with all their outliers and overly attractive distractors and so on. Every single assessment professional who takes a DMV test cannot help but exclaim at the shoddy construction. What's so interesting to me about the DMV tests, though, is that even civilians recognize how bad the tests are, even if they cannot identify the reasons.

Tuesday, January 6, 2009

The Poetry of Item Alignment

When I was an associate editor at CTB McGraw-Hill, back in the day, we editors were all a little grateful for the seasonal winter slowdown. It gave us time to catch our breath, file the stacks that had piled up during bluelines, and reacquaint ourselves with our co-workers. Now that I have been self-employed lo, these many years, I like to keep busy--whatever the time of year. Summer, winter, don't make no never mind to me. Idle hands are the devil's workshop.

Just before Christmas, I was doing a bit of writing, some for National Geographic Extreme Explorer, and some for ETS. (When the NGEE article comes out, I'll be sure to let you know. The other is, of course, confidential.) Then my attention was directed to alignments. Or correlations. The two terms are often used interchangeably; both mean identifying the standard (performance indicator, objective, skill, subskill) assessed or targeted by a specific item or task.

The best case scenario in content development is to write the item or task to a standard. That process is more creative, more organic, in the sense that the item or task may be developed to meet the demands of the standard. But many companies find themselves with a bank or pool of perfectly sound items, and to recycle these items for multiple projects is both efficient and cost-effective.

Alignments can be tricky. Sometimes items are shoehorned into standards that are an obvious bad fit. New aligners are especially prone to falling prey to aligning by key word, which is a big mistake. When aligning, it's imperative to keep in mind the spirit of the law, as opposed to the letter of the law. Think about the task and what the task requires that the student know or do, then review the standards (performance indicators, objectives, skills, subskills), and select the one that cleaves unto those knowledge and skill requirements. There may be more than one; there often is a lot of overlap. If no skill fits, better not to force the fit. How can a bad alignment result in meaningful measurement of skill or knowledge?

What I've noticed in reviewing others' alignments is that I can get a sense not only of the breadth and depth of the aligner's content knowledge, but of the aligner's intuitive feel for the content area (and for language use in general)--of the aligner's capacity for understanding subtle nuance. In a way, it really is like reading poetry.

Which may sound ridiculous, and it would be, if we were discussing a simple skill, such as "Use end punctuation correctly."