Wednesday, September 26, 2012

Math Preparation


The ability to count and all of the ramifications of being able to do so is critical in our present world. Counting leads to measuring. It allows us to predict. It permits us to define the world around use in terms other than our own subjective beliefs. “I know my first 10 letters of the alphabet.” “I know all 400 words in my basic reader.” “I am learning 5 new words each day.” “I read one book each week this summer.” When will you be empowered to the point (exhibit mastery) that cramming for and taking NCLB standardized tests will be, “just a waste of time?”

I learned my math on a sheet of paper. A large square was drawn. The square was divided into four squares. Each of these was marked off with four horizontal and four vertical lines. I could count to 100 by filling in the cells. I could count by 2s, 3s, 4s, and 5s, which was really multiplying by 2, 3, 4, and 5. I could count backwards, which was subtracting. I learned to estimate the answer to any math question so I had some judgment if I might not be right (I actually spotted a failing Marchant calculator. One of the wheels failed to stop at nine when it should have, but continued on to four). Later I learned square and square root. And finally logarithms, where addition is multiplication and subtraction is division (to three significant digits on a slide rule). The experience of learning the multiplication tables not only is essential in math but also carries over into learning other things that are not as limited and well defined. Hand calculators allow students to skip the basics and suffer the consequences.

The oldest, time tested, math downloadable listed by the Educational Software Cooperative teaches counting, addition, and subtraction. Next multiplication was included. Educational games and puzzles have always been popular, I think in part, because the authors have fun creating them. In a free enterprise economy, a successful game can support the author and a business is born. Those who directly benefit pay the cost of development rather than taxpayers or a non-profit that obtained its funding from a taxable free enterprise activity.

Word problems for grades 1-8 can now be presented by software. Unique calculator and statistical software are available to supplement the myriad of hand calculators. Some of this software is free. The same is true for graphing calculators. All of these are under listings and math downloadables.

Fractions and algebra are still available as downloadables from Merit Software even though these are much further developed in its online software. Evallutel Multimedia presents algebra and geometry, and more. Math downloadables are listed under classroom tests, educational games, math, and teaching aids. The online offerings in the prior two posts on reading and writing; vocabulary, phonics and spelling; also include math.

All of these lessons in software are designed to free teachers to teach, and students to learn, the subject rather than to spend time trying to memorize answers to questions they guess will be on a standardized test. Use this software to give students the opportunity to develop the sense of responsibility needed to learn at all levels of thinking. They will then be ready for Knowledge and Judgment Scoring, quantity and quality, and as also in Winsteps and Amplifier.

With subject mastery at the proper levels of thinking, there should be very little concern about passing current NCLB standardized tests that are scored at the lowest levels of thinking. IMHO it is failure on the part of school administrators to understand this miss match that locks many failing schools into continued failure. Preparing for a lower level of thinking scored test using lower level of thinking experiences dooms students to failure now and even more so in the future. In math especially, students need to be able to do and to understand, not just be aware of what a teacher has presented.

Students functioning at lower levels of thinking are dependent upon their teacher for maintaining their knowledge and skill levels. Students functioning at higher levels of thinking are capable of relearning as needed. They need a teacher for direction and for the expansion of their abilities, not for the maintenance of their knowledge and skill. The more you develop students, the easier it is to teach and to learn. This is the reason that well prepared third graders survive and flourish.

Wednesday, September 19, 2012

Vocabulary, Spelling and Phonics


Before assessing reading ability, one must be able to read. Before one can read, one must acquire a vocabulary. Given the complexity of American English, one must also know how to spell.

I did research on this matter with underprepared college freshmen in a remedial biology class. The campus computer system presented a sentence from assigned reading with a word missing. Students could enter one letter at a time to build the complete word (they called it Wheel of Fortune). Many students could not enter hints is this way. They insisted on typing an entire word. They had learned to read by the look-say method that does not stress spelling. The most unusual student read each word out loud, expect for tri-syllabic words or larger, for which she inserted, “buzz”. Her explanation was that long words, “really don’t count”.

My own experience in building a vocabulary was given a gigantic boost after taking a series of aptitude tests at the Johnson O’Connor Research Foundation in New York City during my second year in the USAF. I won an Air Force wide contest to attend the IBM Customer Engineering School at Endicott, New York, but the aptitude testing ranked me in the bottom 10% on vocabulary! “A person’s vocabulary level was the best single measure for predicting occupational success in every area.” And further, “…vocabulary is not innate, … it can be acquired.” 

I then spent several months while on duty in Japan reading through the two-volume edition of The Johnson O’Connor English Vocabulary Builder. It worked! I tested out of freshman English at Mizzou two years later, but enrolled in both semesters anyway. You can still buy the book. Or you can use related software from WordSmart’s. We are not talking flashcards here. The point is to embed a word in a meaningful matrix of relationships, such that it has a meaning, for which the word is the label. The Vocabulary Builder also presents good examples of grammar.

The Educational Software Cooperative (ESC) lists a number of authors teaching vocabulary, spelling and phonics. The key to success is that the student wants to learn. Vocabulary empowers. The high performance teacher has that unique ability to promote student engagement, in person or by way of software (educational games). Crossword puzzles are a fun way of teaching spelling and vocabulary. Crossdown and “other cool sites” let you create puzzles from your current spelling lists and other topics, as well as play online.

The oldest, time proven, phonics, spelling and vocabulary program to be listed under English, Animated Beginning Phonics, uses animation to entice young students to persist. An animated character speaks the word or you record the sound for the student to hear in the next advancement listed under Vocabulary and Spelling. Next comes the software for teaching word recognition, meanings, and origin by Merit Software. Today this software is fully developed online as are the remaining ESC offerings in this category. Spelling and vocabulary are included in the reading and writing online software discussed in the previous post. You must check each list and software category for the one that fits your needs the best.

How students use the software determines what they get out of it. Reading aloud and spelling can be conducted at the lowest levels of thinking with any in-person or software presentation. An empowering vocabulary is created when the student wants to learn and the words are truly labels for meaningful observations and relationships. The reader then experiences what the words represent. The accomplished reader can then command a given situation by use of meaningful words. The world would be so different if people took the time to actually understand one another. IMHO that would mean the end of traditional right count scored multiple-choice NCLB tests. If not the end, at least scoring for quantity and quality (as found in PUP, Winsteps and Amplifire) would be offered as an alternative choice to the current gambling.




Wednesday, September 12, 2012

Reading and Writing Preparation


The standardized testing being created by the two consortia (PARCC and SBAC) and Pearson is predicted to produce comparable test results among states (less cheating at all levels). The next level of assessment is standardized student certification: students passing the test will have predictable performance in future courses and in the workplace. Pearson, based in London, the largest education company in the world, over 70% of the revenues, is making great strides toward this goal (valid assessment must be coupled with appropriate instruction).

All of this must start with students being capable readers. The main problem found to date, indicated by students dropping out of high school, is that they were not good readers by the end of the third grade. They remained passive captives of the classroom until they could no longer stand the boredom; or the prospects of passing a NCLB standardized test to pass a course seemed impossible; or they were missing, disenrolled or triaged to increase the school’s rank.

To actually know something, to make it your own, requires that you do something with that knowledge or skill. Visualization works with both real and imaginary situations. Performance requires doing the real thing or working with a close mock up. Questioning, answering, and verifying do this for the self-educating student.

The passing rate on NCLB standardized tests makes a good example of how this can work out in a complex system of state and federal agencies. Some states visualized an ever-increasing rate of passing, but with each increase a bit smaller then in the previous year. The idea was to maintain the minimum annual yearly increase until the politicians in Washington, DC, would correct the initial error of requiring the impossible standard of 100% of students passing in 2014. Of course, that did not happen.

The politicians are now five years behind in correcting their original errors, and state departments of education must now explain why their new tests are producing such low passing rates, with about the same low scores. The passing rate was a Ponzi scheme at the state level based, as much on luck (the right count scored multiple-choice test), as on student ability. We really need a way to educate individual students and evaluate each for what each one knows and can do. Teaching to the middle of the class and assessing with the popular right marked scored multiple-choice test cannot do that when the result is low scoring tests. The vision was doubly faulted: instruction and assessment.

The Educational Software Cooperative (ESC) was formed in 1992 (incorporated as a non-profit in 1994) to provide a means for teachers and software authors (who were mostly teachers) to empower individual students to become proficient learners (both quantity and quality, knowledge and judgment, are important) regardless of the academic environment of their schooling. Software lessons can also be non-judgmental and have everlasting patience.

ESC members have continued to teach by way of ever changing software that improves in the level of thinking addressed with every advance in technology. The oldest currently listed downloadable reading software is Animated Alphabet for Windows. It teaches letter sounds and vowels with silly animation to prevent boredom when learning at the lowest levels of thinking. Directions are read aloud for pre-readers. The level of thinking is in balance within instruction, learning, and assessment.

The most advanced downloadable reading software listed, such as AceReader, promote reading, fluency and comprehension. I have never seen such a course in all of my schooling. I took the Evelyn Wood’s Reading Dynamics course in 1967 in Honolulu, Hawaii, when on a two year sabbatical with the USDA, Agriculture Research Service. The most successful students in the class performed at phenomenal reading speeds when they were able to change from sub-vocalizing each word to forming visualizations from groups of words. You actually see, experience, what the author is writing about rather than remember the words used. It is a neat experience. Successful students must be efficient and proficient readers. Yet how often do students in public schools get a chance to learn to read at this level, or even be aware that it exists? It was a new experience to me.

Another example of replacing words about a subject with the experience of doing is Smart Science. Sub-vocalizers think they are reading well because they have never experienced a higher level of reading. Cookbook laboratory manual exercises, followed one word at a time (at lower levels of thinking), do not generate the experience of doing science (at higher levels of thinking). The words are a poor replacement for the real thing. Smart Science teaches the important process of science and scientific habits of thought by way of real virtual labs. This is reading, observing, thinking, and writing using all levels of thinking. It is much more than record keeping at the lowest levels of thinking.

The most recent offerings take advantage of the Internet to supplement individual student learning and perform classroom record keeping chores. Merit Software produces award winning software for the home and the classroom that teaches reading and writing. Essential Skills Software produces both CDs and online versions for use in classrooms that are closely aligned with state reading and writing standards. Please check the ESC list for others.

Teaching by way of software is now well developed. The Common Core State Standards has created an environment in which free enterprise can thrive. Good government has a positive effect that costs the taxpayers nothing. The two consortia however are struggling with computer scored essays (at the lowest levels of thinking comparable to human fast scoring), online assessment (with right count scoring at the lowest levels of thinking) and millions of taxpayer dollars.

The Internet pipeline for full online interactive assessment is needed to manage cheating at all levels. Once in place it can then be reversed and used for individual student instruction. At this point I would hope that all levels of thinking could then be accommodated. I also predict the dominate testing company may well become the dominate instruction company. The traditional classroom will no longer be needed. The creativity that produces new educational software will always be needed. 

Currently, to my knowledge gained in writing this post, student ability is still assessed within school courses. Software teaches as a supplement to, or a replacement of, a part of a course. Each year more teacher friendly features are included: record keeping, etc. Even the software I have named as examples have many more features than I have mentioned. Each author, teacher, has a unique approach to creating software. You must check out each offering for the one that fits your needs the best.

At some point software should have the same weight as a correspondence course. This is coming about as a natural experiment in what is now being called  “flipped” instruction. Students do the lower level of thinking portion at home by CD or online. They are then ready to question and discuss at higher levels of thinking, and to know where they may need help, in class.

This method of instruction worked very well in my remedial biology course. The textbook was presented with questions on the campus computer system to help students learn to read by questioning and relating (meaning making). Biweekly multiple-choice tests scored for quantity and quality (accurate, honest, and fair – no guessing required) promoted student development from passive pupil to self-correcting scholar. Today we can also add Winsteps and Amplifire (see previous post).

Wednesday, September 5, 2012

Your Choice of Multiple-Choice Testing


Doing PARCC, SBAC, waiver or no waiver? Your choice of multiple-choice makes a big difference in what you get for your money. Today you have a choice. You are not bound to the traditional, right count scored (RCS) version.

  • Traditional RCS multiple-choice works very well at the mastery level, 90% cut score; it is an easy way to score classroom tests, 75% average score and 60% cut score; it is gambling below a score of 60% (meaningless ranking where quantity and quality scores are identical).
  • Independent quantity and quality scored multiple-choice provides the same freedom for students to report what they trust they know and can do as when using short answer, essay, projects, and reports when scored at all levels of thinking. The quality score can range from that found on a RCS test up to 100%, independently from the quantity score: the number of right marks (both the examinee and the examiner know what the student knows and how well knowledge and judgment are used). No forced guessing or gambling is required, just fair and honest reporting.

You have several ways to implement the student empowering features of independent quantity and quality scoring.

  • Both right count scoring and Knowledge and Judgment Scoring (KJS) are featured in Power Up Plus (PUP). This is a classroom friendly implementation. It allows students to select either right count scoring (at the lowest levels of thinking) or KJS (at all levels of thinking). Research has shown that after two experiences with KJS, over 90% of students switch to KJS. It takes a couple of experiences for them to see and believe that they do better taking responsibility for what they know and can do than just marking a test and hoping for good luck. They then change study habits from memorizing random bits of non-sense (to hopefully match on a RCS test) to making sense of each assignment so they can now correctly answer questions they have not seen before.
  • Winsteps, the software many states have used on NCLB testing, contains a Partial Credit Model (PCM) analysis feature. Using item response theory (IRT), it produces the same scoring as classical test theory (CTT) in PUP. It calculates the unexpectedness of each student mark. PUP now colors the student and test performance charts using the unexpectedness values from Winsteps.
  • Amplifire by Knowledge Factor contains the most powerful implementation of independent quantity and quality scoring. Instead of mixing quantity and quality, half and half, for a test score (as is done in PUP and Winsteps), Confidence Based Learning (the forerunner incorporated into Amplifire) mixed three parts quality with one part quantity (knowledge). This is justified in high-risk occupations and in rigorous academic training (mastery). Amplifier, a patented instructional system, includes fast response coaching in such a timely manner that the seemingly impossible standard set by the high quality requirement can be met in a reasonable amount of time.

Today there is no reason to continue using traditional RMS multiple-choice tests when the average score falls below 75% and the cut score is below 60%. Below these points, the tests tell us nothing useful about student performance (other than a questionable rank on the test). From the same test, same preparation, same scanning, we can also get what each student actually knows and how well that knowledge or skill is used.

We get an insight into the level of thinking being used (teaching and learning) in the classroom. We get a student view of the test as well as teacher and test maker views. Misconceptions are distinguished from difficult questions. We get an insight into the development of the student, what levels of thinking are friendly and useful; and which students are taking charge of learning and reporting (highly teachable, meaning makers); and which students are still waiting for the teacher to teach, to test, and to tell them how many right marks they got gambling on a traditional RCS test.

All of the above test benefits are also available at the state department of education level. One of the difficulties of holding students, teachers, schools, and state departments of education accountable (see prior post) has been the use of a ranking system that has had little to do with what students actually knew or could do at the cut score. The cut score was often selected for political reasons. A valid 90% passing at a mastery level was as unsettling, as 24% passing because the scoring was too low, or 90% passing because the cut score was set too low. The additional information from independent quantity and quality scored testing reduces this problem.

The only change in standardized testing required is to offer Knowledge and Judgment Scoring, or PCM scoring, and RCS on the same test. My experience has been that this eliminates the politics of implementing a different scoring method. It is also crucial to allow students the freedom to make the choice that fits their development. This freedom is consistent with taking responsibility for selecting questions to mark when opting for Knowledge and Judgment Scoring or PCM scoring.

Wednesday, August 29, 2012

NCLB Accountability


The history of accountability from the school and to the state department of education level has been quit varied. Only after running this preposterous natural experiment for ten years is it being challenged in ways that may be effective in either bringing it to an end or in correcting its excesses. Congress created this absurd monstrosity by setting an impossible goal for all students to meet. It then reneged on its oversight responsibility to act in good faith to avoid many of the unintended consequences that occurred (it is now five years behind the time it should have acted on needed changes).

Self-regulation is a lofty idea. It has failed miserably for mortgages, for wall street derivatives, and in the futures market (all of which were presumably being regulated). The same can be expected in state departments of education that must come up with acceptable numbers to obtain federal funding. The two consortia (PARCC and SBAC) promoting the Common Core State Standards have the promise of serving as checks on one another. This is an expensive and ambitious political solution that may have its own down side depending on implementation. If all states will release actual student test scores there will be a way to determine how creative states are in determining passing rates.

The passing rate has been politically exploited in several states. New York is the prize example. Diane Ravitch posted, 21 February 2012, “Whence came this belief in the unerring, scientific objectivity of the tests? Only 18 months ago, New York tossed out its state test scores because the scores were unreliable. Someone in the state education department decided to lower the cut scores to artificially increase the number of students who reach proficient. No one was ever held responsible.”

Michael Winerip posted, 10 June 2012, “Though this may be the worst breakdown in 15 years of state testing, it does not appear that Florida politicians have any interest in figuring out who was responsible. The commissioner? Department officials? Someone at Pearson, the company that scored the writing tests?” Winerip reports further that, “The audit referred to lowering the passing score to 3 as ‘equipercentile equating’”. That is, the score was lowered until the same portion of students passed this year as passed last year. [As I am writing this, the commissioner resigned.]

As is the case with mortgages, derivatives, and futures, it is difficult, in most cases, to say if a crime has been committed or just very poor judgment was exercised, until the chain of events is carefully studied including the false "belief in unerring" test scores. My own explanation here is that research results and application results are not the same thing. In research you predict acceptable results. In application you examine the results for meaningful useful relationships (equipercentile equating to obtain the desired pass rate – any relationship between the ranking on the test and what students actually know or do is mostly coincidental near the cut score).

Has a crime been committed? In the case of New York, and other states that must now “explain” why student tests scores are dropping on new tests, I would say, “Yes.” For states that choreographered an almost perfect, ever slowing, increase in the pass rate for the past 8 to 10 years, the answer is problematic. It can range from outright cheating to self-deception. From equipercentile equating, to selecting test items that produce the desired results, the standard practice for classroom test score management.

On 18 May 2012, Valerie Strauss posted the white paper released by the Central Florida School Board Coalition. This lengthy paper details the unintended consequences and the downright sloppy test items used on their standardized tests. My own software, Power Up Plus (PUP) can pick such items out when run on a notebook computer in a matter of minutes. I am amazed that such items are used considering the millions of dollars spent on development and administration of their tests. I strongly suspect their developmental process.

Cory Doctorow posts, “The Test Item Specifications are the guidelines that are used to write the test questions. If the Science FCAT test is reviewed by the same Content Advisory Committee that reviewed the Test Item Specifications, then it probably has similar errors.” From my experience, a valid test item must assess exactly what it says (concrete level of thinking – what you see is what you get) or be an indicator of knowledge or skill of things in the same class (1 + 3 = 4 to assess addition of integers). Questions that have different right answers based on levels of thinking, socio-economic status, state, religion, politics, ethnicity, and current political correctness are not to be used. That is, stick to the topic, not to what the topic (agenda) or skill may be used for. Where a question is on topic but has different answers related to the above, it should stand. This is part of the broadening effect of education. A recent example is the Missouri constitutional amendment voted on yesterday to protect the religious rights of school children.

In the case of Florida, once again, faulty predictions were made based on some type of research. The entire system (instruction, learning, and assessment) was not fully understood or coordinated with disastrous results. And again, another state education official has resigned. Was this a crime or just a waste of millions of dollars and millions of instructional and learning hours?

 On 24 April 2012 Valerie Strauss posted the National Resolution Protesting High-stakesStandardized Testing that is based on the Texas Resolution Concerning High Stakes, Standardized Testing of Texas Public School Students. These two resolutions combined with the Florida white paper make a strong political protest that may take years to obtain results. The desired results are not specifically stated. We are back to the days of “alternative assessment”: Do something different; but doing something different must be at least as good as what we have, or we lose again; as with the authentic assessment and the portfolio movements.

As well intentioned as all the people are working on this assessment problem, most are still riding their safe steady tricycle, the traditional force-choice multiple-choice test scored at the lowest levels of thinking, that they were exposed to long before it was used for NCLB testing. It actually worked fairly well back then when the average test score was 75% or higher. It has failed miserably when NCLB test score cut scores dropped below 50% (a region where luck of the day determines the rank of pass or fail). Only until these people are willing to get off their old tricycles will they have any interest in getting on a bicycle (where students can actually report what the know and can do).   

Knowledge and Judgment Scoring allows students to individualize standardized tests, to select the level of thinking they will use; to guess at right answers, or to use the questions to accurately report what they actual know and can do. We only need to change the test instructions from, “Mark an answer on each question, even if you must guess” to “Mark an answer only if you can use the question to report something you trust you know or can do.” Change the scoring from “two points per each right mark and zero for each wrong mark” to “zero for each wrong mark, one point for omit (good judgment not to guess – to make a wrong mark) and two points for each right mark (good judgment and right answer)”.

We can now give students the same freedom as given with essay, project, or report to tell us what they actually know and can do. We also have the option to commend students as on other alternative tests, “You did a great job on the questions you marked. Your quality score of 90% is outstanding.” This quality score is independent of the quantity score. You can now honestly encourage traditionally low scoring students for what they can do rather than berate them for what they cannot do (or for their bad luck on traditional forced-choice tests).

The NCLB monster may be controlled by legal action (see prior post), political action, or by just changing its temper from a frightening bully to an almost friendly companion. Just select your breed: Knowledge and Judgment Scoring in PUP, Partial Credit Model in Winsteps, or Amplifier by Knowledge Factor.

Why continue just counting right marks; making unprepared liars out of lucky winners and misclassifying for remediation unlucky losers? We know better now. It no longer has to be that way. “Whence came this belief in the unerring, scientific objectivity of the [dumb, forced-response, guess] tests?” We need to measure what is important. We do not need to make an easy measurement (forced multiple-choice and essay at the lowest levels of thinking) and then try to make the results important (at higher levels of thinking). There is a difference in student performance between riding a tricycle and a bicycle. We cannot hold students responsible for bicycling if they only practice and test on tricycles.

Wednesday, August 22, 2012

Rasch Model Audit Finished


It was over two years ago that I took on the task to audit, make sense of, the Rasch model IRT analysis as performed by Winsteps, the software many state departments of education have used with NCLB standardized testing. It is now 46 posts later (they are being released on weekly intervals). The Winsteps software performs as advertised for psychometricians and test makers. No pixy dust is required. But the developmental environment it creates is problematic.

First is the collection of features provided to cull raw data to fit the perfect Rasch model. It creates a very freewheeling environment in which items and students are culled to get the results to “look right”. This is not a problem for competent experienced operators. In fact it is probably the best set of features for their work. But it is also an invitation to cheat in the hands of desperate state education officials needing results that “looks right” to obtain federal funding. It is an invitation that has been taken by politically motivated city and state officials (see prior posts that prompted the audit).

Early on some officials lost their jobs for making predictions based on research results that were not supported by application results. That no longer happens for a number of reasons. [Edit: It just happened again in Florida.] No one, to my knowledge, has been charged with a crime when test results and cut scores were manipulated for political gain; even when several million dollars were lost and millions of hours of instruction and learning time diverted to cramming for “the test” that at best produced a meaningless ranking, and worst, destroyed the validity of the results for instructional improvement and assessment of teacher performance.

A new threat to schooling is now being developed by predatory testing companies in collaboration with state departments of education, several of which were cheating. Two groups have formed (PARCC and SBAC). This is a good development. Each group can serve as a check on the other. One requirement being voiced outside these two groups is that the actual test scores must be published rather than the passing rate or a ranking of good, better, and best. Whenever the actual test scores are hidden there is no way to know what passing rates or rankings mean. However, the real threat is that, in the name of “formative assessment,” they want to increase standardized testing from once a year to three or four per year. One test has been destructive enough. More would only make things worse except for the predatory testing companies that will publish and score the tests and test preparation materials.

We have a problem. The federal government has shown its inability to improve education in the classroom by bullying schools with money. This failed program is now culminating in a political revolt, and perhaps, in legal action (see below). State departments of education have been too weak to resist being bullied: they needed money. Most public schools are operated as failing establishments by design (lesson plan, teacher presentation, and student performance rank [how well students adjust to the needs of the school’s administrators, staff, and teachers] at the lowest levels of thinking). Schools designed for success are organized around student learning at all levels of thinking. They adjust to the needs of their students. Failing school administrators are more managers than leaders. The result is their teachers are no longer free to be professional teachers (responding to their student's needs). In some states teachers now function as readers of the assigned daily lesson, as coach for "the test", and as day care servants (responding to the perceived needs of the bureaucracy).

There is one branch of government that I have yet to see get involved. That is the office of the state attorney general. That branch of government has the responsibility to prevent tax payers from being ripped off by unethical business practices. Predatory testing companies in conjunction with self-serving state education officials may be about to pull off the biggest scam yet: not one but three or four standardized tests per year per course that, for the most part, will be scored at the lowest levels of thinking (using traditional multiple-choice and high speed evaluated essays).

If they justify this use of multiple tests as “formative assessment”, it will be a patent lie; pure fraud; marketing a product under the guise of a current education fade. Formative assessment occurs in seconds in the heads of students functioning at higher levels of thinking. It occurs in the classroom in minutes between students and teachers (in person and through educational software). It does not occur in weeks or months by way of standardized tests (unless you are promoting the tests). Marion Brady sums the situation up in neat operational terms that everyone can understand: "Do not subject my child to any test that doesn't provide useful, same-day or next-day information about performance."

Are you, or do you know someone, in the office of your state attorney general who is interested in forming a group to act on this matter? Nine-Patch Multiple-Choice, Inc, for-profit, now has facilities to support this action. It can perform multiple-choice test analyses at both lower and higher levels of thinking (right count scoring and Knowledge and Judgment Scoring); classical test theory (CTT) and item response theory (IRT); and offers students the choice between being assessed by traditional guessing or reporting what they trust.  I would suggest that you submit a test data set of 25 to 50 questions and 120 to 300 students for analysis as an entry into the group. We need real test data. Not research or theoretical data. Once a group is formed Nine-Patch Multiple-Choice, Inc. can be re-registered as a non-profit, to give the mission even greater stability. Email rhart@nine-patch.com or phone 1-573-808-5491.

Currently there is no way to audit and verify test results from predatory testing companies or state departments of education. A proposed law could require that random samples be selected, and be examined by an independent company or the office of the state attorney general using appropriate software, that currently runs on a personal computer, and produces meaningful results in minutes rather than months (see next post). 

Richard A. Hart, PhD
Professor of Biology, Emeritus, NWMSU - 1990
Treasurer - 1992
Educational Software Cooperative
Treasurer and Founding Board Member - 1994
Educational Software Cooperative, Inc. (non-profit)
President - 2006
Nine-Patch Multiple-Choice, Inc. (for-profit)
803 Somerset Drive, Columbia, MO 65203-6436 USA
Email:      rhart@nine-patch.com,
Phone:     1-573-808-5491
Organizer - ????
Nine-Patch Multiple-Choice, Inc. (non-profit)

Wednesday, August 15, 2012

Student Quality


We need both quantity and quality, but if a choice must be made, quality generally wins, expect in current academic testing. The traditional right count scored, forced-choice, version of multiple-choice assessment ties quantity and quality together in one meaningless ranking. It extracts the least information from the answer sheets. Because of this, there have been many movements (fads) to improve assessment from alternative assessment, authentic assessment, portfolios, projects, and reports to actual oral and visual presentations. In the end, traditional multiple-guess has always won out for some very good reasons: cheap, fast, easy to do and highly reproducible results.

Multiple-choice assessment does not have to be this way. Just change the instructions a bit and you have assessment at all levels of thinking as well as cheap, fast, easy to do, highly reproducible and meaningful results. Allowing students to accurately report (on multiple-choice tests) what they trust will be of use in further learning and instructing is not something new in 2012.  Geoff Master from Melbourne, Australia, developed the partial credit Rasch model (PCM) that is included in Winsteps prior to 1982.

While teaching at Northwest Missouri State University, USA, along with several 1000 remedial biology students, I developed Knowledge and Judgment Scoring (KJS) in 1981 to obtain an individualized written report from each student that accurately assessed what each student really knew (from lecture, laboratory, and assignments) and on which further meaningful learning could be built. As one faculty member working with pre-med students put, “We know what they know and how well they know”. This method of scoring was crucial in providing the information needed with which to guide each student along the path from passive pupil to active, self-correcting, scholar. It made possible servicing a class of 120 remedial biology students more effectively and with less effort than 24 students in a class with “blue book” exams.

James Bruno made extensive studies in assessment at the University of California in Los Angeles. In 2005, Knowledge Factor patented an educational system (Confidence Based Learning – CBL, now Amplifier) based on his work with great success in the professional development and competency assessment area. Knowledge Factor sets the bar for quality at 75% or higher. KJS sets it at 50%. Traditional multiple-choice sets it at zero (passive scoring – when scoring the finished answer sheet) and at 25% for four-option questions (active scoring – when taking the test).

Both PCM and KJS produce the same test scores. They both also provide estimates of student quality. This illusive property is often discussed as only to be found in “alternative assessments”, alternative to traditional multiple-choice (for the majority of uninformed and un-relearning educational reformers). Quality by alternative assessment is very subjective. Quality by PCM and KJS is not. Quality by PCM and KJS is also highly reproducible.

PCM and KJS produced comparable quality indicators on a remedial general biology test for four students that had a 70% test score. The Student Normal (+) Output values on the table have been corrected by adding 25% to each value to match the Item Normal (+) Output value mean (see the full details on the 3 October 2012 post on the Rasch Model Audit blog, Rasch Model Student Ability and CTT Quality).

PCM and KJS Quality Indicators
Method
Student (70% Test Score)

26
37
40
44
KJS
81%
88%
88%
95%
PCM
68%
76%
76%
88%

These quality indicators cannot be expected to have the exact same values as they include different components. KJS divides the number of right answers by the total number of marks a student makes to estimate quality (% Right). The number of right marks is an indicator of quality. The KJS student test score is a combination of quantity and quality (PUP uses a 1:1 ratio that every student can understand). If a student elects KJS but ends up marking most of the questions, the KJS assessment automatically turns into a traditionally right mark scored test with no penalties (except for the traditional 3 out of 4 wrong when forced to guess).

Knowledge Factor (KF) uses 3/4 for judgment and 1/4 for knowledge when working with high risk occupations (it also uses three-option questions instead of four or five options). This makes sense when setting the value for quality (judgment) three times greater than for quantity (knowledge). The examinee either knows or does not know (and is then coached and trained to seek help). No guessing is allowed when only mastery is the goal. Allowing one airliner to take off directly in the path of one landing is not a good thing.

KJS and KF only see mark counts of 0, 1, and 2. Winsteps combines student ability and item difficulty into one PCM expected score. The perfect Rasch model, implemented by Winsteps, sees combined student ability and item difficulty as probabilities from zero to 1. A question ranks higher if marked right by more able students. A student ranks higher marking more difficult questions than when marking easier questions. The end result is two comparable, but not exactly the same, estimates of quality from the two methods of analysis.

Knowledge Factor optimizes assessment and instruction for mastery in high risk occupations. Winsteps, PCM, is optimized for psychometricians and test makers. Both can be used in the classroom where mastery and the development of high quality students is important, not just a topic of conversation (this is in contrast to just passing). It is in sharp contrast to the traditional failing classroom where instruction and learning are conducted at lower levels of thinking in preparation for NCLB standardized tests.

Knowledge and Judgment Scoring, as presented in Power Up Plus (PUP) is an adaptation of holding students sufficiently accountable that they develop the skills of the self-motivated, self-correcting scholar (question, answers, and verify). Facts change. The skills needed to learn and relearn only develop more with use in a non-threatening environment. PUP provides students with the opportunity to voluntarily select reporting what they trust when they are ready to do so (switch from lower to all levels of think). It does this by scoring both methods: traditional guess testing and KJS. Over 90% of the students I worked with made the change after the second exposure (after two times on their new risky bicycles, where they learned to balance [to be the judge of what they knew], they readily gave up their tricycles). This was a new and empowering experience for many students, “I can do this!”

I have promoted KJS for over 20 years. It provides much of the information now lost using traditional RMS tests. It provides the guidance needed to move students from passive pupils to self-educating high achievers (including the current fad generally expressed as 21st century skills – these skills have always been important for master achievers). But in a highly threatening environment created by federal government bullying, multiple-choice has been given a very bad name. The desire needed to risk, to relearn, that there are two very different multiple-choice assessment methods has been almost squelched.

Until KJS is offered on standardized tests, it still makes a great training ground for preparing for such tests as it makes very clear to each student, during the test (an effective formative assessment willingness to need to know moment), what each student has yet to learn (and what each teacher may need to “reteach” to students willing to learn).

When a student understands, he can answer questions he has never seen before. Students who made the switch in my classes also found they were also doing better in their other classes. Learning and reporting for your own empowerment is a lot more fun than learning for a classroom or standardized test conducted at the lowest levels of thinking (gambling for a passing score).

Both quantity and quality matter in alternative assessments, including PCM and KJS. They are more easily and less expensively assessable when multiple-choice is done right: PCM and KJS. Done right also promotes student development, to be a better learner: a weaver of relationships rather than a cataloger of isolated bits. Multiple-choice done right even guarantees mastery with KF.