Showing posts with label college debate. Show all posts
Showing posts with label college debate. Show all posts

Friday, November 19, 2021

Judge intervention

"[W]hen people argue they move back and forth between strong claims that are weakly defended and weaker claims that are strongly defended." ~ Ross Douthat, https://www.vox.com/21760348/trump-2020-election-republican-party-ross-douthat


What is judge intervention?


Let's start out first by explaining what judge intervention is NOT:


- Judge intervention is not when a judge prefers your opponents' argument to yours. This is, ya know, "judging."

- Judge intervention is not when a judge ridicules your argument. This is rude, but that's about all it is.

- Judge intervention is not when a judge jokingly says, "Extra points to anyone who can avoid saying the President's name," or some such. Silly, and perhaps annoying, but that's about all it is.

- Judge intervention is not when a judge declines to discuss their decision after the debate. In some cases, talking after the debate is disallowed by a tournament; in other case, it may just be the judge's personal preference. So be it.

- Judge intervention is not when a judge grants your opponent favorable treatment (e.g., more prep time than is allowed by the tournament, advice, extra speech time or do-overs, etc.). This is unacceptable behavior and should be reported to an adult in charge of the tournament.

- Judge intervention is not when a judge ignores the debate, falls asleep, or in other ways is derelict in their duty to be a full participant and educator. This is unacceptable behavior and should be reported to an adult in charge of the tournament.

- Judge intervention is not when a judge demeans you personally. This is unacceptable behavior and needs to be brought to an adult's attention, preferably your coach and/or the tournament.

- Judge intervention is not when a judge votes based on the school, sex, gender, sexual orientation, race, ethnicity, or socio-economic class of the debaters, and not the arguments. Let's point out that it's beyond unacceptable and could be illegal.

- Judge intervention is not when a judge creates a hostile environment of harassment or fear or sexual joking. Or implies wins or speaker points can be traded for personal favors unrelated to debate. Beyond unacceptable and definitely illegal.


As you can see, judge intervention is not a catch-all term for rude or inappropriate judging behavior. Judge intervention falls somewhere in the middle of the spectrum: beyond merely rude, but well short of truly inappropriate behavior. Judge intervention is perhaps in some cases murkily unethical, although it would be very, very hard to prove any specific instance as unethical and will probably never result in an overturned ballot.


In many cases, judge intervention is the inevitable result of how YOU debated; it is, in at least some key ways, an inevitable part of judging even good debates.


So, what it judge intervention? Simply put, judge intervention is when a judge has to use their own prior knowledge to resolve an issue in the debate. Consider the following exchange:


Team A: "X is a fact" [no evidence provided]

Team B: "X is not a fact" [no evidence provided]

Judge: "… fudge you all"


In this instance, neither team provided any evidence to back up this point. If X is a crucial fact for the debate, then the judge will have to insert their prior knowledge to resolve it. (If X isn't important, then most judges will happily ignore it as irrelevant so as to avoid intervening.) Of course, this example is not the only one where a judge might be forced to intervene:


Team A: "The sky is green" [no evidence provided]

Team B: "…"

Judge: "Come on, team B!"


While the judge should be open minded, this doesn't extend to being completely gullible. We expect everyone to bring at least some prior knowledge into the debate with them. To assume the judge knew nothing would actually be an exhausting way to debate; in every round, you would spend the bulk of your time explaining how the government worked, that people were bipedal and how they used language to communicate, and so on. To paraphrase the T.V. show Archer, "Why is there conflict in the Middle East? Well, a long time ago there were dinosaurs, then they died, and their bones are what your car uses for food." Giving every bit of backstory would slow debates to a crawl. The judge should be allowed to have some degree of common sense without being accused of wrongdoing, even though it's technically intervening to ignore Team A's insane green-sky claim.


There are in fact many instances in which we should actually expect the judge to intervene and reject your argument, even if the opposing team said nothing against it. Looking at Toulmin's model will help us delineate all the ways. Toulmin's model of an argument contains the following parts:


Claim: your argument

Evidence: your supporting facts or expert opinion

Warrant: the logical connection between evidence and claim

Relevance: the meaning of your claim to the debate


A judge might reject your claim because it's muddled and unintelligible; they might reject it because it contradicts claims you make in other parts of the flow. A judge might reject your claim because it's of the wrong type. There are four types of claims: (1) factual, "X is true"; (2) causal, "If X happens, then Y will result"; (3) value, "X is right and good" or "X is more important than Y"; and (4) definitional, "The word X means…" These types of claims are not interchangeable, so don’t go answering a moral question with a factual claim. Definitions don't really address cause and effect questions…


A judge might reject your argument because you provide no or weak evidence. What if you argue the economy is doing well and your evidence is several years old? What if your source is, unbeknownst to you, a widely discredited conspiracy theorist? There are some real gray areas here. In some instances, we might say, "The other team didn't point that out, so the judge merely needs to accept any evidence I provided at face value," but in other cases, that feels absurdly deferential to a team reading, say, neo-Nazi propaganda or other gibberish. The judge should accept most reasonable seeming sources at face value, even if they perhaps know of a specific flaw with the evidence or source, but what is a reasonable source? Judges may have to walk a very fine line here because, though they are no experts on the topic, they may know more than you do. The line I often draw is,"What would be reasonable for a student at your age and experience level to know?" Do I expect you to know an insane conspiracy theory website when you see one? Yes. Do I expect you to know how respected an academic at a good university is in their field? Not really. So I would throw out the conspiracy website, even if it's unchallenged, while allowing the academic—even if I know something specific about their work—if your opponent says nothing. But if your opponent does say something, I will definitely let my skepticism fly free. My initial, naive presumption you shouldn't know something at your age goes right out the window when your opponent knows it.


A judge might reject your argument for lacking a warrant connecting the evidence back to the claim; the evidence might appear senseless and irrelevant. But the judge might also reject some kinds of warrant that you do provide. We don't expect judges to be subject matter experts in every single field a debate might touch on: they are not scientific experts, legal scholars, linguists, logicians, political theorists and strategists, pollsters and statisticians, and foreign policy gurus and intelligence analysts, all rolled into one. It's quite possible our explanations are too "inside baseball" for a lay judge to follow. Debating is a communicative act, so we need to do a bit of audience analysis. Our explanations, our reasoning, and our warrants need to assume that we are persuading an intelligent but lay audience; after all, our judge is not a subject matter expert but a smart generalist. In other words, yes, it's reasonable that a judge could reject a warrant you've provided that's perfectly intelligible and correct simply because it's over the judge's head—and it's reasonable to expect you to know it would have been.


A judge might believe everything about your argument but simply not get why you believe it's relevant to the debate.


In short, there are many, many reasons why a judge may intervene in a debate. Rather than the problem being that judges intervene, the problem is that you think it's an aberration. Judge intervention isn't an aberration. It happens all the time. It's the norm. It's a normal part of human decision-making to use one's prior knowledge and common sense to evaluate arguments in holistic ways. It's called critical thinking, yo. Judges don't accept everything you spoon-feed to them, and if they did, they wouldn't be able to reach decisions. A truly blank slate judge would be like a computer program that fritzes out before returning an answer; a spinning rainbow; a blue screen of death. Judges are thinking human beings with their own life experiences and assumptions who are trying to be—but not always succeeding at being—open minded to your arguments.


The problem with judge intervention is not that it happens in the first place but that you don't know how to use that fact to your advantage. A great debater knows how any judge, any human being, likes to reason; knows when, why, and how a judge is likely to intervene; and uses that knowledge to win. To be clear, I'm not talking about judge adaptation (changing your speed, style, or argument selection in order to appeal to the judge's preference). I'm talking about using the human psychology of decision making so as to find ways to "stitch" the inevitable gaps in arguments so that the judge won't notice them as flaws.


A winning debater is ultimately a good storyteller, and those stories operate on the micro level (warrants and reasoning) and on the macro level (relevance and voting issues). Those stories are crucial to winning over a judge. Judges need cohesive stories to justify their ballot, and if you don't supply it, the judge will. While you might not do a perfect job on every issue and position in the debate, telling a good macro story enables you to persuade the judge that you are, overall, winning. Likewise, not every claim or piece of evidence is perfect, so a good micro level story enhances your credibility; you paper over any holes in logic by demonstrating an overall cohesive understanding of the disadvantage or the case.


Lacuna in your arguments and direct contradiction by opponents are just inevitable in close debates. The latter, because your opponents are good; the former, because under the pressure of a close debate, you won't get to everything. There will be holes; the judge has to intervene to pull together a cohesive picture.


Consider a close football game, where the final score difference is less than 7 points between the teams. Either team could have won. A bobbled pass that was caught, a different penalty call, a timeout left—any of these could have made the difference in the game. Both teams had a near 50% chance of winning.


So too in debate. In a blowout debate, you are able to put together full, complete arguments and full, complete strategies. Your opponent isn't able to contradict you and isn't prepared to put together a competing narrative. You perhaps have a 90% or 95% chance of winning; there is but a slim chance that your opponent successfully throws a Hail Mary and makes a lucky good argument or that your judge happens to find one of your arguments completely unbelievable (i.e., you've unknowingly used a conspiracy website, so far as the judge is concerned). Pure judge intervention like the latter is rare, but it's just a part of human reasoning. Take the 90% or 95% chance of winning and be glad. There's no way to be 100% airtight in front of a thinking human being with experiences different than your own. They will have different assumptions and thought processes than you do. Mostly airtight is good enough.


A close debate, however, is like the close football game. Any play or argument could make the difference. Your efforts to complete every argument, although valiant, are impossible under the time pressure and constant barrage from your opponents. There will be holes. Therefore, the most important thing you can do is to tell good stories to paper over any holes. You want the inevitable judge intervention to slide your way. You won’t be able to increase your odds to 80% or 90%, but you can perhaps up them to 60% from 50%. In this instance, judge intervention isn't a bad thing, either; it's just part of the round being so close, so use it to your advantage.


When the round isn't close, out-of-clear-blue-sky judge intervention should be rare. It stinks, but all you can do is build your arguments as carefully and as air-tightly as you can—and learn from judges' rare interventions how to improve upon your techniques.


When the round is close, and judge intervention become inevitable, do your best to push the judge to intervene for you. Do this by using solid evidence, being clear and consistent in your positions, and most of all, telling compelling logical stories and impact analyses. Credibility and persuasion matter.

Sunday, March 24, 2019

Scheduled elimination tournament calculator

Single-elimination tournaments have one particular flaw: The tournament can only break powers of 2--i.e., 2, 4, 8, 16, 32, etc., teams can make it into the elimination rounds, unless the tournament decides to do a partial elimination round. For example, let's say the tournament decides to break 20 teams. So, the partial elimination round will involve eight teams debating, with twelve teams sitting around and waiting for two hours. The eight teams debating turns into four teams advancing to the first full elimination round, plus the twelve teams that sat around, making for a perfect bracket of sixteen. It works, but... meh. I don't like that the majority of the elimination-qualified teams did nothing for a whole round--it's kind of unfair that they could scout, plan a new strategy, or go get a nice meal and relax. And I especially don't like it, as a tournament director, that instead of just doing one big elimination round right after preliminary rounds and being done with it, I've got to drag it out into two smaller rounds. Let me explain this one a bit.

Ideally, a tournament breaks exactly one-third of its preliminary teams into elimination rounds. This is the ideal because preliminary rounds use one judge, elimination rounds use three, and well, you get the math. Assuming I have just enough judges for prelims, then breaking one-third of teams will use up all my judges perfectly for the first elimination round. The tournament director can say to every judge, "You must stick around for at least the first elimination round. I need everyone. Then I will start to dismiss judges whose schools have been eliminated." It works out brilliantly if every judge is used in elim round 1, half are needed in elim round 2, a quarter are needed in elim round 3... Smooth and simple.

Now consider the 20 teams breaking to elimination rounds problem. That means I have about 60 teams in prelims, and therefore 30 judges. In the partial elimination round, eight teams debate, so that is four rounds... therefore I need to use twelve of my 30 judges. In the first full elimination round, sixteen teams debate, so eight rounds, meaning I need 24 of my 30 judges. Notice how awkward and weird this has become? Some judges must judge both elim rounds--whom to pick? No one can go home until the first full elimination round is through, so that requires every judge to stay an extra two hours. Many of those judges will have nothing to do for the first two hours--I don't need them for a round. They just have to wait. Is there a better way to break a number of teams that isn't a power of 2?

Double-elimination tournaments can take on any even number of teams, so the above case of 20 is no particular problem, but they run into a different problem very quickly: the double elimination rule usually produces odd numbers of teams during the tournament for some rounds. If the tournament is run with brackets, then you can see how the math works quite easily. With 20 teams to start, ten will be undefeated and ten once-defeated after round 1. After round 2, five will be undefeated, ten will be once-defeated, and five twice-defeated and eliminated. That leaves fifteen teams and the perennial problem: some team has got to get a bye in round 3. Round 3! This is less fair than a team getting a bye in the partial elimination round in the single-elimination tournament. Is there no other option?

My proposed solution is the scheduled-elimination tournament. The plan is quite simple:
  1. Do not eliminate teams that are undefeated.
  2. You must eliminate teams that are twice-defeated.
  3. Decide which once-defeated teams to keep based on speaker points or preliminary seed.
  4. Always keep an even number of teams.
  5. Undefeated teams must debate undefeated teams; once-defeated teams must debate once-defeated teams; one pull-up is allowed.* (see note below for fun substitution!)
In practice, a scheduled-elimination tournament would look quite similar to a double-elimination tournament. Any (even) number of teams could break. There's an undefeated and once-defeated bracket going on in elimination rounds, just like in a double-elimination tournament.* (not necessarily--see note below!) But in many ways, the scheduled-elimination tournament is more similar to a single-elimination tournament: losing one round makes a team eligible for elimination. The tournament could decide to keep most of the once-defeated teams around, or eliminate most of them. It's up to the tournament. A once-defeated team might stick around to win the tournament, but only if the team had high enough speaker points or preliminary seed in order to never be eliminated. This being a mathematical problem, I made a graph to illustrate.




The single-elimination option (green) and the double-elimination option (red) create a lower and upper boundary on possibilities for the scheduled-elimination tournament (anything in the gray area--and yes, I chose gray for its symbolism). So long as the tournament keeps the remaining number of teams in the gray zone, then it has abided by condition 1 and 2 that I specified above. The gray zone represents all the once-defeated teams in the tournament. If the tournament cuts closer to the green curve, it eliminates most of the once-defeated teams. If the tournament goes closer to the red curve, it keeps most of the once-defeated teams. As you can see, the red curve is flat at first--no one in a double-elimination tournament is eliminated after only one round.

You might wonder what the two marked points, (3.32, 2) and (6.16, 2), are. This represents how many rounds each type of tournament will need to have, because two teams remaining leads immediately into the final round. The single-elimination tournament needs to have 3.32 rounds, plus the championship round. About one-third of teams participate in the partial round (eight out of 20), thus the 0.32, then the first full elim round will be sixteen teams, the second round will be eight, the third round will be four, and the fourth round will be two teams--the championship round. Four rounds, plus a partial. Similarly, the double-elimination tournament will need to have 6.16 rounds, plus the championship round. That means you're looking at seven or eight total rounds, including the championship round, depending on how the byes and pull-ups go. The scheduled-elimination tournament would have more than the four plus partial (so really five) elimination rounds of the single-elim tournament but fewer than the seven elimination rounds of the double-elim tournament.

There is considerable choice in that gray zone for how to run a scheduled-elimination tournament. One option would be to run what is almost a single-elimination tournament--but with no partial elimination round to start. Here's an example:


The tournament starts with 20 teams entering: (0, 20). Ten team remain after round 1: (1, 10). (The round itself is really the line segment from (0, 20) to (1, 10): starting with 20 and ending with 10 is round 1's effect.) After round 2, there are five undefeated teams, but the tournament keeps one once-defeated team for a total of six teams: (2, 6). There could be as many as three undefeated teams after round 3, but the tournament keeps four teams on just in case: (3, 4). After round 4, only two teams remain (4, 2), and then round 5 will be the championship round between those two teams. By keeping on perhaps two teams that are once-defeated (after rounds 2 and 3), this method eliminates having the partial elimination round where twelve teams sit around pointlessly, and it has made managing the judging pool much, much more predictable. The tournament will use: 100% of its judging pool for round 1, 50% for round 2, 30% for round 3, 20% for round 4, and 10%--the final three judges--to decide the championship round.

Another option is that a tournament could basically hew as close as possible to running a double-elimination tournament, yet avoid the problem of byes, by using a scheduled-elimination tournament. Here's an example:



As you can see, this tournament eliminates no team after round 1 (all twenty remain), eliminates six teams after round 2 for fourteen remaining, eliminates four teams after round 3 for ten remaining, eliminates four after round 4 for six remaining, eliminates two after round 5 for four remaining, and eliminates two more after round 6 for two remaining. The championship round will be round 7 between the two final teams. (In practice, because of pull-ups, this could potentially violate my second condition in the list above--some twice-defeated teams might stay in. The tournament should probably not cut quite so close to the red curve if it wants to respect this condition. The sequence 20-20-12-8-4-2 might be better in this regard.)

This option lengthened the tournament by two rounds compared to the previous option, but the trade-off is that this tournament eliminated almost no once-defeated teams. In general, it's possible that a once-defeated team goes on to win a scheduled-elimination tournament (e.g., they defeat the remaining undefeated team in the final round and beat them on points) if somewhat unlikely. It's also possible that a once-defeated team survives the cut after round 3, yet even though it wins round 4, the team doesn't survive that post-round 4 cut because its speaker points aren't high enough. This outcome seems reasonable enough to me. Being eliminated based on one loss and low points seems fine to me, although I would generally want my tournaments to stay closer to the two loss and done side of the gray zone. But--it's up the tournament to decide what makes sense for their goals and available time and judges.

I think I would add one other condition to a scheduled-elimination tournament, for a total of six. Condition #6 is: "Once the tournament begins eliminating teams, never increase the number of teams cut after a round above how many were cut after the previous round." In other words, the curve of cuts should flatten out. In the example graph immediately above, the cuts go: -6, -4, -4, -2, and -2. This seems reasonable and straightforward. It seems beyond silly to have the cuts go: -6, -8, -4, -2, -4, -2. Put them in a more sensible order.

I've made the applet available for you to use here: https://ggbm.at/dx9kgnbb. You can change the number of teams, and move the number of teams remaining after each round up or down. The line segments between each round will only show if you meet my sixth condition of eliminating fewer (or the same) number of teams after each round than the previous round.

* Fun addendum: This condition can be swapped out for a different one. The teams do not need to debate within brackets--i.e., several undefeated teams could debate once-defeated teams--so long as no cut is ever more than one-half of the teams remaining (which just seems reasonable and fair). The "no-more-than-half" rule can be substituted for the bracket condition without any risk of violating the first condition to not eliminate undefeated teams. The proof of this is fairly elementary, but let's think through an example first. Say 100% of the teams still in the tournament are undefeated (because you eliminated all the once-defeated teams). After one additional round, 50% will be once-defeated. That turns out to be the worst possible case, so the "no-more-than-half" rule keeps us on the happy side of condition #1.

Let's do this more conclusively with a bit of algebra. Say x% of the teams were undefeated and y% were once-defeated (and obviously x+y=100). If x>y, then y undefeated teams might debate y once-defeated teams, leading to y% of teams remaining undefeated. The remainder of undefeated teams, x-y, will have to debate themselves. So (1/2) (x-y) will also be undefeated through that pathway. That means we have y + (1/2) (x-y) undefeated teams, which simplifies to (1/2) x + (1/2) y, or (1/2) (x+y). Since x+y=100, that means 50% will be undefeated. The worst case scenario is that as many as 50% of the teams are undefeated. Never eliminate more than half of the teams remaining after a round, and the bracket condition can be dropped. It's a huge benefit to be able to drop it! This makes many more rounds possible--so you can avoid schools debating themselves, or opponents debating each other multiple times, until the very end of the tournament. Yay!

Wednesday, December 12, 2018

Tabulation software

Hi all,

I've been thinking about how to run tournaments for many years and publishing articles on it. My published ideas have ranged from geographic mixing, logit scores, and new methods for strength-of-schedule pairing and constrained side equalization assignment.

I've finally gotten around to putting all the ideas into a single, programming-ready document. I'm putting it out there as a Creative Commons Attribution (BY) license, version 4.0. Please feel free to use any ideas contained herein, as long as you attribute me.

Saturday, April 7, 2018

Experimental verification of the logit score

One method for ranking teams that I introduced to the debate community is the logit score. The logit score is derived from a logistic regression. The logit score combines a team's record, speaker points, and its opponents' strength into a single number. Because the logit score factors in record and points, it is performance-based, but that record is adjusted by opponent strength, making the logit score more fair than record alone. A win against a good team is "worth" more than a win against a weak team. If you take the worst opponent a team beat and the best opponent it lost to, and average those together along with the team's average speaker points, then you're approximating the team's logit score. Due to how the logit score is calculated, it is the likeliest team strength that explains its results: its record and its points.

I had previously looked for empirical support for the logit score in a college debate season. I took the real results for the entire season and used them to calculate each team's logit score. I then used those to retrodict the winner in every single match-up that had actually happened, with the higher ranked by logit score team retrodicted to win the round. The logit score did this better than every other ranking method I also tested, slightly edging out median speaker points, and doing better by a goodly margin than the win-loss record. Despite this success, there was the nagging concern that the logit score was being derived from an entire season's worth of information. This empirical support could not show if the logit score would work for a single tournament.

Therefore I set out to do an experiment. I created a simulation tournament in a program, and ran and re-ran it hundreds of times. I tested various tournament conditions, from random prelims to a typical method of power-matching to pre-matching (like a round robin). I looked to see whether in these kind of conditions--using only the information available in a tournament--the logit score fared as well in comparison to record-based rankings and to speaker point-based rankings.

The results are that, in any condition, the logit score is a vast improvement on the win-loss record, but not quite as good as speaker points. It may surprise people to realize that speaker points, even though they vary considerably from judge to judge, are the best information to rank teams. A team's median speaker points isn't affected too much by one judge. Speaker points are rich data when you only have six or eight rounds to rank a team.

However, I believe many in the community would not prefer to use speaker points alone. If nothing else, ignoring wins and losses gives a perverse incentive to teams to speak pretty and ignore winning key arguments. The logit score is a solid, thoughtful compromise. The logit score is based on both wins and points, so there's no perverse incentive to ignore key arguments--nor is there an incentive to ignore effective, mellifluous communication. Although the logit score is slightly less accurate for a single tournament than speaker points alone, the logit score is far more accurate than win-loss record is. The logit score is, in other words, a vast improvement on the status quo method--a compromise in name only.

Saturday, February 11, 2017

The Logit Score: a new way to rate debate teams

I recently published an article on a new debate team-rating method I invented, called the logit score. I hope the logit score will take its place among win-loss record, average speaker points, median speaker points, opponent wins, ranks, and so on as an effective way to rate (and thus rank) debate teams at a tournament.

What is the logit score?


The basic idea is simple: the logit score combines win-loss record, speaker points, and opponent strength into one score using a probability model. In other words, the logit score is the answer to the question, "Given these speaker points and these wins and losses to those particular opponents, what is the likeliest strength of this team?"

Let's take a step back and acknowledge a truth not universally acknowledged in debate: results should be thought of as probabilities, not certainties. A good team won't always beat a bad team--just usually. Off days, unusual arguments, mistakes, and odd judging decisions all contribute to a slight risk of the bad team winning. The truly better team won't always prevail. That means actual rounds need to be thought of as suggesting but not definitively proving which team is better. Team A beats team B. Team A is probably better, but then again, they could have had off day, been surprised by a weird argument, or had a terrible judge. If team A got much, much higher speaker points, it was very likely the better team. If team A only edged out team B by a little bit, then the uncertainty grows.

That's where the logit score comes in. Estimating team A's actual, true strength depends on putting together all of those probabilities and uncertainties into one model. I won't get into the specifics (the details are in the article), but the basic idea is using a logistic regression to put the probabilities for wins and losses to specific opponents as well as specific speaker points received together. The logit score for a team means: "If team A were estimated to be stronger, these results would be a bit more likely, but those other results would be far less likely. If team A were estimated to be weaker, these results would be far less likely, even though those other results would be a bit more likely. This logit score is the proper balance that makes all the results most likely overall." Because it factors in all the results in one probability model, the logit score isn't sensitive to outliers: unusually high or low speaker points, losses to outstanding teams, and wins over terrible teams don't affect the logit score much at all.

Does the logit score have any empirical results to back it up?


Yes. This is the bulk of my article.

I took a past college debate season, used those results to give every team a logit score, and then looked to see how well logit scores "retrodicted" the actual results in a season. That is to say, how often did the higher logit scoring team win rounds against the lower logit scoring team? As a baseline of comparison, I also did the same kind of analysis by ranking the teams by win-loss record.

The logit score rankings got slightly more rounds correct than the win-loss record rankings.

The slightly higher accuracy is not, on its own, a reason to rush to adopt logit scores. It merely proves that the logit scores aren't doing anything crazy. For the most part, the logit scores reshuffles teams ever so slightly with their nearest peers. The moves are slight ups or downs, not drastic shifts.

The real reason to consider using logit scores is that (a) they are less sensitive to outliers, which can matter a lot for a six or eight round tournament; and (b) they factor in more information. Win-loss records only use speaker points as a tiebreaker; it's secondary. Measures of opponent strength usually come third. In other words, a team with a really tough random draw and goes 4-2 as a result of dropping the first two rounds might miss out on breaking if no 4-2s break--win-loss record comes first and opponent strength won't factor in in that scenario. The logit score on the other hand--because wins, points, and opponents are all factored in at once--could reflect that this team is in fact very strong because it only lost two rounds to very good opponents. (See how important it is to be less sensitive to outliers?) More information also rewards well-rounded teams: those that win rounds on squeakingly close decisions and don't receive great speaker points are penalized more under a logit score system than a win-loss-then speaker points-system.

Sunday, January 26, 2014

Mutual judge preference

I read Jim Menick's description of mutual judge preference (M.J.P.) and was surprised about how he would assign judges (the same idea by Mr. Menick also appeared in the Rostrum, vol. 88, issue 2, fall 2013). For those who do not know, mutual judge preference gives every debater at a tournament an opportunity to rate every judge, usually in six categories from 1, most preferred judge, to 6, a strike. After each pairing is set, the tournament uses these ratings to assign judges to each round.

The part that was surprising was how Mr. Menick would rank the ratings, as it were. In order, he would rank a 1-1 (where both debaters rate the judge as 1) as the best option, of course, followed by 2-2, 3-3, 4-4, then 5-5. Then he would circle back to rank a 1-2 as the next best option -- where one debater gives a judge a rating of 1 and the opponent rates that same judge as a 2. Then the order would presumably go 2-3, 3-4, 4-5, then 1-3, 2-4, 3-5, and finally 1-4, 2-5, and 1-5. Mr. Menick's stated objective is to maximize mutuality, that is to say, any possible judge assignment with a difference of 0 in ratings (even a 5-5) is preferable to any judge assignment with a difference of 1 in rating (e.g., a 1-2).

Huh?

I had always understood the rankings would go something like this: 1-1, 1-2, 2-2, 1-3, 2-3, 3-3, 1-4, 2-4, 3-4, 4-4, 1-5, 2-5, 3-5, 4-5, then 5-5. I thought the goal is to maximize each debater's judge preference but not at the cost of screwing the opponent, and thus, a 1-5 judge assignment is ranked below a 4-4 judge assignment. Am I wrong here? Is this not what most people assume mutual judge preference to mean?

As a debater, why do I fill out the judge preference sheet? Under his system, my five most likely options are to get a judge I rated 1, 2, 3, 4, or 5. That seems like a lot of time to invest in considering the judging pool for... not any benefit to me, the debater. Under Russell's system, my six most likely options are to get a judge I rated 1, 2, or 3. That seems worth the time, to get the best half of the judging pool. Actually, if I were a debater and knew Mr. Menick's system was being used, I would not even both consider filling out the judge preference sheet at all. Remember, he points out that if only one debater has filled in the sheet, the judge assigned for that pairing defaults to that debater's picks. So consider the options:


In the two highlighted scenarios, I am better off NOT filling in the preference sheet. Only in the scenario that my opponent and I disagree about every judge am I better off filling in the sheet. How likely is this to happen? Mr. Menick seems to think that the main problem is debating style leading to incompatible judging preferences. I would agree that this happens a lot. But there are also terrible judges out there whom nearly every debater would agree are bad, so the last kind of scenario in the table is truly rare. As a competitor, I am going to hope that my opponents -- even if they have a different style -- will protect me from the truly awful judges and NOT bother to fill in my own judge preference sheet. After all, as the slightly-disagree scenario shows, filling in the sheet only WORSENS my chance of getting a judge I like!

Mr. Menick argues his prioritization of mutuality will lead to debaters getting judges they do not like and force them to adapt. Sure, that is likely true. But why bother with ratings at all? Why not just randomly assign judges? Randomly assigning judges forces judge adaptation at a lot less effort.

Before M.J.P., the tab room decided who the good judges were, and assigned the best judges to the break rounds. M.J.P. was just supposed to avoid the tab room's bias on this issue. So why is M.J.P. a rationale to foist judges I have said I would hate onto me?

Update:

Mr. Menick posted again about M.J.P. I cannot tell whether he was directly responding to me or not, since there were no specific references. Anyway, I think he makes a good case for adapting to lots of different kinds of judges. Which continues to baffle me because it is an argument against any form of M.J.P., not really an argument for his particular form of M.J.P. In other words, if I agree with him, I really ought to prefer random judge assignment, period.

Thursday, November 14, 2013

Debate across the curriculum

There are two interesting curricula I have run across recently. Stanford has an "Reading Like A Historian" curriculum, available free of charge: http://sheg.stanford.edu/rlh. The goal is to ask students to read primary sources carefully, evaluate each speaker's motivations and claims, and arrive at a nuanced, triangulated interpretation of events. They have had excellent results so far.

Sounds a lot like debate to me!

Another curriculum I ran into is Deanna Kuhn's program, which is a philosophy class for middle schoolers. They set this up as an experiment: some kids did philosophy through debate, while the control group kids did philosophy through lecture, reading, and writing alone. The results were quite impressive.

I've been a fan of Deanna Kuhn for at least a decade. She has a new book, Education for Thinking. Here is an excerpt of her writing (but not necessarily her book) from her website:
But aren't children naturally inquisitive? Are inquiry skills something that really need to be developed? The image of the inquisitive preschool child, eager and energetic in her explorations of a world full of surprises, is a compelling one. But the image fades as the child grows older, most often becoming unrecognizable by adolescence, if not middle childhood. What has happened to the "natural" inquisitiveness of early childhood? In part its nurturance into adolescence and adulthood rests on a set of values that parents and teachers must convey and support. But equally important is the channelling of this inquisitive energy into development of the cognitive skills that make for effective inquiry. The skills originate in early childhood, with achievement of the epistemological understanding that knowledge originates in human minds, is fallible, and has the potential for disconfirmation in the face of evidence. Only then does the coordination of theories and evidence that is a hallmark of authentic scientific inquiry become possible. In sum, the so-called "natural" curiosity that infants and young children show about the world around them needs to be enriched and directed by the tools of scientific thinking.

Her book was interesting. She described several investigative tasks; the music club one was easily replicable from her thorough description and sounded quite neat. She also had some interesting things to say about how to structure an introductory debate activity that would be neat for coaches who work with very novice debaters. Dr Kuhn had a key point about bootstrapping: when students knew little about the context and had few critical thinking skills, the teacher needs to have good structure set up so the students can bootstrap their way into both subject and skills.

Dr Kuhn's clearest call for reform was to create a specific focus across all classes on critical thinking skills. As she wrote in the book,
By examining causality in a biological context, a geographical context, a mechanical context, an interpersonal context, a sociological context, and any number of other contexts, students begin to understand and appreciate features of causality itself.

I think the goal of education is not information transfer, per se, but the ability to make a sustained argument, to critically evaluate sources and ideas, and to learn to investigate systematically (which could be a scientific experiment or a comparison of all available historical sources).

Putting the child at the center -- by asking him to debate an issue, to do an experiment, to discover a mathematical idea -- has two main benefits. First, he is actually spending his time on the important task itself (instead of endlessly preparing with rote memorizing for the day -- way off in grad school -- when he might be allowed to investigate anything on his own).

Second, the child is far more engaged in his work, because he has been given agency and purpose. Purpose: answer this question. Agency: it is up to you to answer this question, with some help if you need it, but no one is going to do it for you!

Wednesday, October 23, 2013

Debate topics

I always wished I had gotten a government reform topic, but I never did in my eight years. It seems like these are never used for policy debate. Reform topics seem like these would be great for Public Forum or Parliamentary debate.

On elections:

  • The U.S. House of Representatives districts should be drawn by non-partisan commission, not legislatures.
  • The U.S. House of Representatives should be elected using multi-seat, proportional districts. (I also wrote an earlier post here.)
  • The U.S. Senate should be elected with an open primary system, allowing the top two primary finishers, regardless of party, advance to the general election.
  • The Federal Election Commission should ban private funding of elections, substituting comprehensive public funding.


On structure:

  • The U.S. Speaker of the House should not have the power to delay a vote on a motion if a substantial minority of Representatives desire the vote. (abolish the "Hastert rule")
  • The U.S. Senate should abolish the filibuster.
  • Congressional pay should be suspended if Congress is unable to pass a budget.
  • The U.S. Congress should only be allowed to pass a spending bill if it contains a suitable increase to the debt ceiling.
  • The U.S. Congress should be required to take an immediate up-or-down vote if the President proproses raising the debt ceiling.
  • U.S. Supreme Court justices should be limited to 18-year terms (also here).
  • The U.S. Supreme Court should no longer have the power of constitutional review (also this, this, and this).
  • Amendments to the U.S. Constitution should require a substantially lower threshold for passage.


Substantive:

  • The U.S. federal government should not use block grants to administer federal programs.
  • The U.S. should increase the federal gasoline tax.
  • School funding should be collected and dispensed at the state, not local, level.


I'm sure there a lot more fun topics on general government reform, but these are the ones that occurred to me off the top of my head. Please post additional ones in the comments.

Saturday, October 19, 2013

Scrivener

I must endorse Scrivener, a word processing/book writing program. I've been using it for about a year, and it's fantastic. (And I'm not being paid or being given anything by them.)

The program is available for Mac and PC. It costs $45 for Mac, $40 for PC (and knock about 12% off for the education discount). It's easy to use, and far, far superior to Word for composing long-form files. Three specific uses have occurred to me: 1) for teachers writing a test bank, 2) for math teachers writing a problem set, and 3) for debaters to go paperless (and yes, I think it's better than going paperless by using Word templates, but it depends on the debater's personality and the squad setup).

But first, I'll give you a basic overview of the program. The screen layout is customizable, but this layout shows you three elements of the program: the Binder (left-hand pane), the editing window, and comments/footnotes (right-hand pane).


The document editing window is just like Word's draft or online layout view. The Binder allows you to sort your work into subdocuments. You can go as many levels deep as you want to. Note the research folder -- more on this later. The right pane is the Inspector. Footnotes, comments, meta-tags, and even more tools to annotate your work as you go. Note the Compile button in the top center. This is the "print" function, but unlike Word, there is a lot of control over what you print. You decide which subdocuments to print, whether comments or footnotes print, etc.

For teachers writing a test bank


I have ten versions of the chapter 2 test saved on my computer in Word. I'm not sure how exactly they differ. I rotated some problems out but kept some. I merely reordered some problems to create different forms of the same test for different periods. I would have to open all ten versions and spend two or three hours to sort out this mess.

It's much, much easy to keep a complete test bank in Scrivener and then to print only the questions you need for a given version. Option 1 is to make a Scrivener file for each test and to make each type of question a different document in the Binder. For example, you might have a document for true/false questions, short answer questions, essays, or document-based-questions. You can write the instructions for that type of question once. Then, each question would be a subdocument. You can reorder subdocuments by dragging-and-dropping to create different test forms. Unselect a question from "compile" and it will not print in this year's test but will remain safely stored in the bank. Best yet, you can use the inspector to give each question meta-tags: topic, difficulty, LAST YEAR USED, etc. Furthermore, you can keep all the support documents you want in the Research section of the Binder. For example, you can import PDFs and html files into Scrivener for your document-based-questions, printing them only when the questions you give your students necessitate it. (You can also import your old tests from Word files, so you don't need to retype everything.)

Option 2 is to aim even higher, making one Scrivener file your whole-year test and exam bank. You would need to make a different document for each unit, and then subdivide each of these documents into separate, appropriate subdocuments (such as question type). Since there's no limit on how many levels deep you can go, it seems like making a bank for the entire year should not be a problem. You can keep records of how students do from year to year in Scrivener in a subdocument you never print. You never need to misplace your data again!

For math teachers writing a problem set


Scrivener has a special advantage: MathType can easily be integrated. A further advantage: you can use the comments/meta-tags field to categorize problems by concept or difficulty, to write in the solutions to problems, and to make notes on what problems students find the most challenging. You can also insert tables and graphics for your problems and -- this is the best part -- you can import the source files, in whatever format you have, directly into Scrivener. Data is in Excel? Drag your Excel document in, and it'll live safely in your Scrivener file. GeoGebra file? Drag it in! The GeoGebra file will live safely in the Research section of the Binder.

When you click on an imported file, the appropriate program will start. For example, clicking on a Excel document in the Research section will cause Excel to start and open the file. However, these imported files are not "linked." You can't edit the Excel document and have those changes reflected in Scrivener. You have to copy the updated data from Excel back into your document. And you have to reimport the Excel file itself. But you can delete the original file, and the copy will live in Scrivener, and you never need to worry about losing it. Before I started using Scrivener, I would make my diagrams in GeoGebra, then throw the GeoGebra file away. It just got to be too much to manage all of them. But now with Scrivener, now I can keep all of my GeoGebra files neatly organized.

While I'm on it, GeoGebra is awesome. It's easy to use; it's pretty darn powerful; it integrates geometry and algebra and functions and even statistics; it's stable; and it's free! I make almost all my diagrams, from my Geometry to my B.C. Calculus class, in GeoGebra -- the only exception is 3D stuff. Download it for free here. It's available for PC, Mac, Linux, as an add-on for Chrome (it's a little slow but alright), as a iPad app, as an Android app, and as a Windows tablet app.

Anyway, back to Scrivener. The one downside is that you can't do fancy page layouts. You can resize an image, but you can't rotate it. You can make a table, but you can't combine cells or resize them beyond the default. You can't double-column or do tab stops. Or auto-numbering. Header and footer options are limited. However, you can export your document to a Word file where you could do those things. I think those deficiencies are far outweighed by the positives. Scrivener makes no claim to be a layout program; its goal is to help you write and edit your rough drafts, and mostly, I think its draft quality (as it were) is good enough for making a math problem set!

Debaters going paperless


The Binder can be a tub or accordion file, organized into different sections, e.g., neg > disads > econ da > links > ACA > specific cards or briefs. There are heading and subheading text styles, which could the tags, and the body text style could be the card text. The footnotes could be the citations, and comments the debater's annotations on each card or brief.



Here's the template I created.

Everything is at your fingertips, in one file (no need to open separate Word documents). Scrivener doesn't seem to have any size problems. Because it's designed to write books, it clearly can handle a lot (as opposed to opening up an equivalent number of files in Word, which does seem to tax my computer). The best part is that everything is searchable and search results can be filtered, e.g. a student can find a link card with a specific phrase even if he can't remember which disad it's a part of.

I imagine that a debater would use Scrivener to a) import his research files, in PDF, html, or Word documents; b) use Scrivener to write his briefs; and then c) use Scrivener in round as his document system. I imagine that he would work from a copy in each round, so he doesn't screw anything up permanently by accident. I would want my debaters to create a document for each speech at the top of the Binder, and then drag-and-drop the evidence they need to read into each speech document. The best part is that, when they need to share evidence with opponents, the "compile to PDF" option is awesome. The debaters can select only what they've read, and while the footnotes can be printed, the comments/annotation don't have to be printed.

The only difficulty is with sharing. It's easy to import files, but you can't use Scrivener to collaborate. I could imagine a team of two using one Scrivener file, synced through Google Drive or Dropbox, but I can't imagine a whole squad doing so.

Summary


Scrivener is a great program. It's super easy to import into; it's super easy to use; and it's super easy to export into different formats for final editing for page layout. Plus, they'll give you a free trial.

Thursday, April 18, 2013

Critique article -- response

I have ranted about critiques here; I complained about the current theoretical mish-mash that accompanies most critiques, and I endorsed straightforward critiques of language, thinking, and values, and rejected critiques that are disguised utopian counterplans or linear disadvantages. In my view, a critique is supposed to be an argument about the affirmative failing to prove its case prima facie. Too many critiques I have heard try to spin an implication about in-round discourse into non-sensical utopian impacts about my ballot changing the world, blah blah blah.

I thought Armand Revelin's article in the National Journal of Speech and Debate had the right spirit. He clearly knows his philosophy. Early on, he makes the important distinction that critiques do not function based on real-world impacts (causal chains) but on implications (which are really about logical sufficiency). I thought this summed up his point:

The implication of a kritik . . . may be that the affirmative incoherence amounts to a complete unintelligibility, and that the affirmative does not get access to its solvency and advantage claims because they are founded upon this unintelligibility. (p. 5)

This is exactly right. I would have thrown in the phrase "prima facie" because I am an old-school nerd, and a prima facie case is one where the evidence is logically sufficient to prove the claims and the claims made are sufficient to affirm the topic. The distinction he makes is a crucial one: critiques are different from disadvantages and counterplans because of a different voting issue -- logical lapses, not outcomes. I thought he should have pointed out that there is a third type of voting issue: the procedural voting issue, like topicality or arguments about fiat abuse. It is important to note that critiques are not procedural voting issues. When the judge votes on a critique, she is not voting against the affirmative because its plan is unfair.

While I agree with Mr. Revelin so far, it seems like his guiding idea is that the pressure for an alternative on a critique leads to bad debate, and on this, I disagree. In many examples, the alternative is so clearly implied as to be trivial. On his example of a language critique of the phrase "Native American," surely the alternative is to find a better phrase. Perhaps "First People"? Even when the alternative is not so clear, at root the team responding to the critique is asking about the fairness of the critique: "You say our case is based on sexist assumptions. What are some non-sexist assumptions we could have worked from? Is it even avoidable?" In this way, it is analogous to asking what cases meet a topicality violation. Just as some cases should meet a topicality violation if it is a reasonable violation, the alternative is the argument showing that the critique is reasonable -- that the affirmative case (or the topic or whatever claim is being critiqued) could have been put together in a better way. On the one hand, the alternative should not make or break the critique; we can reject bad logic, even if we cannot imagine better logic. A defense attorney does not have to produce the real criminal. On the other hand, it strikes me as a problem if the critiquing team cannot make some motions in the direction of the alternative, and moreover, it seems like there is usually an alternative implicit in the critique. Few critiques are simply nihilist, and there are good arguments that nihilist critiques make for bad debate.

I think the main reason critiques have become an awful theoretical mish-mash is not the pressure for an alternative, but instead because the community does not stand behind the proposition that dismantling a case is a sufficient reason to vote for a critique. Debaters make a muddle of critiques because they, perhaps rightly, believe that they need to add on impacts to win. This is why Mr. Revelin's assertion that a critique might have an "external impact" deeply troubles me. As far as I can tell, by this he means discursive impacts:

The affirmative can argue for the educational merits of discussing their original advantages within the remote discussion-space of debate. Although this is not the same as weighing the impacts of the advantages themselves, it is substantive and can be compared with the exclusion of such discussions that would follow if a negative kritik team is allowed to exclusively draw attention just to the non-remote discursive harms of their kritik. (p. 7)

Discursive impacts are the root of the problem. They are not impacts in the fiat-link-uniqueness-impact sense; they do not need to exist inside the "box" of hypothetically assuming plan happens. Are discursive impacts to be weighed against fiat-impacts? The debate community has answered yes, but I think the answer should be a firm no. The label "impact" is misleading; discursive impacts are much closer to procedural arguments. The voting issue of a discursive impact is, "Your advocacy is bad for the debate community; it makes us insensitive, or offensive, or unethical," which is quite similar to the voting issue of a procedural argument, "Your plan is unfair and ruins the spirit of this game." The fact that the latter is explicitly about the rules of the game, whereas the former often borders on rejecting the games-playing, does not alter the fact that both are arguments about debating. A rejection of narrow-minded rules is still about rules. Trying to spin a discursive impact into the mould of a fiat-impact leads to all sort of double-think double-ungood. Let us stop calling them discursive impacts, and let us structure them like the procedural arguments they are.

There were two other problems I had with Mr. Revelin's argument. One was about the assertion that the episteme critique always tied back to Heidegger. I would say that there are many critiques of the logical sufficiency of evidence that do not. One that comes to mind is Stephen Jay Gould's outstanding book, The Mismeasure of Man, in which he systemically goes through the history to intelligence testing to demonstrate the long racist history. A case (on some education topic) that used IQ-test-based evidence would be very vulnerable to this critique. But I fail to see how this is Heideggerish. I do think the type of critique Mr. Revelin discusses is a narrow kind of knowledge critique (and a particularly nihilist one, at that).

The other problem was his response to permutations. On counterplans, a permutation is simply the argument that the counterplan is not competitive, that there is no opportunity cost to the plan. On critiques, a permutation is simply the argument that the content of the Affirmative case (the plan, the advantages) can be upheld separately from its advocacy (language, evidence, logic, or values). In other words, the Affirmative argues that we should still like the case but under a different framework. There is an argument that this is severance, so perhaps a permutation is a bad argument.

I am glad others are thinking and pushing at the debate community to re-evaluate critique theory. Thank you, Mr. Revelin, for your thoughtful piece.

Monday, January 7, 2013

Impactranks

I recently heard of Josh Clark's new project impactranks.com. It is basically a coaches' poll. I respect the intention, but I strongly disagree with the method. Does anyone think the B.C.S. does a good job? It is largely based on a coaches' poll. The problem is that coaches do not see enough other teams play (or debate) and end up reflecting the perceived reputation of programs. There are sounder methods to rank teams that reflect actual results: wins, losses, and points. I would very much recommend reading, Who's #1? The Science of Rating and Ranking, by Langville and Meyer, two math professors, for many of the methods. My own method for debate is the weighted win method, explained here and here, in which whom a team beats matters even more than the raw win percentage.

One objection that people might point out is that debaters have so few opponents during the year that there is not enough data to think a mathematical method is any better than a poll. However, I would point out that the same problem exists in college football and professional football, yet mathematical methods based on opponent strength do work reasonably well.

Another objection is that debate has a lot of upsets, because of bad judges, off rounds, or unusual one-time tricks. True, but so does football, and the mathematical methods still work just fine. (In fact, I was even able to calculate an upset rate for college debate -- about 20%.) The other thought that I have had is to use the data on results to evaluate judges at the same time as debaters. Judges who return consistently unusual results (giving wins to worse teams with a lot of regularity) would lower their rating, so a loss from such a judge would not penalize a team by much.

The only difficulty to using these methods is that high school debate records are not stored in a clear format that shows: the two opponents, the judge(s), and the results. But the college records are. So here is my challenge to any reader. Before the N.D.T., I am going to rank all the teams, based solely on the results from the year, and publish that ranking. We will see how many results my rankings correctly predict (given the upset rate of 20%, anything in that neighborhood or better would be excellent). If any reader wants to come up with their own ranking, we will compare the results. You have until March 27th!