Showing posts with label cross-examination debate. Show all posts
Showing posts with label cross-examination debate. Show all posts

Friday, November 19, 2021

Judge intervention

"[W]hen people argue they move back and forth between strong claims that are weakly defended and weaker claims that are strongly defended." ~ Ross Douthat, https://www.vox.com/21760348/trump-2020-election-republican-party-ross-douthat


What is judge intervention?


Let's start out first by explaining what judge intervention is NOT:


- Judge intervention is not when a judge prefers your opponents' argument to yours. This is, ya know, "judging."

- Judge intervention is not when a judge ridicules your argument. This is rude, but that's about all it is.

- Judge intervention is not when a judge jokingly says, "Extra points to anyone who can avoid saying the President's name," or some such. Silly, and perhaps annoying, but that's about all it is.

- Judge intervention is not when a judge declines to discuss their decision after the debate. In some cases, talking after the debate is disallowed by a tournament; in other case, it may just be the judge's personal preference. So be it.

- Judge intervention is not when a judge grants your opponent favorable treatment (e.g., more prep time than is allowed by the tournament, advice, extra speech time or do-overs, etc.). This is unacceptable behavior and should be reported to an adult in charge of the tournament.

- Judge intervention is not when a judge ignores the debate, falls asleep, or in other ways is derelict in their duty to be a full participant and educator. This is unacceptable behavior and should be reported to an adult in charge of the tournament.

- Judge intervention is not when a judge demeans you personally. This is unacceptable behavior and needs to be brought to an adult's attention, preferably your coach and/or the tournament.

- Judge intervention is not when a judge votes based on the school, sex, gender, sexual orientation, race, ethnicity, or socio-economic class of the debaters, and not the arguments. Let's point out that it's beyond unacceptable and could be illegal.

- Judge intervention is not when a judge creates a hostile environment of harassment or fear or sexual joking. Or implies wins or speaker points can be traded for personal favors unrelated to debate. Beyond unacceptable and definitely illegal.


As you can see, judge intervention is not a catch-all term for rude or inappropriate judging behavior. Judge intervention falls somewhere in the middle of the spectrum: beyond merely rude, but well short of truly inappropriate behavior. Judge intervention is perhaps in some cases murkily unethical, although it would be very, very hard to prove any specific instance as unethical and will probably never result in an overturned ballot.


In many cases, judge intervention is the inevitable result of how YOU debated; it is, in at least some key ways, an inevitable part of judging even good debates.


So, what it judge intervention? Simply put, judge intervention is when a judge has to use their own prior knowledge to resolve an issue in the debate. Consider the following exchange:


Team A: "X is a fact" [no evidence provided]

Team B: "X is not a fact" [no evidence provided]

Judge: "… fudge you all"


In this instance, neither team provided any evidence to back up this point. If X is a crucial fact for the debate, then the judge will have to insert their prior knowledge to resolve it. (If X isn't important, then most judges will happily ignore it as irrelevant so as to avoid intervening.) Of course, this example is not the only one where a judge might be forced to intervene:


Team A: "The sky is green" [no evidence provided]

Team B: "…"

Judge: "Come on, team B!"


While the judge should be open minded, this doesn't extend to being completely gullible. We expect everyone to bring at least some prior knowledge into the debate with them. To assume the judge knew nothing would actually be an exhausting way to debate; in every round, you would spend the bulk of your time explaining how the government worked, that people were bipedal and how they used language to communicate, and so on. To paraphrase the T.V. show Archer, "Why is there conflict in the Middle East? Well, a long time ago there were dinosaurs, then they died, and their bones are what your car uses for food." Giving every bit of backstory would slow debates to a crawl. The judge should be allowed to have some degree of common sense without being accused of wrongdoing, even though it's technically intervening to ignore Team A's insane green-sky claim.


There are in fact many instances in which we should actually expect the judge to intervene and reject your argument, even if the opposing team said nothing against it. Looking at Toulmin's model will help us delineate all the ways. Toulmin's model of an argument contains the following parts:


Claim: your argument

Evidence: your supporting facts or expert opinion

Warrant: the logical connection between evidence and claim

Relevance: the meaning of your claim to the debate


A judge might reject your claim because it's muddled and unintelligible; they might reject it because it contradicts claims you make in other parts of the flow. A judge might reject your claim because it's of the wrong type. There are four types of claims: (1) factual, "X is true"; (2) causal, "If X happens, then Y will result"; (3) value, "X is right and good" or "X is more important than Y"; and (4) definitional, "The word X means…" These types of claims are not interchangeable, so don’t go answering a moral question with a factual claim. Definitions don't really address cause and effect questions…


A judge might reject your argument because you provide no or weak evidence. What if you argue the economy is doing well and your evidence is several years old? What if your source is, unbeknownst to you, a widely discredited conspiracy theorist? There are some real gray areas here. In some instances, we might say, "The other team didn't point that out, so the judge merely needs to accept any evidence I provided at face value," but in other cases, that feels absurdly deferential to a team reading, say, neo-Nazi propaganda or other gibberish. The judge should accept most reasonable seeming sources at face value, even if they perhaps know of a specific flaw with the evidence or source, but what is a reasonable source? Judges may have to walk a very fine line here because, though they are no experts on the topic, they may know more than you do. The line I often draw is,"What would be reasonable for a student at your age and experience level to know?" Do I expect you to know an insane conspiracy theory website when you see one? Yes. Do I expect you to know how respected an academic at a good university is in their field? Not really. So I would throw out the conspiracy website, even if it's unchallenged, while allowing the academic—even if I know something specific about their work—if your opponent says nothing. But if your opponent does say something, I will definitely let my skepticism fly free. My initial, naive presumption you shouldn't know something at your age goes right out the window when your opponent knows it.


A judge might reject your argument for lacking a warrant connecting the evidence back to the claim; the evidence might appear senseless and irrelevant. But the judge might also reject some kinds of warrant that you do provide. We don't expect judges to be subject matter experts in every single field a debate might touch on: they are not scientific experts, legal scholars, linguists, logicians, political theorists and strategists, pollsters and statisticians, and foreign policy gurus and intelligence analysts, all rolled into one. It's quite possible our explanations are too "inside baseball" for a lay judge to follow. Debating is a communicative act, so we need to do a bit of audience analysis. Our explanations, our reasoning, and our warrants need to assume that we are persuading an intelligent but lay audience; after all, our judge is not a subject matter expert but a smart generalist. In other words, yes, it's reasonable that a judge could reject a warrant you've provided that's perfectly intelligible and correct simply because it's over the judge's head—and it's reasonable to expect you to know it would have been.


A judge might believe everything about your argument but simply not get why you believe it's relevant to the debate.


In short, there are many, many reasons why a judge may intervene in a debate. Rather than the problem being that judges intervene, the problem is that you think it's an aberration. Judge intervention isn't an aberration. It happens all the time. It's the norm. It's a normal part of human decision-making to use one's prior knowledge and common sense to evaluate arguments in holistic ways. It's called critical thinking, yo. Judges don't accept everything you spoon-feed to them, and if they did, they wouldn't be able to reach decisions. A truly blank slate judge would be like a computer program that fritzes out before returning an answer; a spinning rainbow; a blue screen of death. Judges are thinking human beings with their own life experiences and assumptions who are trying to be—but not always succeeding at being—open minded to your arguments.


The problem with judge intervention is not that it happens in the first place but that you don't know how to use that fact to your advantage. A great debater knows how any judge, any human being, likes to reason; knows when, why, and how a judge is likely to intervene; and uses that knowledge to win. To be clear, I'm not talking about judge adaptation (changing your speed, style, or argument selection in order to appeal to the judge's preference). I'm talking about using the human psychology of decision making so as to find ways to "stitch" the inevitable gaps in arguments so that the judge won't notice them as flaws.


A winning debater is ultimately a good storyteller, and those stories operate on the micro level (warrants and reasoning) and on the macro level (relevance and voting issues). Those stories are crucial to winning over a judge. Judges need cohesive stories to justify their ballot, and if you don't supply it, the judge will. While you might not do a perfect job on every issue and position in the debate, telling a good macro story enables you to persuade the judge that you are, overall, winning. Likewise, not every claim or piece of evidence is perfect, so a good micro level story enhances your credibility; you paper over any holes in logic by demonstrating an overall cohesive understanding of the disadvantage or the case.


Lacuna in your arguments and direct contradiction by opponents are just inevitable in close debates. The latter, because your opponents are good; the former, because under the pressure of a close debate, you won't get to everything. There will be holes; the judge has to intervene to pull together a cohesive picture.


Consider a close football game, where the final score difference is less than 7 points between the teams. Either team could have won. A bobbled pass that was caught, a different penalty call, a timeout left—any of these could have made the difference in the game. Both teams had a near 50% chance of winning.


So too in debate. In a blowout debate, you are able to put together full, complete arguments and full, complete strategies. Your opponent isn't able to contradict you and isn't prepared to put together a competing narrative. You perhaps have a 90% or 95% chance of winning; there is but a slim chance that your opponent successfully throws a Hail Mary and makes a lucky good argument or that your judge happens to find one of your arguments completely unbelievable (i.e., you've unknowingly used a conspiracy website, so far as the judge is concerned). Pure judge intervention like the latter is rare, but it's just a part of human reasoning. Take the 90% or 95% chance of winning and be glad. There's no way to be 100% airtight in front of a thinking human being with experiences different than your own. They will have different assumptions and thought processes than you do. Mostly airtight is good enough.


A close debate, however, is like the close football game. Any play or argument could make the difference. Your efforts to complete every argument, although valiant, are impossible under the time pressure and constant barrage from your opponents. There will be holes. Therefore, the most important thing you can do is to tell good stories to paper over any holes. You want the inevitable judge intervention to slide your way. You won’t be able to increase your odds to 80% or 90%, but you can perhaps up them to 60% from 50%. In this instance, judge intervention isn't a bad thing, either; it's just part of the round being so close, so use it to your advantage.


When the round isn't close, out-of-clear-blue-sky judge intervention should be rare. It stinks, but all you can do is build your arguments as carefully and as air-tightly as you can—and learn from judges' rare interventions how to improve upon your techniques.


When the round is close, and judge intervention become inevitable, do your best to push the judge to intervene for you. Do this by using solid evidence, being clear and consistent in your positions, and most of all, telling compelling logical stories and impact analyses. Credibility and persuasion matter.

Saturday, May 16, 2015

The "How to" of Debate

I decided to make the e-version of my introductory debate textbook free. It's an appropriate guide for middle schoolers or high schoolers.

To download The "How to" of Debate in kindle format, for free, click here. (By the way, the easiest way I know to get an ebook onto any kindle device is to download it to your computer, then email it to your kindle's special email address. Look on Amazon, in your account info, for "manage content and devices." Then go to Settings, and the email address for your kindle device should be under "Personal Document Settings." Make sure your kindle device is enabled to receive emails from your personal email address, then you are ready to go! Just email it to yourself, and a few minutes later, it will sync onto your kindle device. This is much easier than trying to figure out where your device stores kindle files.)

For the paperback version of The "How to" of Debate, click here (still $15.95).

Friday, March 7, 2014

Debating Policies

It also had not occurred to me, until today, to post a copy of my policy debate textbooks. The new, improved version, The "How to" of Debate, is available on lulu. The old version, Debating Policies, is below for free.

Thursday, April 18, 2013

Critique article -- response

I have ranted about critiques here; I complained about the current theoretical mish-mash that accompanies most critiques, and I endorsed straightforward critiques of language, thinking, and values, and rejected critiques that are disguised utopian counterplans or linear disadvantages. In my view, a critique is supposed to be an argument about the affirmative failing to prove its case prima facie. Too many critiques I have heard try to spin an implication about in-round discourse into non-sensical utopian impacts about my ballot changing the world, blah blah blah.

I thought Armand Revelin's article in the National Journal of Speech and Debate had the right spirit. He clearly knows his philosophy. Early on, he makes the important distinction that critiques do not function based on real-world impacts (causal chains) but on implications (which are really about logical sufficiency). I thought this summed up his point:

The implication of a kritik . . . may be that the affirmative incoherence amounts to a complete unintelligibility, and that the affirmative does not get access to its solvency and advantage claims because they are founded upon this unintelligibility. (p. 5)

This is exactly right. I would have thrown in the phrase "prima facie" because I am an old-school nerd, and a prima facie case is one where the evidence is logically sufficient to prove the claims and the claims made are sufficient to affirm the topic. The distinction he makes is a crucial one: critiques are different from disadvantages and counterplans because of a different voting issue -- logical lapses, not outcomes. I thought he should have pointed out that there is a third type of voting issue: the procedural voting issue, like topicality or arguments about fiat abuse. It is important to note that critiques are not procedural voting issues. When the judge votes on a critique, she is not voting against the affirmative because its plan is unfair.

While I agree with Mr. Revelin so far, it seems like his guiding idea is that the pressure for an alternative on a critique leads to bad debate, and on this, I disagree. In many examples, the alternative is so clearly implied as to be trivial. On his example of a language critique of the phrase "Native American," surely the alternative is to find a better phrase. Perhaps "First People"? Even when the alternative is not so clear, at root the team responding to the critique is asking about the fairness of the critique: "You say our case is based on sexist assumptions. What are some non-sexist assumptions we could have worked from? Is it even avoidable?" In this way, it is analogous to asking what cases meet a topicality violation. Just as some cases should meet a topicality violation if it is a reasonable violation, the alternative is the argument showing that the critique is reasonable -- that the affirmative case (or the topic or whatever claim is being critiqued) could have been put together in a better way. On the one hand, the alternative should not make or break the critique; we can reject bad logic, even if we cannot imagine better logic. A defense attorney does not have to produce the real criminal. On the other hand, it strikes me as a problem if the critiquing team cannot make some motions in the direction of the alternative, and moreover, it seems like there is usually an alternative implicit in the critique. Few critiques are simply nihilist, and there are good arguments that nihilist critiques make for bad debate.

I think the main reason critiques have become an awful theoretical mish-mash is not the pressure for an alternative, but instead because the community does not stand behind the proposition that dismantling a case is a sufficient reason to vote for a critique. Debaters make a muddle of critiques because they, perhaps rightly, believe that they need to add on impacts to win. This is why Mr. Revelin's assertion that a critique might have an "external impact" deeply troubles me. As far as I can tell, by this he means discursive impacts:

The affirmative can argue for the educational merits of discussing their original advantages within the remote discussion-space of debate. Although this is not the same as weighing the impacts of the advantages themselves, it is substantive and can be compared with the exclusion of such discussions that would follow if a negative kritik team is allowed to exclusively draw attention just to the non-remote discursive harms of their kritik. (p. 7)

Discursive impacts are the root of the problem. They are not impacts in the fiat-link-uniqueness-impact sense; they do not need to exist inside the "box" of hypothetically assuming plan happens. Are discursive impacts to be weighed against fiat-impacts? The debate community has answered yes, but I think the answer should be a firm no. The label "impact" is misleading; discursive impacts are much closer to procedural arguments. The voting issue of a discursive impact is, "Your advocacy is bad for the debate community; it makes us insensitive, or offensive, or unethical," which is quite similar to the voting issue of a procedural argument, "Your plan is unfair and ruins the spirit of this game." The fact that the latter is explicitly about the rules of the game, whereas the former often borders on rejecting the games-playing, does not alter the fact that both are arguments about debating. A rejection of narrow-minded rules is still about rules. Trying to spin a discursive impact into the mould of a fiat-impact leads to all sort of double-think double-ungood. Let us stop calling them discursive impacts, and let us structure them like the procedural arguments they are.

There were two other problems I had with Mr. Revelin's argument. One was about the assertion that the episteme critique always tied back to Heidegger. I would say that there are many critiques of the logical sufficiency of evidence that do not. One that comes to mind is Stephen Jay Gould's outstanding book, The Mismeasure of Man, in which he systemically goes through the history to intelligence testing to demonstrate the long racist history. A case (on some education topic) that used IQ-test-based evidence would be very vulnerable to this critique. But I fail to see how this is Heideggerish. I do think the type of critique Mr. Revelin discusses is a narrow kind of knowledge critique (and a particularly nihilist one, at that).

The other problem was his response to permutations. On counterplans, a permutation is simply the argument that the counterplan is not competitive, that there is no opportunity cost to the plan. On critiques, a permutation is simply the argument that the content of the Affirmative case (the plan, the advantages) can be upheld separately from its advocacy (language, evidence, logic, or values). In other words, the Affirmative argues that we should still like the case but under a different framework. There is an argument that this is severance, so perhaps a permutation is a bad argument.

I am glad others are thinking and pushing at the debate community to re-evaluate critique theory. Thank you, Mr. Revelin, for your thoughtful piece.

Friday, February 15, 2013

Making debaters focus on arguments, not cards

Every time my debaters say, "I found a good card," I cringe. Is there a good argument in the card, perhaps?

I have found that my debaters, especially my beginning debaters, get overwhelmed when trying to do research to put an argument or case together. One part of the difficulty is that there are so much new information coming at them (they are learning the content about their topic); another part of the difficulty is that the structure of an argument is new to them, and so they have trouble (a) interpreting the structure of the arguments they are reading (and trouble evaluating these arguments) and (b) imagining how to use these arguments for their own case. The result is that beginning debaters tend to focus on the conclusions in a quotation, missing the arguments that get the author there, and thus essentially are guilty of an appeal to authority.

But the same thing happens, in a different way, when an advanced debater turns in a file with fifty uniqueness cards that all repeat one basic fact. I have a hard time persuading advanced debaters that a ten-card file might be the most excellent disadvantage they have seen all year. I have tried different techniques to get my point across, most notably handing a debater back a card he says is good with the instruction to highlight every argument and fact but not any claim. This can help the advanced debater who reads without reading, but it does not help the beginning debaters in the moment of research.

Rather than ask themselves, "Does this argument or fact help me?", beginning debaters -- maybe even many advanced debaters, too -- focus on whether something is "card"-able. The beginning debaters spin themselves into fits worrying about the threshold for cutting something, mostly ignoring the content. (And my injunction that, if the fact or argument is good but the author is lousy and unclear, then they should look for and will find the same fact by a better writer is completely ignored.) Direct quotation is a labor-saving device: it is easier to copy and paste than paraphrase, and it is easier to verify the claim (tag) is a correct interpretation of the author's work. But it can short-circuit the thinking process. I have been sorely tempted to prohibit my debaters from direct quotation in debating, forcing them to paraphrase, to get them to focus on the actual, essential arguments and facts they are gleaning from their research. (It is certainly a great technique for practice rounds.)

Instead, I have come up with argument sheets. I tell them that the first stage of research is merely about mapping out the terrain; they can fill in the sheet and save links or PDFs but not cut cards. Of course, I created different sheets for the three types of claims: factual (uniqueness or harms), causal (links or solvency), and value and philosophical claims. If a student is researching uniqueness for a disadvantage, I give her the first kind of sheet, which directs her to pay attention to what facts she finds, how those facts were collected by the original researcher (e.g., statistics do not fall from the sky), and what the facts mean for her claim. For example, perhaps she is making a claim about the effect of a law for her inherency. I want her to point to specific text in the law or to point to Congressmen's statements, and to recognize that each method has limits. With causal arguments, I want students to pay attention to the complexity of causation. I want them to move away from the simple, pat story or scenario; I want them to think about effects being overdetermined. The hardest arguments for most debaters to analyze are value arguments. I want my debaters to pay special attention to the broader philosophical ideas that are tied into, as well as to clearly think about what is included and what is excluded in a philosophical judgment. The idea I am trying to hammer home is that philosophical concepts are all about distinctions. (I am reminded of the old joke about a philosopher, asked how he liked his job: "I make some distinctions. It's a living.")

Below is an example of what filled-in sheets might look like.

Argument sheet (example)



After they have filled in a sheet, then we will discuss the case and organize the key ideas into an idea map or web. (I have found that they can not jump straight from reading, without note-taking, to idea-mapping.)

Finally, they go back to start to cut cards. I have found that debaters now tend to cut cards on factual claims that are shorter than before, but that are more likely to have some context for the collection or meaning of the data presented. I have found that debaters now tend to cut cards with causal claims that are longer than before -- they are trying to include more complex analysis. And I have found that debaters are now less likely to cut cards making value claims that are just conclusions. The students look a little more carefully for cards that link an issue to a big philosophical concept. Which is to say, the debaters are more thoughtful and discriminating.

Thursday, January 3, 2013

Literature base

I'm going to enter into cantankerous old man territory by making this post: the National Forensic League is not doing a great job with Lincoln-Douglas topics, but it is doing a decent job with Public Forum topics. My problem is the scope of the topic literature bases. It is unreasonable (and counter-productive to education) to give students a big topic to research and little time to do it. Students learn best when they can thoroughly explore a topic and are able to prepare on most of the key arguments. Of course, because debate is a competitive activity, some opponents will always search out unusual, squirrelly arguments, but my point is that being caught off-guard should be a rare experience for debaters. If the topic is too broad to prepare and debaters are regularly caught off-guard, then the value of research and preparation is undermined.

Given that topics are used for one month in PF, two months in LD, and ten months in Policy debate, I think the literature bases should follow a 1:2:10 proportion for these various formats. In other words, the length of time a topic is used should be roughly commensurate with how big the literature base is. This only makes sense: LD debaters might prepare for four to six topics during the year, so it only makes sense that they would invest about one-fifth the research and preparation time into each one as policy debater do; PF debaters might prepare for nine to eleven topics, so one-tenth the work seems fair. How would the N.F.L. do against this benchmark?

I started with the last five years' policy topics. The N.F.L. does NOT write the policy debate topics; those are written by the National Federation of High School Associations. Generally, the policy topics are just about the right breadth to spend a whole year researching. This result is probably no accident, because their process requires the topic framers to consider the literature base qualitatively, and to a lesser extent, quantitatively. To make a very rough gauge of the size of the literature bases, I typed in key terms and terms of art into ProQuest. ProQuest is a general use research database, containing both news articles and some academic journal articles, which is commonly available to high school students. I chose search terms that seemed appropriate to find the core articles of each topic. I make no apologies for the very provisional method; don't infer too much from this. I think this gives a sense of scale for comparison, but not much else.

Policy topics:

2012-2013 topic: "The United States federal government should substantially increase its transportation infrastructure investment in the United States."
Search terms: "transportation infrastructure" and ("United States" or "U.S.")
Results: 16,069

2011-2012 topic: "The United States federal government should substantially increase its exploration and/or development of space beyond the Earth’s mesosphere."
Search terms: ("space exploration" or "development of space") and ("United States" or "U.S.")
Results: 25,657

2010-2011 topic: "The United States federal government should substantially reduce its military and/or police presence in one or more of the following: South Korea, Japan, Afghanistan, Kuwait, Iraq, Turkey."
Search terms: ("military presence" or "police presence") and ("United States" or "U.S.") and ("South Korea" or Japan or Afghanistan or Kuwait or Iraq or Turkey)
Results: 17,954

2009-2010 topic: "The United States federal government should substantially increase social services for persons living in poverty in the United States."
Search terms: "social services" and poverty and ("United States" or "U.S.")
Results: 22,133

2008-2009 topic: "The United States federal government should substantially increase alternative energy incentives in the United States."
Search terms: "alternative energy" and ("United States" or "U.S.")
Results: 36,728

There is a wide variation from topic to topic. But the mean of 24,000 articles seems about right and a reasonable place to start from.

How do PF topics compare? They are in the right ballpark.

2012-2013 PF topics:

Sept. topic: "Congress should renew the Federal Assault Weapons Ban."
Search terms: "assault weapons ban"
Results: 3,367

Oct. topic: "Developed countries have a moral obligation to mitigate the effects of climate change."
Search terms: "developed countries" and ("mitigate" or "mitigation") and "climate change"
Results: 2,546

Nov. topic: "Current U.S. foreign policy in the Middle East undermines our national security."
Search terms: ("U.S. foreign policy" or "United States foreign policy") and "Middle East" and "national security"
Results: 3,105

Dec. topic: "The United States should prioritize tax increases over spending cuts."
Search terms: ("tax increase" or "spending cut") and "fiscal cliff" and ("United States" or "U.S.")
Results: 1,200 (66,279 without "fiscal cliff" term)

Jan. topic: "On balance, the Supreme Court decision in Citizens United v. Federal Election Commission harms the election process."
Search terms: "Citizens United" and election
Results: 3,885

Feb. topic: "On balance, the rise of China is beneficial to the interests of the United States."
Search terms: China and ("United States interests" or "U.S. interests")
Results: 5,756
Search terms: "rise of China" and ("United States" or "U.S.")
Results: 3,350


The China topic seems to be dangerously large, but the remaining seven topics have an average of 2,800, just about one-tenth of a policy topic. How do LD topics compare?

2012-2013 LD topics:

Sept./Oct. topic: "The United States ought to extend to non-citizens accused of terrorism the same constitutional due process protections it grants to citizens."
Search terms: terrorist and "due process" and ("United States" or "U.S.")
Results: 5,505

Nov./Dec. topic: "United States ought to guarantee universal health care for its citizens."
Search terms: "universal health care" and ("United States" or "U.S.")
Results: 10,696

Jan./Feb. topic: "Rehabilitation ought to be valued above retribution in the United States criminal justice system."
Search terms: (rehabilitation or retribution) and "criminal justice" and ("United States" or "U.S.")
Results: 10,546

With the exception of the first topic, they are all quite large topics. The health care topic is essentially the 1993-1994 policy debate topic! The criminal justice topic is similar to the Jan./Feb. 2011 LD topic, except that one was limited by its focus on juveniles.

What specific recommendations would I make to the N.F.L.?

First, when writing topics, please consider the size and quality of the literature base. Perhaps develop some standard statistics to gauge the size of the literature base. If the potential topic generates too many hits, narrow the topic in some way; if the potential topic generates too few hits, broaden it.

Second, please report out the statistics to us members whenever we vote between different wordings of the same topic. It would be helpful to know which version is the more narrowly worded.

Friday, July 13, 2012

Meta-debate tournament tabulation post

I have been at this blog for three and a half years. I have discussed neat calculus, geometry, and statistics problems; I have laid out some discussions for a critical thinking course; but in about half the posts, I have discussed tournament tabulation procedures. I have been working through a lot of ideas as I try to make a case for building a new generation of programs.

I just received a new book today, Who's #1? The Science of Rating and Ranking, by Langville and Meyer, two math professors. They state that they were frustrated that the methods they discuss in their book were not collected in any one other single book. I can say I share that frustration. At an initial look through it, the book is amazing. A lot of the methods they discuss I have looked at in one form or another in considering a tab program; many I have not. It is both validating and humbling. I have a lot of work to do! So, for quite a while, I am going silent on tab programs while I read this book. I will still post on neat mathematics problems and critical thinking course discussion ideas.

Before I go temporarily silent on the tab program topic, though, I thought it would be good to summarize what I have written so far. My positions have evolved in three and a half years considerably.

I started out looking at whether high-low and high-high pairings created fair schedules for teams, specifically looking at opponent wins. I looked at the Harvard tournament, a set of big national tournaments, and what was possible in theory. In general, I think the case is pretty clear that even at big tournaments, where constraints should not be an issue, the traditional methods fail to deliver fair schedules -- even within a bracket. Teams break with easy schedules; teams don't break with hard schedules.

I looked for a way to pair within brackets that I felt was more fair. The idea I hit on was strength-of-schedule pairings: a team would get an opponent that would balance out its schedule difficulty compared to other teams in the same bracket. This is a method only a computer can do, since it requires simultaneously evaluating whether team A is a good opponent for team B (as defined above) AND whether team B is a good opponent for team A. Keeping track of both ratchets up the complexity beyond a human's hands. It is still an understandable method, just too many calculations for a person to do. I tried it on a small tournament and a large tournament and found that it does indeed work to even out the schedules.

That problem "solved," I started thinking about evaluating a team's strength (and therefore also a team's schedule strength) in a more sophisticated ways than just wins and losses. I looked at graph theory, but for all its promise for some tiebreakers in round robins, it is not a good method for regular tournaments. I looked at one weighted wins scheme, where the points per win decline for each subsequent round, but this system is not good. I next looked at a weighted wins scheme where a team receives extra "points" for its defeated opponents' wins and loses points for its defeating opponents' losses. This does seem to work well for a simple method, and it lead me to think about more complex ways of getting, from the data, a team's strength on the affirmative and strength on the negative. I have also been thinking about how reliable even the best methods are. How many "upsets" are there in debate?

Along the way, I have also thought about side assignment here. I have written about how many teams will break here, here, and here. And I have sparred with A Numbers Game on whether there is topic side bias (sorry, it looks to me like novices muck up the average; varsity debaters get closer to parity) or judge side bias (again, sorry, it looks alright to me).

So where have I landed? I started out thinking all I wanted to do was propose a different within brackets pairing algorithm, which then expanded to thinking about measuring a team's strength and opponent strength. But a very early post on round robins planted the seed: preliminary rounds don't have to be elim rounds. They don't need winners and losers brackets. Teams need to meet all their opponents, or failing that, a good cross-section. So the fairest statement of what I believe now is that tournaments should ditch the brackets and start making sure that teams get mixed up by skill level and geography in prelims; let a sophisticated algorithm do the mixing, not random chance; let the best teams break, and use elims like they always have been: to pick the champion. And I think N.F.L. Nationals ought to start.

Sunday, June 24, 2012

Debate shouldn't be sophistry

In high school humanities, that is, English, history, and perhaps a semester of civics, ethics, or politics, a student is supposed to pick up five skills:
  1. empathy for others, especially those who are different; 
  2. ethical reflection; 
  3. a understanding of social and/or government functioning; 
  4. expository writing skills: putting together a coherent argument with evidence; and 
  5. critical reading skills: analyzing texts carefully. 
It has always seemed to me that doing competitive debate does a really good job teaching a student all but the first. However, I have grown more concerned as I age that what debate teaches is very deep but misses some topics. I haven't become a gray-haired proponent of "cultural literacy" in my dottage: I still think that because debate topics are current, they are inherently more interesting and relevant to students than fusty old treatises. If you think high school students ought to read Aristotle and Plato fully, then I'm not your man. I believe they're far more likely to enjoy learning a little philosophy in order to apply it to current controversies. No, my concern is that, while it is good that debaters learn a lot about their topics, they are missing a few key ideas they ought to be taught specifically. My key recommendation is that second-year debaters take some kind of general critical thinking course that fills in these gaps.

One area I've written about before is logic. It is shocking to me that debaters would not know Toulmin's model, how to model definitions with Venn diagrams, or how to model causal arguments with Ishikawa diagrams. While it's not necessary to get into the complexities of propositional logic and truth tables, the basics ought to be covered.

Another area I've mentioned before too is statistics. Debaters are woefully uninformed about how statistics are collected and interpreted. To be blunt, they'll quote just about anyone, whether or not the method makes a lick of sense. Nothing too much is required: just a little knowledge about sampling methods, sources of sample bias, how to understand different correlational studies (like regressions), and how to understand experiments.

But where I think debaters come closest to outright sophistry is on critical arguments. Critical authors can shed light on complex issues. At their best, debaters quote these authors' anthropological, cultural, economic, and psychological investigations to unpack the standard assumptions most public policy authors approach a topic with. The critical authors can help all of us get behind the positions to see how our assumptions and values shape what we believe and advocate for. At their worst, debaters quote these authors to confuse and bury their opponents under a flurry of ten-cent words.

Why is it that outlandish impacts rule the day in policy debate? This is a bit of an overstatement. But I might argue that the development of critiques was probably driven by outlandish brink-based impacts. Many critiques re-introduce linear impacts into the debate round, i.e., endemic problems that are worsened by the plan. For example, racism exists before the plan, but plan makes it worse. Yet no judge would vote for this argument stated as a simple linear disadvantage. Add on the ten-cent words about critical legal studies and it's a go. The problem for debate theory is that the voting issue on a critique brings in a lot of philosophical and debate theoretical baggage: Is the judge voting against the Affirmative's discourse? Or is he voting against the Affirmative because the false assumptions identified by the critique show the case's logic is suspect (e.g., racist arguments undermine a civil rights plan)? Or is the judge simply voting that the critique's "impact" is bad (e.g., racism bad), like a linear disadvantage?

I believe many critiques only work by stringing together a theoretically inconsistent position: the link is about the topic itself, almost a counterwarrant (a guaranteed link, no matter what the plan actually does, e.g., attempting to improve women's literacy in Africa is based on sexist assumptions); the answer to comparative arguments is about pre-fiat, in-round discourse; and the implication is about real-world, post-fiat, non-comparative assessments of the critiqued impact. I can see some truth in the criticism that critiques are utopian counterplans: even thinking about the topic leads to racism -- just look at how the Affirmative speaks -- but by rejecting this, we can magically live in a world where racism is alleviated. Anyway, back to the main point:

If critiques are just bringing in linear impacts, why not just cut straight to running a linear disadvantage? The answer is that deprived of all the hoopla and sleight of hand (of critique theory and the critical author's fancy words), I believe a linear impact would never win a round.

Part of my solution is to teach every debater in a critical thinking course and to grind to a halt the competitive advantage of poorly-argued critiques. If a debater knows how to construct an argument, she can hammer away at the gaps in a critique. If a debater has a good sense of the basic flavors of philosophy and recognizes a nihilist argument when he hears one, great. Furthermore, many debaters go on to be judges, so getting this basic knowledge in them early is important.

But the second part of my solution is to ask judges to (a) stop rewarding sophistry and (b) recognize that linear disadvantages ought to win sometimes. This is not a theory problem or a rules problem. There is nothing that says linear disadvantages are weak. It is not a problem in the topic literature -- debaters are usually stretching the literature as it is. I am skeptical of the "literature checks abuse" argument. Most debaters' positions are undercut by their own authors; no one thinks the horrible impact scenarios are very likely. We the debate community need to recognize that:

     impact = mag x (duration) x prob / timeframe.

Humans are bad at estimating probability. But debaters and judges ought to make an effort to do better. Unfortunately, as it is, we've decided between offensive and defensive arguments that defense never wins. Never.

If debate could keep critiques straightforward and bring back good old-fashioned linear disadvantages, I would have a lot more confidence saying to other humanities teachers that competitive debating does indeed teach social/governmental knowledge, expository writing skills, and critical reading skills. As it is, I worry that we're doing our own thing that's becoming too divorced from reality.

Friday, April 13, 2012

Weighted wins 2

I've been interested in using weighted wins as a statistic for a while. The idea is that by considering the strength of an opponent, an algorithm could handicap the results: a win against a good opponent counts for more than a win against a middling opponent. Of course, the algorithm has to base the calculation of how good an opponent is on how well it did against its opponents, so the algorithm must be based on the whole set of results. Therefore, I looked at a debate tournament and an N.F.L. season. The process can continue infinitely, although usually, after a couple times, things settle down and the handicapping doesn't change much upon further iterations. This is a kind of Markov chain technique.

One assumption that this depends on is that a team has an invariant strength. Clearly, this is suspect for debate (differing strengths on the affirmative and negative sides) and the other N.F.L. (the defense and offense are, literally, two different teams). Is it possible to use the same idea but adapt it to recognize the split-strengths?

I input total offensive yards from each 2011 N.F.L. game. A lot of yards against a weak defense is good; a lot of yards against a strong offense is better. By using only this information -- offensive yards for each team vs. each opponent -- a few matrix operations yielded a weighted offensive yards statistic and a defensive strength score for each team's offense and defense, respectively:


For example, against the "average" defense, the Saints would have earned 460 yards. The Steelers defense would have cut that to 82%, to 377 yards. Or so say these statistics. They never actually played.

How well did these statistics work? Well, as predictions, terribly. But, compared to reputable sports statisticians, fairly well. This is a comparison of my rank versus football outsiders rank (before the playoffs began) [offense on left, defense on right]:


On offensive, the difference was on average 3 ranks. On defense, the difference was on average 5.3 ranks. So, the rankings I generated from nothing other than actual yards in regular season games compared moderately closely to the ranks based on a complex calculation they call the DVOA (Defense-adjusted Value Over Average).

The same idea would work for debate tournaments: affirmative strength and negative strength could be treated as two separate variables for each team.

Sunday, September 13, 2009

Topic side bias

A Numbers Game did some interesting work on the side bias of various college topics, here, specifically, controlling for team strength. In that spirit, I decided to use a different method and see how the results compared.

I used a matched pairs method: for each team, there is a matched pair of results: that team's win percentage on the affirmative, and that team's win percentage on the negative. If the two results show no difference, then the team did equally well (or equally poorly) on both sides of the topic. If the two results do show a difference, there are three possible explanations: (1) the team isn't equally strong on both sides of the topic, e.g., the 2A isn't as good as the 2N; (2) the team hit an unequal set of opponents on the two sides; or (3) there is side bias on the topic. When one looks at all the teams on a topic, (1) is unlikely because the whole point is to control for team strength by assuming that the average team is equally strong on either side, (2) cancels out when one looks at the entire pool, and (3) is left as the most plausible outcome. Although this method still relies on the assumption of invariant strength (that a team has a fixed strength, the same on both sides of the topic, unchanging throughout the year), so does any other method that attempts to control for team strength. With those disclaimers, here are the results:


The third column shows the mean of the matched pairs computation: for each team, I subtracted its negative win percentage from its affirmative win percentage, and I averaged this score over all the teams that year. The fourth column shows a calculated (not the actual) affirmative win rate. They compare closely to the actual rates A Numbers Game already found. The two that are highlighted differ slightly. The affirmative win rate for the China topic my analysis suggests is slightly lower than the actual rate. The affirmative win rate for the courts topic my analysis suggests that the negative had an advantage, while the actual rate showed an affirmative advantage.


Category 1 is roughly the 0-50th percentile (in terms of rounds of competition); category 2 is 50-75th; category 3 is 75-87th; category 4 is 87-94th; and category 5 is 94-100th. You can see the results clearly in both: the less experienced teams had greater success on the negative; the more experienced teams did better (relatively or absolutely) on the affirmative. The reason why the originally calculated affirmative win rate was too low was because there are so many more less experienced teams that bring down the average matched comparison -- but they debate few rounds, so they do not have a big effect on the total ballot count. (I looked at the same tables for other years, but there were no patterns as clear as '05-06 and '06-07.)

One final note: All of the whole years' analyses are significant to at least the 95% level except for '06-07, which is only significant at about the 85% level. I used a 1-sample t-test, since the distributions are more or less normal:



(This is '08-09, but all the distributions look like this.) In fact, the distribution looks a little tighter than the normal curve (given the population's mean and standard deviation), since so many teams have 0% aff-neg win spread. The cumulative frequency graph is even more persuasive:


The blue line represents a normal CDF, given the population's ('08-09) mean and standard deviation. Anything below the line on the left or above the line on the right is tighter than normal. The reason is that the standard deviation is pulled way out by the teams that competed for few rounds (who had very high variability in the spread, from -1 to 1). For '08-09 for example, the standard deviation for the whole population is 0.31; excluding teams with fewer than 9 rounds experience, 0.22.

Based on A Numbers Game's question, I created a Lorenz curve for '08-09:


A few data points in words:

The top 5% of teams debated 19% of all rounds.
The top 10% of teams debated 34% of all rounds.
The top 20% of teams debated 54% of all rounds.
The top half of teams debated 83% of all rounds.

Monday, March 2, 2009

Debate tournament math

Here's how a high school or college debate tournament works: for the first two rounds of debating, each team is randomly assigned an opponent; for the third round, the winners of the first two rounds are assigned other winners as opponents, while losers debate losers. This system continues for several preliminary rounds, "power matching" teams against opponents with the same record of wins and losses, until the top "brackets" with winning records (7-0s and 6-1s, for example) move on to elimination rounds. Thus, the preliminary rounds are a type of Swiss system tournament, a format that is used in chess competition, too. The number of teams in each final bracket follows a perfect binomial distribution (plus or minus one team or two for odd numbers that require a team to be "pulled up" from a lower bracket).

This approach is generally felt to be fair, although it is a recognized problem that a good team could lose the first or second round and would have easier opponents all the way through. How often does this happen? A visualization helps:


Click on image for more detail.

These are the preliminary varsity policy debate results at the 2009 Harvard invitational high school tournament. (I took out names because I don't want to seem like I'm ragging on any school; I'm really just interested in the math.) Each row represents a different bracket -- the 7-0 at the top, 6-1s one row down, etc., and the 0-7 at the bottom -- and each row is sorted best speaker points (left) to worst speaker points (right) in that bracket. Each arrow represents one actual debate between two teams, pointing to the winner but in the loser's row color. Every single round is there, but I bolded the rounds that the top nine teams won. You can see how differently the top teams (the 7-0 and 6-1s) got that record. Some 6-1s, circled, defeated at least three 5-2 or better teams. Other 6-1s, in squares, defeated only one or no 5-2s. The 6-1 on the far left defeated not one team in the top 20%. Perhaps they could have, but they never even faced off against one. They made it into elimination rounds on the basis of an easier schedule than any other 6-1.

Let me make it absolutely clear, I'm not criticizing the folks who run the Harvard tournament. They do a fine job. The problem is not with their execution. I'm sure that at every point, the 6-1s were given proper opponents for their records; the problem is that some of those opponents went on to lose many of their remaining rounds and revealed their weakness. The problem is the method, which is only as good as the current record of each team accurately reflects its true strength. Since this information can't be known in advance, the only solution so far has been to repeat the process many, many times to thoroughly test and properly rank each team in preliminary rounds. Potentially, what you're looking at above is a raw sort that still contains some errors, like ABCEDLFGJIMNP... it's getting better, but there's still a need for further sorting. Consider it this way: the first round is supposed to determine whether a letter is in the first half of the alphabet or not, by picking up two letters at the same and determining which comes first. Generally speaking, this works, and A, B, C, etc., are likely to end up in the first-half pile. But what happens if the letters you pick up to compare are T and W? T will be misleadingly placed in the first-half pile, and you hope that this doesn't happen two, or three, or seven times in a row, but clearly, it can and did happen, and a team made into the top 6% without ever facing an opponent in the top 20%.

Randomness isn't enough. There needs to be an element added to power-matching that controls for strength of schedule. If you need further convincing, here are the 5-2s highlighted:

Click on image for more detail.

The circled 5-2s defeated at least one other 5-2. (It's hard to see those blue arrows, so click on the image for expansion first.) The 5-2s in squares defeated only one or two 4-3s or better -- that is, they made it into the top 20% and elimination rounds on the basis of defeating only one or two teams in the top 40%. That's quite a disparate schedule: debating other 5-2s and several 4-3s, or debating a few 4-3s and then several teams that are weaker.