One method for ranking teams that I introduced to the debate community is the logit score. The logit score is derived from a logistic regression. The logit score combines a team's record, speaker points, and its opponents' strength into a single number. Because the logit score factors in record and points, it is performance-based, but that record is adjusted by opponent strength, making the logit score more fair than record alone. A win against a good team is "worth" more than a win against a weak team. If you take the worst opponent a team beat and the best opponent it lost to, and average those together along with the team's average speaker points, then you're approximating the team's logit score. Due to how the logit score is calculated, it is the likeliest team strength that explains its results: its record and its points.
I had previously looked for empirical support for the logit score in a college debate season. I took the real results for the entire season and used them to calculate each team's logit score. I then used those to retrodict the winner in every single match-up that had actually happened, with the higher ranked by logit score team retrodicted to win the round. The logit score did this better than every other ranking method I also tested, slightly edging out median speaker points, and doing better by a goodly margin than the win-loss record. Despite this success, there was the nagging concern that the logit score was being derived from an entire season's worth of information. This empirical support could not show if the logit score would work for a single tournament.
Therefore I set out to do an experiment. I created a simulation tournament in a program, and ran and re-ran it hundreds of times. I tested various tournament conditions, from random prelims to a typical method of power-matching to pre-matching (like a round robin). I looked to see whether in these kind of conditions--using only the information available in a tournament--the logit score fared as well in comparison to record-based rankings and to speaker point-based rankings.
The results are that, in any condition, the logit score is a vast improvement on the win-loss record, but not quite as good as speaker points. It may surprise people to realize that speaker points, even though they vary considerably from judge to judge, are the best information to rank teams. A team's median speaker points isn't affected too much by one judge. Speaker points are rich data when you only have six or eight rounds to rank a team.
However, I believe many in the community would not prefer to use speaker points alone. If nothing else, ignoring wins and losses gives a perverse incentive to teams to speak pretty and ignore winning key arguments. The logit score is a solid, thoughtful compromise. The logit score is based on both wins and points, so there's no perverse incentive to ignore key arguments--nor is there an incentive to ignore effective, mellifluous communication. Although the logit score is slightly less accurate for a single tournament than speaker points alone, the logit score is far more accurate than win-loss record is. The logit score is, in other words, a vast improvement on the status quo method--a compromise in name only.
Essays on education, debate, and math instruction; neat math problems; and whatever else I get around to.
Showing posts with label rankings. Show all posts
Showing posts with label rankings. Show all posts
Saturday, April 7, 2018
Sunday, January 4, 2015
100th Post!!
My hundredth post! When I started 5 and a half years ago, I never imagined I would get here. It turns out that I have written a post, on average, every 20 days. I thought for this special occasion, I would go back to one of my original reasons for starting this blog: my dislike for the traditional methods of power matching in debate tournaments. My opinion was -- and still is -- that power matching doesn't give each debate team a fair experience at the tournament. Many debate teams make it to elimination rounds without facing good opponents. My solution was to create a strength-of-schedule pairing that power-matched but also attempted to even out schedule strength. That was five years ago. Since that time, I have come to believe that it's better to abandon power matching rather than try to improve it. The alternative is random prelims. Random prelims works for the N.S.D.A. Nationals (formerly the N.F.L.) and can even be improved with geographic mixing.
I decided to test it out with an experiment. How do different pairing methods compare at producing the actual ranking of the teams? Obviously, one only has actual rankings in an experiment. To start, I generated 200 random teams, giving each one a true strength. The true strengths were in a Normal distribution with an average of 27 points and standard deviation of 1 point. This is realistic, based on previous empirical analysis I've done. Next, I paired the teams against each other using one of four pairing methods. Each team's performance could deviate from its true strength by a random number that followed a Normal distribution, average of 0, standard deviation of 1 point. In other words, most teams would perform within +/- 1 point of their true strength about 68% of the time. This may seem like a lot but is realistic from the same empirical analysis. This deviation in performance accounts for off-rounds, surprise strategies, and judging variability (e.g., point trolls), too.
Based on the two team's strengths and factoring in their random deviations from their true strengths, I decreed a winner. Then I set up the next round using the stated pairing method. After all six rounds, I calculated each team's win/loss record, total speaker points, and median speaker points (more on this in a moment). I ran the same tournament four times, one for each pairing method: (1) simple random, (2) random within win/loss bracket, (3) high-high power matched, and (4) high-low power matched. I used the same round 1 pairing for all four methods to give them all even starting conditions. For each one of the methods, I used the results at the end to calculate a traditional ranking (win/loss, then total speaker points, then median speaker points) from 1 to 200 and also a "median points" ranking (first, median speaker points, then total speaker points, then win/loss record) from 1 to 200 -- and compared them to the true rankings. The results kind of blew my mind and switched my perspective around.
A few caveats for the nit-pickers: yes, I ignored low-point wins. Those aren't too frequent and, as you will see, including them would only make my case stronger. And yes, I ignored side constraints and I pretended like the teams were from 200 different schools. Again, using those constraints would only strengthen my case. Without further ado, here are the results:
This is the r-squared, the coefficient of determination you might have learned about in Intro to Stats, of the traditional and median points rankings to the actual ranking for each pairing method I tested out. A high r-squared is good; it means the listed ranking closely corresponds to the truth.
A couple of things to draw your attention to: (a) the median points rankings do not change much for any pairing method; (b) the median points rankings are higher than or almost equal to the traditional rankings for every pairing method; and (c) the traditional rankings are closest to true rankings for the high-low pairing, then the random within brackets pairing, then the simple random pairing, and lastly the high-high pairings.
It is actually worth looking at that last one:
Notice the clear pattern? In a high-high pairing, a ton of decent teams get screwed by getting several very hard opponents and therefore have terrible records -- these are the outliers that are very low on the y-axis (indicating true strength) but on the right side of the x-axis (indicating very poor records). Notice the one team in the far bottom right: pity the poor team that had true ranking of 19 but ended up with a 1-5 record and ranked 183rd by the traditional tiebreakers. Enragingly, a lot of weak teams somehow squeak by to great records. Notice the one team in the upper left: 148th in truth, but given several easy opponents, ending up 5-1 and ranked 21st by the traditional tiebreakers. Visually, you can see how unjust the whole high-high pairing is when coupled with using win/loss record as the primary criterion for ranking, as it is in the traditional method. The median points ranking does not suffer from the same problem; even the good team that gets several tough opponents and ends up 1-5 is not penalized in the rankings, as long as that team continued to earn high points in each one of its rounds.
For comparison, here is what the best correlation looked like:
To be sure, the correlation is far from perfect. But that's just about variability in the teams' in-round performances compared to their true strengths (that random deviation score I added). In other words, what you are looking at is just the off-days, surprises, and crappy judging that is unavoidable. It isn't really possible to do better than about 0.82 or 0.83 -- that's why the median rankings have about the same correlation, no matter what the pairing method.
On the other hand, the traditional rankings are very sensitive to the pairing method. Why? A team's record depends on both its true strength and the opponents it faces! In the high-high pairing method, many teams get unfairly hard or unfairly easy opponents. The method drives down the correlation between true strength and record by screwing some and blessing others. However, in the high-low pairing method, the assignment of opponents pushes up the correlation between true strength and record -- better teams face weaker opponents, so get a few more easy wins.
It can be a bit hard to interpret what these correlations means, so I also calculated the mean absolute deviations for each pairing method and ranking. For each team, I took its traditional rank and its true rank, found the difference, and took the absolute value. Then I averaged those to produce the mean absolute deviation (MAD). I also did the same thing for the median points rankings.
For example, for the random within bracket pairing method, the median points ranking had a MAD of 18.66. That means, on average, the median ranking was off by 18.66 places from the truth. The lower the MAD, the better.
The exact same patterns appear as in the correlations table. In general, the best we can hope for is to be within about 20 places of the truth. Given that debaters have off-rounds, and that our sample size is only six rounds, this isn't terrible: 20/200 is 10%. The true ranking is probably +/- decile from the median points ranking. Notice that the traditional rankings are sensitive to the pairing method in the exact same pattern. If one uses the traditional criteria for ranking, then the high-low pairing is best. So, was I wrong five years ago?
In both of the two tables I've given so far, the high-low pairing method plus traditional ranking was marginally superior to any pairing method plus median points ranking. But the problem is that it is not equally important to rank any team correctly. It is more important to get the top teams right. Enter the weighted rule. As I did for the MAD, I took each team's traditional rank, subtracted its true rank, and took the absolute value. But before I averaged, I divided by the team's true rank. Thus, getting a good team's results wrong by a lot was worth big negative points; getting a weak team's results by a lot was worth a few negative points. The results:
The pattern is almost the same as before, except that... high-low pairings and traditional rankings is worse (higher score) than random pairings with median rankings. This means that the high-low plus traditional combination made more mistakes ranking the best teams than the random plus median combination.
1. If you are doing high-high power matching, STOP IT RIGHT NOW. Even one round of high-high power matching is harmful. You are screwing many teams over.
2. Consider using the median points ranking instead of the traditional ranking.
On T.R.P.C., it means putting the "drop two high - drop two low speaker points" as the first criterion for ranking. (For a three- or four-round tournament, the "drop high - drop low" option is equivalent to the median. For a five- or six-round tournament, the double-drop option is equivalent to the median. For a seven- or eight-round tournament, the triple-drop option is equivalent to the median.) You can make win/loss record the second or third criterion.
All the data from my experiment show that the median ranking is simply more accurate, no matter how you pair the tournament. The win/loss record is too variable.
3. Consider not doing high-low power matching either.
It is enormously time intensive to run a power-matched tournament. In some cases, power-matching adds 2-3 hours for a six-round tournament: 30-45 minutes after rounds 2, 3, 4, and 5 -- although one or two or those lag times might occur during a food break that had to happen anyway. But 2-3 hours might be used in other ways... say, to squeeze in an extra round. Another round would, in fact, yield more data and would improve the accuracy of the results far more than stopping frequently to power match. And, as the experiment data show, high-low pairings do not improve the accuracy any more than simply switching over to median rankings. (Furthermore, I suspect that high-low pairings plus traditional rankings' accuracy peaks at around five to six preliminary rounds; my suspicion is that for longer tournaments, the accuracy starts to go down again because the brackets start to get too small.)
High-low power matching does have something to argue for it: teams get to see more opponents of similar ability levels (to themselves). But there's a counterargument: random pairings enable teams to see a wide cross-section of opponents' skill levels, and better gauge where they fall on the spectrum. Getting your butt kicked can inspire striving, and besides, your tournament should have a novice and JV division for teams that are afraid of the best opponents in the top division.
However, if you feel like high-low power matching is something you want to preserve but you do want to speed up your tournament, then go to lag-powering. For example, round 3 would be power matched, but only off of the results of round 1. You can slip round 3 pairings under the doors while round 2 is wrapping up, cutting your turnaround time drastically. You probably won't push down your accuracy too much (see how well random within brackets plus traditional rankings compares) if you lag-power -- but especially not if you use median points rankings.
4. Do everything you can to help your judges give speaker points more consistently.
If speaker points are more accurate than records, that means we ought to put more weight on speaker points AND strive to make them seem less arbitrary. Brief training sessions at the beginning of the tournament for less experienced judges, clearly delineated rubrics for speaker points, or scoring grids for various attributes of speaking all help!
I used to think that power matching started as the best way people had, when tabbing on notecards, to improve the accuracy of tournament results. Maybe that is why it got started, but as we can see, all it does is bring accuracy to parity with median rankings. Is there any other reason tabbers might have started to use power matching?
Then it dawned on me: power matching reduces the likelihood that two teams have met before will be randomly drawn against each other in later rounds. Team A might meet team B in round 1, win, then lose round 2. Team B might do the opposite and win round 2, giving the two opponents a possibility of being randomly drawn against each other for round 3 -- but overall, power matching makes it less likely than simply randomly assigning everyone in one big pool. When you're tabbing on notecards, it speeds things up considerably if this is a rare occurrence. Maybe power matching began, not with HH or HL but with the random within brackets pairing. In other words, how we pair double elimination tournaments (undefeateds and down-ones in two separate brackets, randomly assigned in each) might have gotten translated to all preliminary rounds at all tournaments.
Just a suspicion. It does make sense: high-low (or the awful high-high) brackets are hard to do on notecards, but random within brackets is easy to do.
I decided to test it out with an experiment. How do different pairing methods compare at producing the actual ranking of the teams? Obviously, one only has actual rankings in an experiment. To start, I generated 200 random teams, giving each one a true strength. The true strengths were in a Normal distribution with an average of 27 points and standard deviation of 1 point. This is realistic, based on previous empirical analysis I've done. Next, I paired the teams against each other using one of four pairing methods. Each team's performance could deviate from its true strength by a random number that followed a Normal distribution, average of 0, standard deviation of 1 point. In other words, most teams would perform within +/- 1 point of their true strength about 68% of the time. This may seem like a lot but is realistic from the same empirical analysis. This deviation in performance accounts for off-rounds, surprise strategies, and judging variability (e.g., point trolls), too.
Based on the two team's strengths and factoring in their random deviations from their true strengths, I decreed a winner. Then I set up the next round using the stated pairing method. After all six rounds, I calculated each team's win/loss record, total speaker points, and median speaker points (more on this in a moment). I ran the same tournament four times, one for each pairing method: (1) simple random, (2) random within win/loss bracket, (3) high-high power matched, and (4) high-low power matched. I used the same round 1 pairing for all four methods to give them all even starting conditions. For each one of the methods, I used the results at the end to calculate a traditional ranking (win/loss, then total speaker points, then median speaker points) from 1 to 200 and also a "median points" ranking (first, median speaker points, then total speaker points, then win/loss record) from 1 to 200 -- and compared them to the true rankings. The results kind of blew my mind and switched my perspective around.
A few caveats for the nit-pickers: yes, I ignored low-point wins. Those aren't too frequent and, as you will see, including them would only make my case stronger. And yes, I ignored side constraints and I pretended like the teams were from 200 different schools. Again, using those constraints would only strengthen my case. Without further ado, here are the results:
This is the r-squared, the coefficient of determination you might have learned about in Intro to Stats, of the traditional and median points rankings to the actual ranking for each pairing method I tested out. A high r-squared is good; it means the listed ranking closely corresponds to the truth.
A couple of things to draw your attention to: (a) the median points rankings do not change much for any pairing method; (b) the median points rankings are higher than or almost equal to the traditional rankings for every pairing method; and (c) the traditional rankings are closest to true rankings for the high-low pairing, then the random within brackets pairing, then the simple random pairing, and lastly the high-high pairings.
It is actually worth looking at that last one:
Notice the clear pattern? In a high-high pairing, a ton of decent teams get screwed by getting several very hard opponents and therefore have terrible records -- these are the outliers that are very low on the y-axis (indicating true strength) but on the right side of the x-axis (indicating very poor records). Notice the one team in the far bottom right: pity the poor team that had true ranking of 19 but ended up with a 1-5 record and ranked 183rd by the traditional tiebreakers. Enragingly, a lot of weak teams somehow squeak by to great records. Notice the one team in the upper left: 148th in truth, but given several easy opponents, ending up 5-1 and ranked 21st by the traditional tiebreakers. Visually, you can see how unjust the whole high-high pairing is when coupled with using win/loss record as the primary criterion for ranking, as it is in the traditional method. The median points ranking does not suffer from the same problem; even the good team that gets several tough opponents and ends up 1-5 is not penalized in the rankings, as long as that team continued to earn high points in each one of its rounds.
For comparison, here is what the best correlation looked like:
To be sure, the correlation is far from perfect. But that's just about variability in the teams' in-round performances compared to their true strengths (that random deviation score I added). In other words, what you are looking at is just the off-days, surprises, and crappy judging that is unavoidable. It isn't really possible to do better than about 0.82 or 0.83 -- that's why the median rankings have about the same correlation, no matter what the pairing method.
On the other hand, the traditional rankings are very sensitive to the pairing method. Why? A team's record depends on both its true strength and the opponents it faces! In the high-high pairing method, many teams get unfairly hard or unfairly easy opponents. The method drives down the correlation between true strength and record by screwing some and blessing others. However, in the high-low pairing method, the assignment of opponents pushes up the correlation between true strength and record -- better teams face weaker opponents, so get a few more easy wins.
It can be a bit hard to interpret what these correlations means, so I also calculated the mean absolute deviations for each pairing method and ranking. For each team, I took its traditional rank and its true rank, found the difference, and took the absolute value. Then I averaged those to produce the mean absolute deviation (MAD). I also did the same thing for the median points rankings.
For example, for the random within bracket pairing method, the median points ranking had a MAD of 18.66. That means, on average, the median ranking was off by 18.66 places from the truth. The lower the MAD, the better.
The exact same patterns appear as in the correlations table. In general, the best we can hope for is to be within about 20 places of the truth. Given that debaters have off-rounds, and that our sample size is only six rounds, this isn't terrible: 20/200 is 10%. The true ranking is probably +/- decile from the median points ranking. Notice that the traditional rankings are sensitive to the pairing method in the exact same pattern. If one uses the traditional criteria for ranking, then the high-low pairing is best. So, was I wrong five years ago?
In both of the two tables I've given so far, the high-low pairing method plus traditional ranking was marginally superior to any pairing method plus median points ranking. But the problem is that it is not equally important to rank any team correctly. It is more important to get the top teams right. Enter the weighted rule. As I did for the MAD, I took each team's traditional rank, subtracted its true rank, and took the absolute value. But before I averaged, I divided by the team's true rank. Thus, getting a good team's results wrong by a lot was worth big negative points; getting a weak team's results by a lot was worth a few negative points. The results:
The pattern is almost the same as before, except that... high-low pairings and traditional rankings is worse (higher score) than random pairings with median rankings. This means that the high-low plus traditional combination made more mistakes ranking the best teams than the random plus median combination.
What are the take-aways?
1. If you are doing high-high power matching, STOP IT RIGHT NOW. Even one round of high-high power matching is harmful. You are screwing many teams over.
2. Consider using the median points ranking instead of the traditional ranking.
On T.R.P.C., it means putting the "drop two high - drop two low speaker points" as the first criterion for ranking. (For a three- or four-round tournament, the "drop high - drop low" option is equivalent to the median. For a five- or six-round tournament, the double-drop option is equivalent to the median. For a seven- or eight-round tournament, the triple-drop option is equivalent to the median.) You can make win/loss record the second or third criterion.
All the data from my experiment show that the median ranking is simply more accurate, no matter how you pair the tournament. The win/loss record is too variable.
3. Consider not doing high-low power matching either.
It is enormously time intensive to run a power-matched tournament. In some cases, power-matching adds 2-3 hours for a six-round tournament: 30-45 minutes after rounds 2, 3, 4, and 5 -- although one or two or those lag times might occur during a food break that had to happen anyway. But 2-3 hours might be used in other ways... say, to squeeze in an extra round. Another round would, in fact, yield more data and would improve the accuracy of the results far more than stopping frequently to power match. And, as the experiment data show, high-low pairings do not improve the accuracy any more than simply switching over to median rankings. (Furthermore, I suspect that high-low pairings plus traditional rankings' accuracy peaks at around five to six preliminary rounds; my suspicion is that for longer tournaments, the accuracy starts to go down again because the brackets start to get too small.)
High-low power matching does have something to argue for it: teams get to see more opponents of similar ability levels (to themselves). But there's a counterargument: random pairings enable teams to see a wide cross-section of opponents' skill levels, and better gauge where they fall on the spectrum. Getting your butt kicked can inspire striving, and besides, your tournament should have a novice and JV division for teams that are afraid of the best opponents in the top division.
However, if you feel like high-low power matching is something you want to preserve but you do want to speed up your tournament, then go to lag-powering. For example, round 3 would be power matched, but only off of the results of round 1. You can slip round 3 pairings under the doors while round 2 is wrapping up, cutting your turnaround time drastically. You probably won't push down your accuracy too much (see how well random within brackets plus traditional rankings compares) if you lag-power -- but especially not if you use median points rankings.
4. Do everything you can to help your judges give speaker points more consistently.
If speaker points are more accurate than records, that means we ought to put more weight on speaker points AND strive to make them seem less arbitrary. Brief training sessions at the beginning of the tournament for less experienced judges, clearly delineated rubrics for speaker points, or scoring grids for various attributes of speaking all help!
A changed perspective
I used to think that power matching started as the best way people had, when tabbing on notecards, to improve the accuracy of tournament results. Maybe that is why it got started, but as we can see, all it does is bring accuracy to parity with median rankings. Is there any other reason tabbers might have started to use power matching?
Then it dawned on me: power matching reduces the likelihood that two teams have met before will be randomly drawn against each other in later rounds. Team A might meet team B in round 1, win, then lose round 2. Team B might do the opposite and win round 2, giving the two opponents a possibility of being randomly drawn against each other for round 3 -- but overall, power matching makes it less likely than simply randomly assigning everyone in one big pool. When you're tabbing on notecards, it speeds things up considerably if this is a rare occurrence. Maybe power matching began, not with HH or HL but with the random within brackets pairing. In other words, how we pair double elimination tournaments (undefeateds and down-ones in two separate brackets, randomly assigned in each) might have gotten translated to all preliminary rounds at all tournaments.
Just a suspicion. It does make sense: high-low (or the awful high-high) brackets are hard to do on notecards, but random within brackets is easy to do.
Thursday, April 17, 2014
A presentation on ranking methods
Here is the annotated PowerPoint from a recent presentation I gave on ranking methods.
Here's a post on debate tabulation. And here's one on debate upsets. A post about weighted wins. Another post about network graphs.
Here's a post on debate tabulation. And here's one on debate upsets. A post about weighted wins. Another post about network graphs.
Saturday, December 21, 2013
Academic journal rankings
I recently read an article on academic journal (and article) rankings (and another one here). You might expect, given that I am a connoisseur of ranking methods, that I would support the mission to rank academic journals (and articles). I do not. No matter the method, ranking the "impact" is an over-simplification to popularity. I found out from The Atlantic that the system of impact ranking had been invented by librarians in the '70s merely to evaluate which journals were the most important in each field. Now tenure committees are using them to decide whether an academic has done good work. Huh? Shouldn't a tenure committee actually, you know, read the articles in question to decide whether the work is good? Isn't that the committee's job: to bring expert opinion to bear and to make subtle, careful, and thoughtful judgments?
The problem, of course, is that there are actually two different variables at play, rather than "impact." The real variables are trustworthiness and importance. If the research method is sound, if the authors understood the literature correctly and asked reasonable follow-up questions, if the data are interpreted correctly, then the research is trustworthy. There is a lot of trustworthy research that is done, period. But there is a lot of trustworthy research that is never, ever published because it fails to meet a high threshold of importance. Novel results that turn a field on its head are important. Results that create a new field of research are important. Results that really expand or complicate a field are important. Current academic journals are heavily biased to publish results that are important by this definition -- novel, theory-creating or -expanding or -re-defining. This tendency is one that makes a lot of sense, but in the aggregate, it means that published research is less likely to be trustworthy, since "important" results can often reflect a design flaw in the research. At the very, very least, academic committees ought to score articles along both axes -- trustworthiness and importance -- and journals (or e-publications) ought to publish a lot MORE trustworthy, UNimportant articles. I, for one, would like to know, "Hey, 50 researchers had good but boring results with this theory," to know that an idea has been confirmed and re-confirmed by experiment before building upon it. I would be especially interested to know that a theory had been tested and failed over and over again, so I would not waste my time.
But there is an even more radical solution. What if e-publication of research took full advantage of the tools that already exist to make publication as beneficial as possible all around, to really build an easy-to-use knowledge engine? Three specific capabilities come to mind: 1) links, 2) data, and 3) commenting. The links part is clear: if enough authors went to e-publication, we could track down an article's set of citations more easily than currently. The data part is clear, too: Google docs' spreadsheet and Tableau, for example, are both ways that researchers could publish their raw data by embedding it in the article. Often I have wished I could look at the raw data and run the numbers for myself. Errors do happen. And I think it would cut down on the statistical chicanery if the norm was to publish raw data as well as the polished tables and p-values.
The best addition of all would be to add commenting. What if I could read an article with comments by other experts in the field? (Of course, there would need to be some kind of threshold set for ability to comment.) If I could read a paragraph on the research method, then immediately see several other expert researchers' key concerns about it, I would be in a better place to judge the method as trustworthy or not. If I could read a paragraph about the theory being advanced, then immediately see several other experts' alternatives to said theory, again, I would be in a better place to judge the importance of the article. If a specific paragraph was linked to or bookmarked by a bunch of other academic articles, well that would tell be even more. Comments, being qualitative, would provide a wealth of information that the quantitative information about pageviews, citations, etc., can never do. Yet it seems like the direction universities (specifically, tenure committees) want to go is simply ranking. Ranks make sense for sports, or debate teams, because it is a closed system with only one outcome: wins. Academic publication is unlike this and ought to be "scored" differently, or not even really scored at all. My fantasy is that articles -- one per webpage -- would be the quanta of a much wider ranging debate in the links, highlighting, quotations, line-item comments, and end-of-article comments and reviewer votes. My fantasy is that a flawed article is obviously labeled as flawed, so that readers could immediately tread with caution. If you're more cynical, then at least you could hope for a system where readers know which articles represent solid, consensus-based work and which are contentious. The current system, where flawed work is almost never retracted and a reader has to work hard to find out whether an article has been retracted or discredited, serves no one's interest.
Maybe, as this article implies, the problem is that journals are doing what's best for the journal (and these publishers are for-profit), not what's best for science. Here's one more article, on the brave new world of publishing that might await. Another one, on the statistical importance of publishing null results; finally, one more on the problems with PLOS.
The problem, of course, is that there are actually two different variables at play, rather than "impact." The real variables are trustworthiness and importance. If the research method is sound, if the authors understood the literature correctly and asked reasonable follow-up questions, if the data are interpreted correctly, then the research is trustworthy. There is a lot of trustworthy research that is done, period. But there is a lot of trustworthy research that is never, ever published because it fails to meet a high threshold of importance. Novel results that turn a field on its head are important. Results that create a new field of research are important. Results that really expand or complicate a field are important. Current academic journals are heavily biased to publish results that are important by this definition -- novel, theory-creating or -expanding or -re-defining. This tendency is one that makes a lot of sense, but in the aggregate, it means that published research is less likely to be trustworthy, since "important" results can often reflect a design flaw in the research. At the very, very least, academic committees ought to score articles along both axes -- trustworthiness and importance -- and journals (or e-publications) ought to publish a lot MORE trustworthy, UNimportant articles. I, for one, would like to know, "Hey, 50 researchers had good but boring results with this theory," to know that an idea has been confirmed and re-confirmed by experiment before building upon it. I would be especially interested to know that a theory had been tested and failed over and over again, so I would not waste my time.
But there is an even more radical solution. What if e-publication of research took full advantage of the tools that already exist to make publication as beneficial as possible all around, to really build an easy-to-use knowledge engine? Three specific capabilities come to mind: 1) links, 2) data, and 3) commenting. The links part is clear: if enough authors went to e-publication, we could track down an article's set of citations more easily than currently. The data part is clear, too: Google docs' spreadsheet and Tableau, for example, are both ways that researchers could publish their raw data by embedding it in the article. Often I have wished I could look at the raw data and run the numbers for myself. Errors do happen. And I think it would cut down on the statistical chicanery if the norm was to publish raw data as well as the polished tables and p-values.
The best addition of all would be to add commenting. What if I could read an article with comments by other experts in the field? (Of course, there would need to be some kind of threshold set for ability to comment.) If I could read a paragraph on the research method, then immediately see several other expert researchers' key concerns about it, I would be in a better place to judge the method as trustworthy or not. If I could read a paragraph about the theory being advanced, then immediately see several other experts' alternatives to said theory, again, I would be in a better place to judge the importance of the article. If a specific paragraph was linked to or bookmarked by a bunch of other academic articles, well that would tell be even more. Comments, being qualitative, would provide a wealth of information that the quantitative information about pageviews, citations, etc., can never do. Yet it seems like the direction universities (specifically, tenure committees) want to go is simply ranking. Ranks make sense for sports, or debate teams, because it is a closed system with only one outcome: wins. Academic publication is unlike this and ought to be "scored" differently, or not even really scored at all. My fantasy is that articles -- one per webpage -- would be the quanta of a much wider ranging debate in the links, highlighting, quotations, line-item comments, and end-of-article comments and reviewer votes. My fantasy is that a flawed article is obviously labeled as flawed, so that readers could immediately tread with caution. If you're more cynical, then at least you could hope for a system where readers know which articles represent solid, consensus-based work and which are contentious. The current system, where flawed work is almost never retracted and a reader has to work hard to find out whether an article has been retracted or discredited, serves no one's interest.
Maybe, as this article implies, the problem is that journals are doing what's best for the journal (and these publishers are for-profit), not what's best for science. Here's one more article, on the brave new world of publishing that might await. Another one, on the statistical importance of publishing null results; finally, one more on the problems with PLOS.
Labels:
academic journals,
academic research,
open source,
rankings
Monday, January 7, 2013
Impactranks
I recently heard of Josh Clark's new project impactranks.com. It is basically a coaches' poll. I respect the intention, but I strongly disagree with the method. Does anyone think the B.C.S. does a good job? It is largely based on a coaches' poll. The problem is that coaches do not see enough other teams play (or debate) and end up reflecting the perceived reputation of programs. There are sounder methods to rank teams that reflect actual results: wins, losses, and points. I would very much recommend reading, Who's #1? The Science of Rating and Ranking, by Langville and Meyer, two math professors, for many of the methods. My own method for debate is the weighted win method, explained here and here, in which whom a team beats matters even more than the raw win percentage.
One objection that people might point out is that debaters have so few opponents during the year that there is not enough data to think a mathematical method is any better than a poll. However, I would point out that the same problem exists in college football and professional football, yet mathematical methods based on opponent strength do work reasonably well.
Another objection is that debate has a lot of upsets, because of bad judges, off rounds, or unusual one-time tricks. True, but so does football, and the mathematical methods still work just fine. (In fact, I was even able to calculate an upset rate for college debate -- about 20%.) The other thought that I have had is to use the data on results to evaluate judges at the same time as debaters. Judges who return consistently unusual results (giving wins to worse teams with a lot of regularity) would lower their rating, so a loss from such a judge would not penalize a team by much.
The only difficulty to using these methods is that high school debate records are not stored in a clear format that shows: the two opponents, the judge(s), and the results. But the college records are. So here is my challenge to any reader. Before the N.D.T., I am going to rank all the teams, based solely on the results from the year, and publish that ranking. We will see how many results my rankings correctly predict (given the upset rate of 20%, anything in that neighborhood or better would be excellent). If any reader wants to come up with their own ranking, we will compare the results. You have until March 27th!
One objection that people might point out is that debaters have so few opponents during the year that there is not enough data to think a mathematical method is any better than a poll. However, I would point out that the same problem exists in college football and professional football, yet mathematical methods based on opponent strength do work reasonably well.
Another objection is that debate has a lot of upsets, because of bad judges, off rounds, or unusual one-time tricks. True, but so does football, and the mathematical methods still work just fine. (In fact, I was even able to calculate an upset rate for college debate -- about 20%.) The other thought that I have had is to use the data on results to evaluate judges at the same time as debaters. Judges who return consistently unusual results (giving wins to worse teams with a lot of regularity) would lower their rating, so a loss from such a judge would not penalize a team by much.
The only difficulty to using these methods is that high school debate records are not stored in a clear format that shows: the two opponents, the judge(s), and the results. But the college records are. So here is my challenge to any reader. Before the N.D.T., I am going to rank all the teams, based solely on the results from the year, and publish that ranking. We will see how many results my rankings correctly predict (given the upset rate of 20%, anything in that neighborhood or better would be excellent). If any reader wants to come up with their own ranking, we will compare the results. You have until March 27th!
Wednesday, January 20, 2010
Differing opponent strengths
Below is a graphic of a traditionally-run tournament.

The horizontal axis shows each team's final strength; the vertical axis shows each team's average opponent strength; the size of the bubble shows, for each team, the standard deviation of its opponents' strengths. A small bubble represents a team that debated opponents that were all very close together in strength. A large bubble represents a team that debated a wide cross-section of opponents, some weak and some strong.
I think about how a debate tournament ought to look, if it's paired fairly. It seems to me that every team ought to have a good cross-section of opponents. Thus, a fair tournament would be like a partial round robin. We would know that 3-3 teams were truly middle-of-the-pack because of their abilities, not because they got an unfair draw. The bubbles in the diagram would be bigger (each team sees a true cross-sections of opponents) and closer to the horizontal line (average opponent strength for each team would be closer to the overall average opponent strength).
It's relatively easy to pair a tournament like this, even on the fly. To pair a round, you can look at each team's opponents and decide what is missing so far. After three rounds, a team might have debated a 0-3, a 2-1, and a 3-0 opponent; they would now debate a 1-2 opponent. It is true that the opponent records change after the fourth round, but the process is repeated, and by the end, most teams will debate a decent cross-section of opponents from 0-6 to 6-0. Of course, traditional tournaments do not do this; teams debate opponents within brackets. Why?
The reason is that brackets increase the accuracy of rankings. Consider a 4-2 team. Does it deserve to break? If the tournament pairs it against a representative cross-section, this team would debate a 6-0 opponent, a 5-1, a 4-2, a 3-3, etc. There's only one opponent with an equal record -- but it's precisely the comparisons to very similarly-abled opponents that shed the most accurate information about a team's true strength. In a brackets system, the same team would likely debate several 4-2 opponents. There are more points of comparison, allowing for finer rankings. The downside, though, is that a team could go through the preliminary rounds of the tournament debating opponents that are all at exactly the same level. It seems to me like something valuable would be lost.
Of course, these two virtues -- fairness and accuracy -- trade off. You can't maximize both. But there are several ways to get a reasonable equilibrium. For example, pair odd rounds to have every team debate a reasonable cross-section of opponents (i.e., across brackets), and pair even rounds to increase accuracy of rankings (i.e., within brackets).

The horizontal axis shows each team's final strength; the vertical axis shows each team's average opponent strength; the size of the bubble shows, for each team, the standard deviation of its opponents' strengths. A small bubble represents a team that debated opponents that were all very close together in strength. A large bubble represents a team that debated a wide cross-section of opponents, some weak and some strong.
I think about how a debate tournament ought to look, if it's paired fairly. It seems to me that every team ought to have a good cross-section of opponents. Thus, a fair tournament would be like a partial round robin. We would know that 3-3 teams were truly middle-of-the-pack because of their abilities, not because they got an unfair draw. The bubbles in the diagram would be bigger (each team sees a true cross-sections of opponents) and closer to the horizontal line (average opponent strength for each team would be closer to the overall average opponent strength).
It's relatively easy to pair a tournament like this, even on the fly. To pair a round, you can look at each team's opponents and decide what is missing so far. After three rounds, a team might have debated a 0-3, a 2-1, and a 3-0 opponent; they would now debate a 1-2 opponent. It is true that the opponent records change after the fourth round, but the process is repeated, and by the end, most teams will debate a decent cross-section of opponents from 0-6 to 6-0. Of course, traditional tournaments do not do this; teams debate opponents within brackets. Why?
The reason is that brackets increase the accuracy of rankings. Consider a 4-2 team. Does it deserve to break? If the tournament pairs it against a representative cross-section, this team would debate a 6-0 opponent, a 5-1, a 4-2, a 3-3, etc. There's only one opponent with an equal record -- but it's precisely the comparisons to very similarly-abled opponents that shed the most accurate information about a team's true strength. In a brackets system, the same team would likely debate several 4-2 opponents. There are more points of comparison, allowing for finer rankings. The downside, though, is that a team could go through the preliminary rounds of the tournament debating opponents that are all at exactly the same level. It seems to me like something valuable would be lost.
Of course, these two virtues -- fairness and accuracy -- trade off. You can't maximize both. But there are several ways to get a reasonable equilibrium. For example, pair odd rounds to have every team debate a reasonable cross-section of opponents (i.e., across brackets), and pair even rounds to increase accuracy of rankings (i.e., within brackets).
Labels:
brackets,
debate tournaments,
fairness,
rankings,
round robins
Tuesday, August 4, 2009
Network graph theory and debate tournaments
I recently saw A Numbers Game's charts of debate tournaments, which look like modified network graphs, and it inspired me to post about some thoughts I'd had a few years ago about graph theory and debate tournaments. (Network) graph theory is a fascinating branch of mathematics, and it is directly applicable to debate tournaments. Basically, a network graph is anything that maps all the pathways (edges) between some nodes (vertices). A network graph of a debate tournament shows the match-ups (edges) between the teams at the tournament (vertices). An edge without an arrow would just show a match-up between two teams; an edge with an arrow would indicate the winner of it. A network graph can show the results of all six (or eight or whatever) preliminary rounds simultaneously. You can see a lot of interesting patterns that would be hard to notice any other way; I've used network graphs as a way to illustrate how a tournament played out.
You might think, since graphs are a complete mapping of all the preliminary round results, that they could be used to rank all the teams at a tournament from first seed to worst. However, if you tried to use a graph this way, you'd run into three different problems. Graphs are best for illustrative purposes (and one alternative use I'll suggest at the end).
The first problem with trying to use a graph to seed the teams is that you still need a lot of tiebreakers. Here's a simple example:

As indicated by the graph, team A beat B and C, and teams B and C both beat D. The problem is that B and C can't be ranked against each other from this information alone. Given the results, the final rankings must contain the sequences {A, B, D} and {A, C, D}, but both {A, B, C, D} and {A, C, B, D} satisfy this. The problem is that, except for round robins, there will always be teams that didn't meet at the tournament, and thus, pairs of vertices without edges.
Ok, so, you still need tiebreakers, but a network graph as a ranking mechanism presents a second, bigger problem: "contradictions." Let's say we run a really small tournament, and we must force A and B to debate twice. Here's one possible outcome that creates a contradiction:

A (aff.) beat B, but B (aff.) beat A. If we try to use the network graph to rank these two teams, we'd be forced to give two contradictory statements. (It's worth remembering that a team's strength is not invariant: teams can be better on one side than the other, they can have off rounds or lucky rounds, and judges also play a factor.) This is simple enough to resolve: we can use speaker points to rank one team higher than the other. Here's a slightly more complex version of the same problem:

A beat B (rd 1), B beat C (rd 2), and C beat A (rd 3). There's a contradiction here also. It will have to be resolved by "upsetting" one result, that is, a team must be ranked below an opponent whom they beat. For example, a final ranking might be {A, B, C}, which means that C's victory over A was "upset" or "overruled" because A had more overall wins, better speaker points, etc., than B or C. Bear in mind that doing a graph didn't create this problem; it merely exposed it. We're mostly oblivious to it because cume sheets show wins, points, etc., and don't provide a network graph. Look carefully through the results of nearly any tournament, and I'd be willing to beat you can find at least a few teams who are ranked below (based on overall record, speaks, etc.) an opponent they beat. Here's another example:

Any such type of contradiction, no matter how many teams, is known as a simple cyclical graph, if there's only one cyclical pathway. The good news is that, however long a simple cyclical graph, only one result has to be upset: {A, B, C, D} only upsets D's victory over A. (A complex cyclical graph would contain multiple, intersecting cycles and be a much bigger headache, but my hunch is that they are extremely rare.)
The third problem with using a network graph as a ranking mechanism is perhaps its biggest, if not mathematically, then for the debate community to stomach:

Notice the problem? There are no cycles; everything flows in only one direction. The rankings, according to the graph, are unambiguously {A, B, C, D, E, F}. Here it is again:

C beat D, but at any tournament, D would be ranked higher, because C is 1-2 and D is 2-1. In the situation I graphed, C is the better team but had a tougher schedule. (It's also possible to create an example in which D is the better team but just had an off-round or a crazy judge.) There's no cycle in the graph, so there's no need to have an upset, but I believe every debater and coach would say D should be ranked higher. Total wins as the first method of ranking is sacrosanct to the debate community, and with good reason. Total wins is clear and unambiguous; a network graph is abstract and complex. I know I would never want a tournament to start interpreting results with a network graph; it would create too much room for error when deciding breaks.
This last problem is insurmountable. I think the lesson is that we need to take every effort to make schedule strength more equitable, but there's no mathematical way to jigger results to somehow weight or re-interpret an unfair set of pairings. The cat's out of the bag, and we just have to say, "Sorry, tournaments aren't always fair," at that point. The bottom line is that network graphs will never and ought not be used for tournament rankings. But there's is an alternative use for graph theory: as a tiebreaker at round robins.
Let's say a six-team round robin has these final results:
Team Record
A 4-1
B 3-2
C 3-2
D 3-2
E 2-3
F 0-5
There's a three-way tie for second place. My suggestion is that round robins can use an idea from graph theory as the first tiebreaker, before resorting to speaker points. I believe that direct results ought to be respected as much as possible, and in round robins, you have direct results for every head-to-head match up. Who cares that team C spoke prettier than B or D if C lost both of those rounds? Here's the graph of this hypothetical round robin:

Now, you may notice that C beat A, which suggests that C is the best of the three, or you may focus on the fact that D beat both B and C, which suggests that D is the best. Both may catch your eye, but there's a way to quantify an exact result. But first, you need to turn the graph into an adjacency matrix, like so:

A column represents a team's wins; a row, its losses. Hence, reading down column A, you see that team A beat B, D, E, and F; reading across column A, you see that team A lost to C. The adjacency matrix contains all the information that the graph does; it's merely another representation of the same data. You might also notice that there's an easy way to double-check the matrix:

You can add to make sure you have the right number of wins in the column and losses in the row for each team. (The adjacency matrix is easy to do for multi-ballot round robins; just put the ballots won into the correct spot for each team.)
Now, we need to test possible rankings. There are six possibilities created by the three-way tie: {A, B, C, D, E, F}, {A, B, D, C, E, F}, {A, C, B, D, E, F}, {A, C, D, B, E, F}, {A, D, B, C, E, F}, and {A, D, C, B, E, F}. We create an adjacency matrix for each possible ranking, a matrix of the hypothetical results that perfectly consistent the ranking order with no upsets, ties, or ambiguity. For example, for the possible ranking {A, B, C, D, E, F}, the matrix is:

The matrix shows results that would be perfectly consistent with the ranking. Now we can calculate a score for how well this ranking corresponds to the actual tournament results:

Subtract the possible ranking matrix [R] from the actual tournament results [T] and count the -1s. (You can count the +1s also; they are symmetrical.) According to this ranking, there were four "upsets": that C actually beat A, that D actually beat B, that D actually beat C, and that E actually beat D. If a tournament accepts the ranking {A, B, C, D, E, F}, it "overturns" four direct results.
How do the other rankings compare? Using the same method for each "prediction," the rankings can be compared:
Rank order "Upsets"
{A, B, C, D, E, F} 4
{A, B, D, C, E, F} 3
{A, C, B, D, E, F} 5
{A, C, D, B, E, F} 4
{A, D, B, C, E, F} 2
{A, D, C, B, E, F} 3
The best ranking is obviously {A, D, B, C, E, F}. It upsets the fewest direct results. In fact, the two that are upset make perfect sense:

that C (3-2) actually beat A (4-1) and that E (2-3) actually beat D (3-2). In a sense, these were unavoidable (or you might even say real) upsets, not artifacts created by a poor ranking. Anyway, the point is, the three-way tie could be broken in this case without resorting to speaker points. This algorithm would be relatively easy to program into a tab program for round robins.
Of course, some ties are cannot be resolved by this method, namely, if the ties are really contradictions created by cycles. There's still a place for speaker points and other tiebreakers.
You might think, since graphs are a complete mapping of all the preliminary round results, that they could be used to rank all the teams at a tournament from first seed to worst. However, if you tried to use a graph this way, you'd run into three different problems. Graphs are best for illustrative purposes (and one alternative use I'll suggest at the end).
The first problem with trying to use a graph to seed the teams is that you still need a lot of tiebreakers. Here's a simple example:

As indicated by the graph, team A beat B and C, and teams B and C both beat D. The problem is that B and C can't be ranked against each other from this information alone. Given the results, the final rankings must contain the sequences {A, B, D} and {A, C, D}, but both {A, B, C, D} and {A, C, B, D} satisfy this. The problem is that, except for round robins, there will always be teams that didn't meet at the tournament, and thus, pairs of vertices without edges.
Ok, so, you still need tiebreakers, but a network graph as a ranking mechanism presents a second, bigger problem: "contradictions." Let's say we run a really small tournament, and we must force A and B to debate twice. Here's one possible outcome that creates a contradiction:

A (aff.) beat B, but B (aff.) beat A. If we try to use the network graph to rank these two teams, we'd be forced to give two contradictory statements. (It's worth remembering that a team's strength is not invariant: teams can be better on one side than the other, they can have off rounds or lucky rounds, and judges also play a factor.) This is simple enough to resolve: we can use speaker points to rank one team higher than the other. Here's a slightly more complex version of the same problem:

A beat B (rd 1), B beat C (rd 2), and C beat A (rd 3). There's a contradiction here also. It will have to be resolved by "upsetting" one result, that is, a team must be ranked below an opponent whom they beat. For example, a final ranking might be {A, B, C}, which means that C's victory over A was "upset" or "overruled" because A had more overall wins, better speaker points, etc., than B or C. Bear in mind that doing a graph didn't create this problem; it merely exposed it. We're mostly oblivious to it because cume sheets show wins, points, etc., and don't provide a network graph. Look carefully through the results of nearly any tournament, and I'd be willing to beat you can find at least a few teams who are ranked below (based on overall record, speaks, etc.) an opponent they beat. Here's another example:

Any such type of contradiction, no matter how many teams, is known as a simple cyclical graph, if there's only one cyclical pathway. The good news is that, however long a simple cyclical graph, only one result has to be upset: {A, B, C, D} only upsets D's victory over A. (A complex cyclical graph would contain multiple, intersecting cycles and be a much bigger headache, but my hunch is that they are extremely rare.)
The third problem with using a network graph as a ranking mechanism is perhaps its biggest, if not mathematically, then for the debate community to stomach:

Notice the problem? There are no cycles; everything flows in only one direction. The rankings, according to the graph, are unambiguously {A, B, C, D, E, F}. Here it is again:

C beat D, but at any tournament, D would be ranked higher, because C is 1-2 and D is 2-1. In the situation I graphed, C is the better team but had a tougher schedule. (It's also possible to create an example in which D is the better team but just had an off-round or a crazy judge.) There's no cycle in the graph, so there's no need to have an upset, but I believe every debater and coach would say D should be ranked higher. Total wins as the first method of ranking is sacrosanct to the debate community, and with good reason. Total wins is clear and unambiguous; a network graph is abstract and complex. I know I would never want a tournament to start interpreting results with a network graph; it would create too much room for error when deciding breaks.
This last problem is insurmountable. I think the lesson is that we need to take every effort to make schedule strength more equitable, but there's no mathematical way to jigger results to somehow weight or re-interpret an unfair set of pairings. The cat's out of the bag, and we just have to say, "Sorry, tournaments aren't always fair," at that point. The bottom line is that network graphs will never and ought not be used for tournament rankings. But there's is an alternative use for graph theory: as a tiebreaker at round robins.
Let's say a six-team round robin has these final results:
Team Record
A 4-1
B 3-2
C 3-2
D 3-2
E 2-3
F 0-5
There's a three-way tie for second place. My suggestion is that round robins can use an idea from graph theory as the first tiebreaker, before resorting to speaker points. I believe that direct results ought to be respected as much as possible, and in round robins, you have direct results for every head-to-head match up. Who cares that team C spoke prettier than B or D if C lost both of those rounds? Here's the graph of this hypothetical round robin:

Now, you may notice that C beat A, which suggests that C is the best of the three, or you may focus on the fact that D beat both B and C, which suggests that D is the best. Both may catch your eye, but there's a way to quantify an exact result. But first, you need to turn the graph into an adjacency matrix, like so:

A column represents a team's wins; a row, its losses. Hence, reading down column A, you see that team A beat B, D, E, and F; reading across column A, you see that team A lost to C. The adjacency matrix contains all the information that the graph does; it's merely another representation of the same data. You might also notice that there's an easy way to double-check the matrix:

You can add to make sure you have the right number of wins in the column and losses in the row for each team. (The adjacency matrix is easy to do for multi-ballot round robins; just put the ballots won into the correct spot for each team.)
Now, we need to test possible rankings. There are six possibilities created by the three-way tie: {A, B, C, D, E, F}, {A, B, D, C, E, F}, {A, C, B, D, E, F}, {A, C, D, B, E, F}, {A, D, B, C, E, F}, and {A, D, C, B, E, F}. We create an adjacency matrix for each possible ranking, a matrix of the hypothetical results that perfectly consistent the ranking order with no upsets, ties, or ambiguity. For example, for the possible ranking {A, B, C, D, E, F}, the matrix is:

The matrix shows results that would be perfectly consistent with the ranking. Now we can calculate a score for how well this ranking corresponds to the actual tournament results:

Subtract the possible ranking matrix [R] from the actual tournament results [T] and count the -1s. (You can count the +1s also; they are symmetrical.) According to this ranking, there were four "upsets": that C actually beat A, that D actually beat B, that D actually beat C, and that E actually beat D. If a tournament accepts the ranking {A, B, C, D, E, F}, it "overturns" four direct results.
How do the other rankings compare? Using the same method for each "prediction," the rankings can be compared:
Rank order "Upsets"
{A, B, C, D, E, F} 4
{A, B, D, C, E, F} 3
{A, C, B, D, E, F} 5
{A, C, D, B, E, F} 4
{A, D, B, C, E, F} 2
{A, D, C, B, E, F} 3
The best ranking is obviously {A, D, B, C, E, F}. It upsets the fewest direct results. In fact, the two that are upset make perfect sense:

that C (3-2) actually beat A (4-1) and that E (2-3) actually beat D (3-2). In a sense, these were unavoidable (or you might even say real) upsets, not artifacts created by a poor ranking. Anyway, the point is, the three-way tie could be broken in this case without resorting to speaker points. This algorithm would be relatively easy to program into a tab program for round robins.
Of course, some ties are cannot be resolved by this method, namely, if the ties are really contradictions created by cycles. There's still a place for speaker points and other tiebreakers.
Subscribe to:
Posts (Atom)