Showing posts with label team strength. Show all posts
Showing posts with label team strength. Show all posts

Sunday, January 4, 2015

100th Post!!

My hundredth post! When I started 5 and a half years ago, I never imagined I would get here. It turns out that I have written a post, on average, every 20 days. I thought for this special occasion, I would go back to one of my original reasons for starting this blog: my dislike for the traditional methods of power matching in debate tournaments. My opinion was -- and still is -- that power matching doesn't give each debate team a fair experience at the tournament. Many debate teams make it to elimination rounds without facing good opponents. My solution was to create a strength-of-schedule pairing that power-matched but also attempted to even out schedule strength. That was five years ago. Since that time, I have come to believe that it's better to abandon power matching rather than try to improve it. The alternative is random prelims. Random prelims works for the N.S.D.A. Nationals (formerly the N.F.L.) and can even be improved with geographic mixing.

I decided to test it out with an experiment. How do different pairing methods compare at producing the actual ranking of the teams? Obviously, one only has actual rankings in an experiment. To start, I generated 200 random teams, giving each one a true strength. The true strengths were in a Normal distribution with an average of 27 points and standard deviation of 1 point. This is realistic, based on previous empirical analysis I've done. Next, I paired the teams against each other using one of four pairing methods. Each team's performance could deviate from its true strength by a random number that followed a Normal distribution, average of 0, standard deviation of 1 point. In other words, most teams would perform within +/- 1 point of their true strength about 68% of the time. This may seem like a lot but is realistic from the same empirical analysis. This deviation in performance accounts for off-rounds, surprise strategies, and judging variability (e.g., point trolls), too.

Based on the two team's strengths and factoring in their random deviations from their true strengths, I decreed a winner. Then I set up the next round using the stated pairing method. After all six rounds, I calculated each team's win/loss record, total speaker points, and median speaker points (more on this in a moment). I ran the same tournament four times, one for each pairing method: (1) simple random, (2) random within win/loss bracket, (3) high-high power matched, and (4) high-low power matched. I used the same round 1 pairing for all four methods to give them all even starting conditions. For each one of the methods, I used the results at the end to calculate a traditional ranking (win/loss, then total speaker points, then median speaker points) from 1 to 200 and also a "median points" ranking (first, median speaker points, then total speaker points, then win/loss record) from 1 to 200 -- and compared them to the true rankings. The results kind of blew my mind and switched my perspective around.

A few caveats for the nit-pickers: yes, I ignored low-point wins. Those aren't too frequent and, as you will see, including them would only make my case stronger. And yes, I ignored side constraints and I pretended like the teams were from 200 different schools. Again, using those constraints would only strengthen my case. Without further ado, here are the results:


This is the r-squared, the coefficient of determination you might have learned about in Intro to Stats, of the traditional and median points rankings to the actual ranking for each pairing method I tested out. A high r-squared is good; it means the listed ranking closely corresponds to the truth.

A couple of things to draw your attention to:  (a) the median points rankings do not change much for any pairing method; (b) the median points rankings are higher than or almost equal to the traditional rankings for every pairing method; and (c) the traditional rankings are closest to true rankings for the high-low pairing, then the random within brackets pairing, then the simple random pairing, and lastly the high-high pairings.

It is actually worth looking at that last one:


Notice the clear pattern? In a high-high pairing, a ton of decent teams get screwed by getting several very hard opponents and therefore have terrible records -- these are the outliers that are very low on the y-axis (indicating true strength) but on the right side of the x-axis (indicating very poor records). Notice the one team in the far bottom right: pity the poor team that had true ranking of 19 but ended up with a 1-5 record and ranked 183rd by the traditional tiebreakers. Enragingly, a lot of weak teams somehow squeak by to great records. Notice the one team in the upper left: 148th in truth, but given several easy opponents, ending up 5-1 and ranked 21st by the traditional tiebreakers. Visually, you can see how unjust the whole high-high pairing is when coupled with using win/loss record as the primary criterion for ranking, as it is in the traditional method. The median points ranking does not suffer from the same problem; even the good team that gets several tough opponents and ends up 1-5 is not penalized in the rankings, as long as that team continued to earn high points in each one of its rounds.

For comparison, here is what the best correlation looked like:


To be sure, the correlation is far from perfect. But that's just about variability in the teams' in-round performances compared to their true strengths (that random deviation score I added). In other words, what you are looking at is just the off-days, surprises, and crappy judging that is unavoidable. It isn't really possible to do better than about 0.82 or 0.83 -- that's why the median rankings have about the same correlation, no matter what the pairing method.

On the other hand, the traditional rankings are very sensitive to the pairing method. Why? A team's record depends on both its true strength and the opponents it faces! In the high-high pairing method, many teams get unfairly hard or unfairly easy opponents. The method drives down the correlation between true strength and record by screwing some and blessing others. However, in the high-low pairing method, the assignment of opponents pushes up the correlation between true strength and record -- better teams face weaker opponents, so get a few more easy wins.

It can be a bit hard to interpret what these correlations means, so I also calculated the mean absolute deviations for each pairing method and ranking. For each team, I took its traditional rank and its true rank, found the difference, and took the absolute value. Then I averaged those to produce the mean absolute deviation (MAD). I also did the same thing for the median points rankings.


For example, for the random within bracket pairing method, the median points ranking had a MAD of 18.66. That means, on average, the median ranking was off by 18.66 places from the truth. The lower the MAD, the better.

The exact same patterns appear as in the correlations table. In general, the best we can hope for is to be within about 20 places of the truth. Given that debaters have off-rounds, and that our sample size is only six rounds, this isn't terrible: 20/200 is 10%. The true ranking is probably +/- decile from the median points ranking. Notice that the traditional rankings are sensitive to the pairing method in the exact same pattern. If one uses the traditional criteria for ranking, then the high-low pairing is best. So, was I wrong five years ago?

In both of the two tables I've given so far, the high-low pairing method plus traditional ranking was marginally superior to any pairing method plus median points ranking. But the problem is that it is not equally important to rank any team correctly. It is more important to get the top teams right. Enter the weighted rule. As I did for the MAD, I took each team's traditional rank, subtracted its true rank, and took the absolute value. But before I averaged, I divided by the team's true rank. Thus, getting a good team's results wrong by a lot was worth big negative points; getting a weak team's results by a lot was worth a few negative points. The results:


The pattern is almost the same as before, except that... high-low pairings and traditional rankings is worse (higher score) than random pairings with median rankings. This means that the high-low plus traditional combination made more mistakes ranking the best teams than the random plus median combination.

What are the take-aways?


1. If you are doing high-high power matching, STOP IT RIGHT NOW. Even one round of high-high power matching is harmful. You are screwing many teams over.


2. Consider using the median points ranking instead of the traditional ranking.

On T.R.P.C., it means putting the "drop two high - drop two low speaker points" as the first criterion for ranking. (For a three- or four-round tournament, the "drop high - drop low" option is equivalent to the median. For a five- or six-round tournament, the double-drop option is equivalent to the median. For a seven- or eight-round tournament, the triple-drop option is equivalent to the median.) You can make win/loss record the second or third criterion.

All the data from my experiment show that the median ranking is simply more accurate, no matter how you pair the tournament. The win/loss record is too variable.


3. Consider not doing high-low power matching either.

It is enormously time intensive to run a power-matched tournament. In some cases, power-matching adds 2-3 hours for a six-round tournament: 30-45 minutes after rounds 2, 3, 4, and 5 -- although one or two or those lag times might occur during a food break that had to happen anyway. But 2-3 hours might be used in other ways... say, to squeeze in an extra round. Another round would, in fact, yield more data and would improve the accuracy of the results far more than stopping frequently to power match. And, as the experiment data show, high-low pairings do not improve the accuracy any more than simply switching over to median rankings. (Furthermore, I suspect that high-low pairings plus traditional rankings' accuracy peaks at around five to six preliminary rounds; my suspicion is that for longer tournaments, the accuracy starts to go down again because the brackets start to get too small.)

High-low power matching does have something to argue for it: teams get to see more opponents of similar ability levels (to themselves). But there's a counterargument: random pairings enable teams to see a wide cross-section of opponents' skill levels, and better gauge where they fall on the spectrum. Getting your butt kicked can inspire striving, and besides, your tournament should have a novice and JV division for teams that are afraid of the best opponents in the top division.

However, if you feel like high-low power matching is something you want to preserve but you do want to speed up your tournament, then go to lag-powering. For example, round 3 would be power matched, but only off of the results of round 1. You can slip round 3 pairings under the doors while round 2 is wrapping up, cutting your turnaround time drastically. You probably won't push down your accuracy too much (see how well random within brackets plus traditional rankings compares) if you lag-power -- but especially not if you use median points rankings.


4. Do everything you can to help your judges give speaker points more consistently.

If speaker points are more accurate than records, that means we ought to put more weight on speaker points AND strive to make them seem less arbitrary. Brief training sessions at the beginning of the tournament for less experienced judges, clearly delineated rubrics for speaker points, or scoring grids for various attributes of speaking all help!


A changed perspective


I used to think that power matching started as the best way people had, when tabbing on notecards, to improve the accuracy of tournament results. Maybe that is why it got started, but as we can see, all it does is bring accuracy to parity with median rankings. Is there any other reason tabbers might have started to use power matching?

Then it dawned on me: power matching reduces the likelihood that two teams have met before will be randomly drawn against each other in later rounds. Team A might meet team B in round 1, win, then lose round 2. Team B might do the opposite and win round 2, giving the two opponents a possibility of being randomly drawn against each other for round 3 -- but overall, power matching makes it less likely than simply randomly assigning everyone in one big pool. When you're tabbing on notecards, it speeds things up considerably if this is a rare occurrence. Maybe power matching began, not with HH or HL but with the random within brackets pairing. In other words, how we pair double elimination tournaments (undefeateds and down-ones in two separate brackets, randomly assigned in each) might have gotten translated to all preliminary rounds at all tournaments.

Just a suspicion. It does make sense: high-low (or the awful high-high) brackets are hard to do on notecards, but random within brackets is easy to do.

Tuesday, July 3, 2012

Probability of upsets

A team has an average strength or skill level, which is how well we expect it to debate in a typical round. This is the same as the team's tournament-long average strength (teams probably improve during the course of the entire season). But a team's strength is also variable: in any given round, it might debate better or worse than its average. This variability should follow a normal distribution. When two teams debate, either might debate above or below its average. How to model this?


The horizontal axis shows possible performances of team 1, based on a normal distribution centered at 0 (indicating an exactly average performance for team 1 based on its average strength). The vertical axis shows possible performances of team 2, again a normal distribution around 0, the average-strength performance.

Let's say that team 1 is significantly stronger than team 2. In order for team 2 to win, it must have a much better than average performance -- and team 1 would have to have a much worse than average performance. In other words, only some of the possible results in quadrant 2 would result in a team 2 win, like so:


The red cases highlight the upsets. Rare indeed, because team 1 must underperform and team 2 must overperform. As an alternative, consider the scenario that team 1 and team 2 are evenly matched. In this world, team 2 wins about 50% of the time:


Mathematically, it is simple to model this with a logistic function. If difference = team 1 strength - team 2 strength, then the formula for the probability of team 1 winning is


where k depends on the units in which strength is measured and just how variable the teams' performances are. The value of k is an empirical research question that could change from season to season. The logistic function looks like this:


As the difference gets larger, team 1 is stronger and more likely to win, approaching 100%. As the difference turns negative, team 1 is weaker and less likely to win, approaching 0%. And at a difference of 0, the teams are even, and the odds are 50-50.

I analyzed the 2010-2011 season for open/varsity policy debate for CEDA/NDT data. I looked at each team's strength, using the easy-to-understand measure of weighted wins, expressed as an expected win percentage for a season (so, 62% means that a team is expected to win 62% of its rounds in an entire season, adjusted slightly from its actual win percentage by schedule strength). Then I analyzed all the rounds that happened, based on the difference in the two teams' strengths, as either wins (for the higher rated team) or upsets (for the lower rated team).

I found that about 20% of rounds were upsets. This is close to football's 25% or so. But of course, most of the upsets occur when the teams are fairly close in rating. Here are the results:


So, for example, when the difference in the ratings was greater than 0.5 but less than 0.55, the higher rated team won 97.3% of the time. This is obviously a significant difference in the teams' strengths: a team rated at 82% weighted wins versus a team weighted at 30% weighted wins! It is hardly surprising that this is such a lock. At the other extreme, when the difference in the ratings is greater than 0.1 but less than 0.15, the higher rated team only wins about 59% of the time. These are close rounds, nearly toss-ups. A difference of 0.2 seems to be the tipping point: above this, there are few upsets.

Here is the same data in graph form:


A line of best fit is modeled. Using the formula above, my best guess is that k is about 6.5.

Thursday, December 31, 2009

A new measure of team strength: weighted wins

I've been thinking about and working for a while on a more accurate method of estimating a team's strength. Bear in mind, I'm not talking about the method for generating final preliminary rankings. The final ranking method is unlikely to ever change, which is not really a bad thing. There's a reason we're all rightfully attached to it: wins and total speaker points may not be the most accurate way to assess a team's strength, but it does seem the most just: those are the wins and points a team earned. So, I'm interested in how to estimate a team's strength only in order to make better power matches during the prelims, not to decide which teams break or don't break. As a second caveat, let me state that there's no way to say objectively that rankings are "correct"; a good ranking method is a good estimate of team strength, which varies anyway from round to round. All you can do is look at whether a ranking method yields some common sense results.

With those caveats stated, here's the first problem with win/loss record as a measure of team strength: no team debates a representative sample of the teams at the tournament. Every team debates six opponents out of n teams at the tournament. Except for round robins and very small tournaments, the proportion isn't very large. At a normal-sized tournament, it might be under 10%. Recognizing this, tournaments do not use randomly selected opponents. Brackets select a subset of opponents for teams to debate, and the results are more informative than if opponent selection is random. While there are many pathways through the tournament (e.g., WWWLLL versus WLWLWL), usually the key is what caliber of opponent a team beats and by what caliber of opponent a team is beaten. For example, a team who beats a 2-4 and loses to a 4-2 will, because of the way the brackets work, most often end up with a 3-3 record, revealing that this team is probably in the middle third of the tournament. Of course, sometimes the brackets don't work perfectly in this way; for example, a team that beats a 4-2 might end up with a 3-3 record. Clearly, in this case, the overall win/loss record is not an accurate reflection of one or both teams' strength.

It is possible to create many different, more accurate measures of team strength that account for schedule strength using complicated formulas. But I think a measure needs to be relatively easy to understand and transparent, if it's to be adopted. The measure I developed (from a suggestion from Steve Gray) is weighted wins/weighted losses. Let's say team A beats teams B, C, and D, and loses to team G: 3 wins, 1 loss. But what if team B was a 3-1 team, team C was a 2-2 team, team D was an 0-4 team, and team G was a 3-1 team? The win against B ought to count for more than the win against D. The weighted measures would give team A exactly 8 "wins" (3 wins + 3 + 2 + 0) and 2 "losses" (1 loss + 1): it gains an extra "win" for every win of each opponent it beats (B, +3; C + 2; and D, + 0) and incurs an extra "loss" for every loss of each opponent that defeats it (G, - 1). As a first step, this already makes an enormous difference in assessing a team's true strength. Teams that defeat good opponents have more weighted wins than teams that defeat mediocre opponents, even if they have the same win/loss record. (Of course, it's possible that a team that has only been paired against mediocre opponents is actually very good -- which will get sorted out through power-matching!)

The method becomes enormously powerful if it re-iterates: for example, team A is now treated as having 8 wins and 2 losses, just as all its opponents are shown with their weighted wins and losses. Let's say B has a weighted record of 5-2; C, 4-4; D, 0-7; and G, 6-2. In the second iteration, team A will have 12 re-weighted wins (3 wins + 5 + 4 + 0) and 3 losses (1 loss + 2). The process can continue to be re-iterated until a reasonable stopping point, say, a team has as many or more weighted wins than there are teams at the tournament (or as many losses)! Based on this rule, the process will re-iterate about log(n) times for an n-team tournament.

For simplicity's sake, I would turn the weighted wins and weighted losses into one statistic:

where r is the number of rounds at the tournament. The first term will create something like a handicapped win percentage.

I ran this method on a small four-round tournament, which took only three iterations. You can see the results here:


In only one case, highlighted yellow, did the method rank a team above someone that it beat. I highlighted in green three teams that dramatically moved up under this ranking versus a traditional wins-speaker points method and in pink four teams that dramatically moved down. Based on their schedule strengths, these all seem pretty defensible to me.

Some might point out that there's an inherent difficulty created by "upsets," where good teams happen to get knocked off by bad teams. How much is the good team "punished," or pushed down in the rankings, by that loss? I thought a different kind of data set, where the teams play a much more representative sample, would show the basic sanity of the weighted wins approach. I used it to rank the 2009 N.F.L. regular season because there are so many "upsets" in football (about 25%! -- much higher than debate tournaments, where there are about 5%). You can see how well the method I described handles the unusual losses:


As you can see, it does a reasonable job, despite lots of unusual losses. [Note: weighted wins is scaled differently here for an unimportant technical reason (I was confounded by how to deal with the multiple games teams play against the same opponents), but the method is the same.] Here is a second post on weighted wins.