Showing posts with label weighted wins. Show all posts
Showing posts with label weighted wins. Show all posts

Monday, December 29, 2014

N.F.L rankings

I didn't get a chance last year to do my N.F.L. rankings, but here they are far this year. I made a slight tweak and used a logistic function to rate each win. A small point differential resulted in a half-win to each time. A large point differential resulted in one team getting much closer to a win (nearly 1) and the other getting a loss (closer to 0). However, with a logistic function, the differential does not result in a linear change. Basically, I set it up so that anything beyond a two-possession game -- for example, one team wins by 16 points -- is above a 0.90 for the winning team. Running up the score further pushes the win closer to a 1 in an ever slower fashion.

The next step was to use these adjusted results to do a strength-of-schedule adjustment.

Here is my logistic function:

win=0 if differential≤3.511+e-(diff-3.5)/5

And here are my rankings:



The link to download it. No real surprises, except that the Ravens are ranked in my method ahead of the Bengals, Colts, and Steelers.

My wild-card round predictions: Ravens over Steelers (the only "surprise"), Colts over Bengals, Cowboys over Lions, Cardinals over Panthers.

---

Update, 1/5/2015: Got the Ravens, Colts, and Cowboys right. Got the Panthers wrong (although this should have been the one I was the most confident about). Score: 3/4.

Next predictions for the divisional playoffs, using the unrevised list above: Patriots over Ravens, Broncos over Colts, Packers over Cowboys, Seahawks over Panthers.

Update, 1/13/2015: Got the Patriots, Packers, and Seahawks right. Got the Colts wrong. Score: 3/4.

Next predictions: Seahawks over Packers, Patriots over Colts.

Update: 1/21/2015: Got both right. Score: 2/2.

Next predictions: Seahawks over Patriots in the Superbowl.

Update: Wrong Superbowl prediction. 0/1.

Total score: 8/11, or 73%.

Monday, January 7, 2013

Impactranks

I recently heard of Josh Clark's new project impactranks.com. It is basically a coaches' poll. I respect the intention, but I strongly disagree with the method. Does anyone think the B.C.S. does a good job? It is largely based on a coaches' poll. The problem is that coaches do not see enough other teams play (or debate) and end up reflecting the perceived reputation of programs. There are sounder methods to rank teams that reflect actual results: wins, losses, and points. I would very much recommend reading, Who's #1? The Science of Rating and Ranking, by Langville and Meyer, two math professors, for many of the methods. My own method for debate is the weighted win method, explained here and here, in which whom a team beats matters even more than the raw win percentage.

One objection that people might point out is that debaters have so few opponents during the year that there is not enough data to think a mathematical method is any better than a poll. However, I would point out that the same problem exists in college football and professional football, yet mathematical methods based on opponent strength do work reasonably well.

Another objection is that debate has a lot of upsets, because of bad judges, off rounds, or unusual one-time tricks. True, but so does football, and the mathematical methods still work just fine. (In fact, I was even able to calculate an upset rate for college debate -- about 20%.) The other thought that I have had is to use the data on results to evaluate judges at the same time as debaters. Judges who return consistently unusual results (giving wins to worse teams with a lot of regularity) would lower their rating, so a loss from such a judge would not penalize a team by much.

The only difficulty to using these methods is that high school debate records are not stored in a clear format that shows: the two opponents, the judge(s), and the results. But the college records are. So here is my challenge to any reader. Before the N.D.T., I am going to rank all the teams, based solely on the results from the year, and publish that ranking. We will see how many results my rankings correctly predict (given the upset rate of 20%, anything in that neighborhood or better would be excellent). If any reader wants to come up with their own ranking, we will compare the results. You have until March 27th!

Friday, April 13, 2012

Weighted wins 2

I've been interested in using weighted wins as a statistic for a while. The idea is that by considering the strength of an opponent, an algorithm could handicap the results: a win against a good opponent counts for more than a win against a middling opponent. Of course, the algorithm has to base the calculation of how good an opponent is on how well it did against its opponents, so the algorithm must be based on the whole set of results. Therefore, I looked at a debate tournament and an N.F.L. season. The process can continue infinitely, although usually, after a couple times, things settle down and the handicapping doesn't change much upon further iterations. This is a kind of Markov chain technique.

One assumption that this depends on is that a team has an invariant strength. Clearly, this is suspect for debate (differing strengths on the affirmative and negative sides) and the other N.F.L. (the defense and offense are, literally, two different teams). Is it possible to use the same idea but adapt it to recognize the split-strengths?

I input total offensive yards from each 2011 N.F.L. game. A lot of yards against a weak defense is good; a lot of yards against a strong offense is better. By using only this information -- offensive yards for each team vs. each opponent -- a few matrix operations yielded a weighted offensive yards statistic and a defensive strength score for each team's offense and defense, respectively:


For example, against the "average" defense, the Saints would have earned 460 yards. The Steelers defense would have cut that to 82%, to 377 yards. Or so say these statistics. They never actually played.

How well did these statistics work? Well, as predictions, terribly. But, compared to reputable sports statisticians, fairly well. This is a comparison of my rank versus football outsiders rank (before the playoffs began) [offense on left, defense on right]:


On offensive, the difference was on average 3 ranks. On defense, the difference was on average 5.3 ranks. So, the rankings I generated from nothing other than actual yards in regular season games compared moderately closely to the ranks based on a complex calculation they call the DVOA (Defense-adjusted Value Over Average).

The same idea would work for debate tournaments: affirmative strength and negative strength could be treated as two separate variables for each team.

Thursday, December 31, 2009

A new measure of team strength: weighted wins

I've been thinking about and working for a while on a more accurate method of estimating a team's strength. Bear in mind, I'm not talking about the method for generating final preliminary rankings. The final ranking method is unlikely to ever change, which is not really a bad thing. There's a reason we're all rightfully attached to it: wins and total speaker points may not be the most accurate way to assess a team's strength, but it does seem the most just: those are the wins and points a team earned. So, I'm interested in how to estimate a team's strength only in order to make better power matches during the prelims, not to decide which teams break or don't break. As a second caveat, let me state that there's no way to say objectively that rankings are "correct"; a good ranking method is a good estimate of team strength, which varies anyway from round to round. All you can do is look at whether a ranking method yields some common sense results.

With those caveats stated, here's the first problem with win/loss record as a measure of team strength: no team debates a representative sample of the teams at the tournament. Every team debates six opponents out of n teams at the tournament. Except for round robins and very small tournaments, the proportion isn't very large. At a normal-sized tournament, it might be under 10%. Recognizing this, tournaments do not use randomly selected opponents. Brackets select a subset of opponents for teams to debate, and the results are more informative than if opponent selection is random. While there are many pathways through the tournament (e.g., WWWLLL versus WLWLWL), usually the key is what caliber of opponent a team beats and by what caliber of opponent a team is beaten. For example, a team who beats a 2-4 and loses to a 4-2 will, because of the way the brackets work, most often end up with a 3-3 record, revealing that this team is probably in the middle third of the tournament. Of course, sometimes the brackets don't work perfectly in this way; for example, a team that beats a 4-2 might end up with a 3-3 record. Clearly, in this case, the overall win/loss record is not an accurate reflection of one or both teams' strength.

It is possible to create many different, more accurate measures of team strength that account for schedule strength using complicated formulas. But I think a measure needs to be relatively easy to understand and transparent, if it's to be adopted. The measure I developed (from a suggestion from Steve Gray) is weighted wins/weighted losses. Let's say team A beats teams B, C, and D, and loses to team G: 3 wins, 1 loss. But what if team B was a 3-1 team, team C was a 2-2 team, team D was an 0-4 team, and team G was a 3-1 team? The win against B ought to count for more than the win against D. The weighted measures would give team A exactly 8 "wins" (3 wins + 3 + 2 + 0) and 2 "losses" (1 loss + 1): it gains an extra "win" for every win of each opponent it beats (B, +3; C + 2; and D, + 0) and incurs an extra "loss" for every loss of each opponent that defeats it (G, - 1). As a first step, this already makes an enormous difference in assessing a team's true strength. Teams that defeat good opponents have more weighted wins than teams that defeat mediocre opponents, even if they have the same win/loss record. (Of course, it's possible that a team that has only been paired against mediocre opponents is actually very good -- which will get sorted out through power-matching!)

The method becomes enormously powerful if it re-iterates: for example, team A is now treated as having 8 wins and 2 losses, just as all its opponents are shown with their weighted wins and losses. Let's say B has a weighted record of 5-2; C, 4-4; D, 0-7; and G, 6-2. In the second iteration, team A will have 12 re-weighted wins (3 wins + 5 + 4 + 0) and 3 losses (1 loss + 2). The process can continue to be re-iterated until a reasonable stopping point, say, a team has as many or more weighted wins than there are teams at the tournament (or as many losses)! Based on this rule, the process will re-iterate about log(n) times for an n-team tournament.

For simplicity's sake, I would turn the weighted wins and weighted losses into one statistic:

where r is the number of rounds at the tournament. The first term will create something like a handicapped win percentage.

I ran this method on a small four-round tournament, which took only three iterations. You can see the results here:


In only one case, highlighted yellow, did the method rank a team above someone that it beat. I highlighted in green three teams that dramatically moved up under this ranking versus a traditional wins-speaker points method and in pink four teams that dramatically moved down. Based on their schedule strengths, these all seem pretty defensible to me.

Some might point out that there's an inherent difficulty created by "upsets," where good teams happen to get knocked off by bad teams. How much is the good team "punished," or pushed down in the rankings, by that loss? I thought a different kind of data set, where the teams play a much more representative sample, would show the basic sanity of the weighted wins approach. I used it to rank the 2009 N.F.L. regular season because there are so many "upsets" in football (about 25%! -- much higher than debate tournaments, where there are about 5%). You can see how well the method I described handles the unusual losses:


As you can see, it does a reasonable job, despite lots of unusual losses. [Note: weighted wins is scaled differently here for an unimportant technical reason (I was confounded by how to deal with the multiple games teams play against the same opponents), but the method is the same.] Here is a second post on weighted wins.

Tuesday, December 15, 2009

Pascal's triangle (modified)

I started thinking about this problem in a specific debate context (for a specific application), before realizing that it isn't so useful after all. Still, the math is quite interesting.

Here's the original question that got me thinking: Is there a way to assign weighted wins -- round 1 wins count for so much, round 2 wins count for something different, and so on -- such that a tournament can produce a ranking that is consistent with the actual results, such that if team A beats team B, and they finish with the same win-loss record, A must be ranked above B? (Thanks to Steve Gray for coming up with the original idea to use weighted wins this way.) The point values for each round must be decreasing for this to work. If team A beats team B in round 1 for 1 point, then all subsequent point values must be smaller, or else B would have the chance to tie (with an equal number of wins) or surpass A (with an equal number of wins -- of course, B should surpass A if it wins more total rounds than A). I tested out a specific set of weighted win points that meets both of these conditions: {1, 0.9, 0.89, 0.889, 0.8889, 0.88889}. I'll write more on why this particular set works later in this post. The results of the test follows:


Pick any point where you might compare two teams, and you can see that you can always tell who got there by winning the earlier round(s). As you can see from the bottom row, it is possible to distinguish every one of the 64 possible pathways (2^6) that a team could take through the tournament (e.g., WLLWWL), and the results line up nicely in order -- that is, an earlier win always ranks a team higher (e.g., WWLWLL is ranked higher than WLLWWL). The earlier win is ranked higher, because that way, there is no possibility a team is ranked lower than someone they beat. NB: This only works if everyone debates within brackets. It fails to hold as a true assumption for pull-up rounds, so this system would fail utterly for round robins.

However, it is not the pull-up problem that scuppers this system (I think it's somewhat solvable for one-win pull-up rounds with half-points). The problem is that this system produces a ranking that is consistent with the actual win-loss results; it does not necessarily produce the most desirable ranking. True, teams are never ranked below someone they beat, but this system overdoes it: teams are ranked by the order in which they won rounds. What about an excellent 4-2 team with great speaker points who loses the first two random rounds to the top 4-2 teams? To the good, this excellent 4-2 team will be ranked below the top two 4-2 teams (because LLWWWW is worth fewer points than, for example, WLWWWL and LWWWLW); to the excess, this team will be ranked below EVERY 4-2 team who happened to win either of the first two rounds. It's throwing the baby out with the bath water. Furthermore, it makes no sense to use this kind of weighted wins as a third or fourth tie-breaker. If speaker points are the first tie-breaker, then it becomes possible to rank a team below someone they have beaten. It's an all-or-nothing solution, and no one would think that the rankings it produces are good.

The math behind this problem turns out to be more interesting than the originally intended application. Some patterns emerge in any set of numbers that work for weighted wins. In the set of numbers {a, b, c, d, e, }, the terms must be decreasing, so a > b > c > d > e > f  is the first condition. However, not just any set of decreasing terms will work. For example, if the weighted win points were {1, 0.9, 0.8, 0.7, 0.6, 0.5}, the system breaks down: for two 2-0 teams, a round 1 and round 6 win totals 1.5 points, but it should beat a round 2 and round 3 win, which totals 1.7 points. Therefore, the second, trickier condition is that the decreases must decrease, that is, while the slope is negative (negative first derivative) the set must have upward concavity (positive second derivative). Going to my original set of numbers {1, 0.9, 0.89, 0.889, 0.8889, 0.88889}, the terms are decreasing, meeting the first condition, and the decrease is decreasing {-0.1, -0.01, -0.001, -0.0001, -0.00001}, meeting the second condition.

Looking at the chart, every unmarked place is one where any set of numbers that meets only the first condition will be true. For example, on the line of round 3, as long as a > b > c, then a+b > a+c > b+c; it's analytically true. Every marked place is one where a set of numbers that meets only the first condition might fail; only sets meeting both conditions must be true. For example, on the line of round 4, a+d must be > b+c only if the second condition is also met. I made up a term for a set that meets the conditions of this problem: an "inequality ordered set." The definition would be: a set of numbers, such that for any two equally-sized subsets created without repetition of terms, the subset with the single largest term always has the largest sum. For example, a+f > b+c and a+e+f > b+c+d.

I wondered whether it is possible to create an infinitely long "inequality ordered set." Two possible ways to create such an infinite set: {1, (1/x), (1/x^2), ...} and {1, 1-(1/x), 1-(1/x)-(1/x^2), ...}. For the first kind, the smallest x that still works is about 2. For the second kind, the smallest x that still works is about 1.25. I'm working on writing a proof of these.