Showing posts with label algorithms. Show all posts
Showing posts with label algorithms. Show all posts

Thursday, October 9, 2008

More info on APR

I want to expand and emphasize a statement I made in my earlier post on how APR works. I wrote:

a simple sort of teams from largest to smallest power index is used to generate a power rankings

This is important. Each team's position in the power ranking is based only on the relative values of their respective power index values generated by APR. In particular:

  • No consideration is given to the ranking a team had in the previous week. It is possible (and in fact, not unusual) for teams to fall even though they won, or rise even though they lost.

  • No consideration is given to win-loss records, and it's very possible for, say, a 4-2 team to be ranked below a 1-5 team.

An example

I know it's possible for a 4-2 team to be ranked below a 1-5 team, because it happened during the 2007 season.

After the week 6 games of the 2007 season, the Baltimore Ravens were 4-2. The ESPN power rankings had them in the #7 spot. But APR ranked them #25. Below the 1-5 Falcons, the 1-4 Bengals, and the 1-4 Saints. In all, APR had 12 teams with .500 or worse records ranked above the Ravens that week.

Why should a team with a winning record be ranked so low? As usual with APR, the Ravens had played a some weak teams, including losses to the Browns and Bengals, and a 2-point victory over a very weak 49ers squad.

APR is meant to be a predictive measure of power (e.g. how well teams perform the following week). In this case, the low ranking was clearly justified, as the Ravens lost their next 9 games (including giving the Dolphins their one win of the season). The only game the Ravens won after week 6 was in week 17, against a Steelers team playing Charlie Batch at QB.

Conclusion:

This is exactly the kind of thing APR is meant to show: the power of each team as it continues to play games in the season. If you read this description of the SRS ranking algorithm, the discussion makes a distinction between predictive systems (which team is more likely to win their next game) and retrodictive systems (which team accomplished more in the past). APR (like SRS) is a predictive system, and should be approached as such.

Wednesday, October 8, 2008

Algorithms: How APR works, and what it's trying to do

For those of you coming here from footballoutsiders.com, let me say right away that APR is not supposed to be a replacement or alternative to DVOA, DPAR, or any of their other metrics.

The APR algorithm:

  1. Initially, every team is assigned a power index of 1.0

  2. For each team, a power index detailed below) is assigned to it for each game played.

  3. For each team, the game power values for that team are then combined together using a weighted average (old games weighted less than new games) to give a new power index for the team.

  4. Steps 2 and 3 are repeated a number of times to create a feedback loop.

The power for a played game is computed according to the following formula:

gamePower = resultPower × marginPower × opponentPower
Where:
  • resultPower is a constant based on whether the team won or lost, and whether they were the visiting or home team, with the following constraints:

    road win > home win > tie > road loss > home loss
  • marginPower is a constant based on the margin of victory (or loss). Right now, margin power is assigned in the following ranges:

    • won by 14 or more
    • won by 7 or more
    • won (or lost) by 6 or fewer points
    • lost by 7 or more
    • lost by 14 or more
  • opponentPower is the power index of the opponent team.

Using the power index values

Once power index values have been generated, a simple sort of teams from largest to smallest power index is used to generate a power rankings. Similarly, picks for the following week's games are made by chosing the team with the largest power index value for games played to that point.

Evaluating design choices

The efficacy of the algorithm (as well as choices for the various constants) is judged strictly on how well it does picking games on the historical data set of NFL games played. I have made attempts to give more power for wider margins of victory (or loss); such attempts have lead to fewer games picked correctly and were discarded. This means teams can generate a only limited amount of power playing very weak teams, even with blow-out wins.

The design of the APR algorithm has also lead to a related phenomenon I call "power by association" (or PBA, if you like TLAs): When a weak team plays a strong team close (especially with the weak team on the road), the weak team will increase in power, even if they lose (and the strong team will decrease in power, even if they win).

The APR ranking system is by no means perfect. It doesn't take into account injuries, break-out players, fluke wins, how different teams match up, the strength of individual units within a team, or the strength of teams in different game situations. At the end of the season, it doesn't take into account that some teams will clinch their post-season fate early, and elect to rest key starters for one or more games.

Still, it does (to my way of thinking, anyway) a remarkably good job at picking games, and often reveals over- and under-rated teams before ESPN or other subjective-based power ranking systems notices.

Update: and, if you've made it this far, be sure to read this post, which expands on the consequences of APR's design.

Saturday, March 29, 2008

Historical performance of Isaacson-Tarbell

In the column he describes it in, TMQ claimed Isaacson-Tarbell finished the 2007 season 183-84, which is a pretty impressive 68.5% accuracy. However, in my analysis, I was only able to account for 176 correct picks (65.9%). I have triple-checked my results, so I'm going to blog them.

Here's a graph of how the Isaacson-Tarbell Predictor performs on historical data, from 1960 to 2007 (along with the same "Home Team" data as posted below).

Click on the image to open a full-sized version

Note that ITP seems to have its best years in the 60's and early 70's; about the same time the "Home Team" strategy gives its worst results. Then, during the late 70's, they close together, giving what looks to be much more correlated results.

ITP looks to be an interesting algorithm, but not quite the home-run TMQ makes it out to be (you can see why he didn't mention its performance on the 2006 season).

A few notes on methodology

  • There are a number of tie games in the historical data, particularly before the advent of the overtime period for regular season games in 1974. Tie games are counted as a "push" for pick algorithms. As with the NFL, for the purposes of computing a winning percentage, push results are counted as half wins, half losses.

  • TMQ's description of Isaacson-Tarbell reads
    Best Record Wins; If Records Equal, Home Team Wins.
    Instead of "best record", I used the easier-to-deal-with "best winning percentage". The only time this will matter is when two undefeated (or two winless) teams match up, and one has had their bye, and the other hasn't. Coincidentally, this situation came up in the 2007 season, when the 8-0 Patriots met the 7-0 Colts. They both have the same winning percentage (100%), but the Patriots have the better record (one more win).

Thursday, March 27, 2008

Historical results of picking the home team

Here's a graph of the home team winning percentage, for NFL games (and AFL, prior to 1970), from 1960 to 2007.

Click on the image to open a full-sized version

A few things stand out:

  • From 1960 to 1976, only 4 years (1963, 1969, 1970, and 1973) are above the mean. There were 4 really bad years (1962, 1965, 1968, 1972). There were only two relatively good years (1969 and 1973). From 1974-2007, the worst year was 2006 (53.9%).

  • Since 1986, the year-to-year difference has been in the range -5% < δ < +5%.

    . From 1960 to 1986, there were 11 swings of 5% or wider, including three (1968 to 1969: +13.3%, 1972 to 1973: +12.1%, and 1985 to 1986: -9.8%) of 9.8% or wider.

  • It is surprising to me that the home-team winning percentage varies so much, even in the modern era. For a 267 game season (256 regular season and 11 playoff games), a 10% swing represents nearly 27 games.

    The disadvantages of being a road team seem pretty static: they still have to spend time travelling, they still have to spend time away from home, away from familiar comforts, away from team facilities...

Wednesday, March 26, 2008

Some more game-picking algorithms

In his February 15, 2008 column, Gregg Easterbrook (aka the Tuesday Morning Quarterback) described a couple of objective ways to pick football games.

The first method is the model of simplicity: always pick the home team. This yielded a correct prediction 152 times out of 267 games (56.9%, including the playoffs). Some historical analysis shows this can be surprisingly good (63.9% correct in 1985) or surprisingly bad (47.5% NFL and AFL combined, in 1968). Still, at the very least, the "Pick the Home Team" method provides a base-line point of comparison for other techniques.

The second method Easterbrook discusses is a method I will call the Isaacson-Tarbell Predictor: always pick the team with the better record (in case of a tie, choose the home team). TMQ claims this yielded a correct prediction 183 times out of 267 games (68.5%, including the playoffs). This is pretty spectacular, considering (as TMQ points out), none of the 8 experts featured on ESPN's "Expert Picks" did as well (Jaws came closest, at 68.3%).

Some questions for further investigation: how good is ITB (the Isaacson-Tarbell Predictor)? Was 2007 a fluke season? Can some variant of SRS, or some other power index algorithm do better?

Stay tuned for further details.