Tag Archives: data science

Nate Silver’s Iowa Caucus Predictions

Political season is data science season! Here is some more on Nate Silver’s forecasting methods. If you are reading this in real time (Sunday January 31), by tomorrow night we will find out what actually happens. I will reproduce some graphics here – these are all from the FiveThirtyEight site, so please thank me for the free advertising and don’t send me to copyright jail.

For Clinton vs. Sanders, here is Nate’s average of polls as of today. He gives more recent polls greater weighting, and also adjusts somehow for bias shown in the same polls in the past.

Average of polls: Clinton 48.0% vs. Sanders 42.7%

Now, this is within the 4-6% “margin of error” reported by most polls. (I find this easier to find on the RealClearPolitics site, although curiously it lists margins of error for Democratic polls but not Republican ones. RealClearPolitics does a straight-up poll average without all the corrections that today is Clinton 47.3% vs. Sanders 44%. So all the corrections don’t make an enormous difference.) I can’t easily and quickly find information on whether the “margin of error” is a standard error or a confidence interval or what, but generally when the polls are within the margin of error the media tends to report it as a “statistical tie” or dead heat. And that is exactly what they are saying in this case.

Nate Silver does a set of simulations – it sounds very complicated, but in essence I assume he takes his adjusted poll average for each candidate, some measure of spread like standard error, then runs a whole bunch of simulations. Which leads to results like this:

Clinton-Sanders Simulation

http://projects.fivethirtyeight.com/election-2016/primary-forecast/iowa-democratic/

Based on this, Nate Silver gives Clinton an 80% chance of winning Iowa and Sanders only a 20% chance.

So what’s interesting is that you have the average of polls (48-43 or 47-44 depending on source), which everyone says is a statistical tie. You have Silver’s predicted result (50-43) based on a large number of simulations, and then you have the resulting odds considering both the predicted result and the spread in the predictions (80-20). In other words, the computer is generating random numbers and 80% of simulations end up favoring Clinton. Of course in real life the dice get rolled only once, but these odds seem pretty good for Clinton.

Meanwhile, the Trump-Cruz contest is similarly close in the polls (30-25 in favor of Trump), but the predicted result (26-25 in favor of Trump) and odds (48-41 in favor of Trump) are much closer. From a quick glance, this appears to be because the spreads are much wider. I don’t know why that would be the case – presence of more viable candidates on the Republican side? Or maybe there is just more variability in the polls and nobody actually knows why.

Republican Iowa Caucus simulation

http://projects.fivethirtyeight.com/election-2016/primary-forecast/iowa-republican/

 

 

what is a blizzard?

According to Five Thirty Eight,

Three factors are required for a storm to be classified as a blizzard at a particular place, besides falling or blowing snow:

1. Sustained winds or frequent wind gusts of 35 mph or greater

2. Visibility under a quarter-mile

3. These conditions must persist for three hours.

This definition is the same whether you’ve got 1 inch or 40 on the ground.

what is a p-value?

Five Thirty Eight has a video of statisticians trying to explain what a p-value is. Well, what’s disturbing to me is that they won’t really try. Then again, the maker of the video very well may have cherry picked the most entertaining answers. I can’t reproduce the research so I have no way of knowing.

Here’s another article slamming the humble p-value. It’s true, there will always be some false positives if the data set is large enough. As an engineer, I try to use statistics to back up (or not) a tentative conclusion I have reached based on my understanding of a system. I will question a statistically significant result using my understanding of a system. That way both statistics and system thinking can reinforce and make each other stronger, rather than our relying exclusively on one or the other. Another way to think about this is that as data sets grow and our traditional engineering system analysis methods are just taking too long to apply, we can use statistics to weed out a lot of the data that is clearly just noise, and then focus our brains on a reduced data set that we are pretty sure contains the signal, although we know there are some false positives in there. So i say relax, use statistics, but don’t expect statistics to be a substitute for thinking. Thinking still works.

 

how do you value data?

This article lists six ways a company or organization can try to value its data:

  1. Intrinsic value of information. The model quantifies data quality by breaking it into characteristics such as accuracy, accessibility and completeness.
  2. Business value of information. This model measures data characteristics in relation to one or more business processes. Accuracy and completeness, for example, are evaluated, as is timeliness…
  3. Performance value of information…measures the data’s impact on one or more key performance indicators (KPIs) over time
  4. Cost value of information. This model measures the cost of “acquiring or replacing lost information.”
  5. Economic value of information. This model measures how an information asset contributes to the revenue of an organization.
  6. Market value of information. This model measures revenue generated by“selling, renting or bartering” corporate data

Another article says that algorithms are becoming less valuable as data becomes more valuable.

Google is not risking much by putting its algorithms out there.

That’s because the real secret sauce that differentiates Google from everybody else in the world isn’t the algorithms—it’s the data, and in particular, the training data needed to get the algorithms performing at a high level.

“A company’s intellectual property and its competitive advantages are moving from their proprietary technology and algorithms to their proprietary data,” Biewald says. “As data becomes a more and more critical asset and algorithms less and less important, expect lots of companies to open source more and more of their algorithms.”

 

Edward Tufte

Here’s a fun interview with Edward Tufte, insult comic and author of The Visual Display of Quantitative Information. Here are a couple of his snappy retorts:

…highly produced visualizations look like marketing, movie trailers, and video games and so have little inherent credibility for already skeptical viewers, who have learned by their bruising experiences in the marketplace about the discrepancy between ads and reality (think phone companies)…

…overload, clutter, and confusion are not attributes of information, they are failures of design. So if something is cluttered, fix your design, don’t throw out information. If something is confusing, don’t blame your victim — the audience — instead, fix the design. And if the numbers are boring, get better numbers. Chartoons can’t add interest, which is a content property. Chartoons are disinformation design, designed to distract rather than inform. Thus they reduce the credibility of your presentation. To distract, hire a magician instead of a chartoonist, for magicians are honest liars…

Sensibly-designed tables usually outperform graphics for data sets under 100 numbers. The average numbers of numbers in a sports or weather or financial table is 120 numbers (which hundreds of million people read daily); the average number of numbers in a PowerPoint table is 12 (which no one can make sense of because the ability to make smart multiple comparisons is lost). Few commercial artists can count and many merely put lipstick on a tiny pig. They have done enormous harm to data reasoning, thankfully partially compensated for by data in sports and weather reports. The metaphor for most data reporting should be the tables on ESPN.com. Why can’t our corporate reports be as smart as the sports and weather reports, or have we suddenly gotten stupid just because we’ve come to work?

It’s a very interesting point, actually, that people are willing to look at very complex data on sports sites, really study it and think about it, and do that voluntarily, considering it fun rather than boring, hard work. It’s child-like in a way – I mean in a positive sense, that for children the world is fresh and new and learning is fun. What is the secret of not shutting down this ability in adults. I think it’s context.

more on automated data synthesis

Here’s another article from Environmental Modeling and Software about automated synthesis of scattered research results:

We describe software to facilitate systematic reviews in environmental science. Eco Evidence allows reviewers to draw strong conclusions from a collection of individually-weak studies. It consists of two components. An online database stores and shares the atomized findings of previously-published research. A desktop analysis tool synthesizes this evidence to test cause–effect hypotheses. The software produces a standardized report, maximizing transparency and repeatability. We illustrate evidence extraction and synthesis. Environmental research is hampered by the complexity of natural environments, and difficulty with performing experiments in such systems. Under these constraints, systematic syntheses of the rapidly-expanding literature can advance ecological understanding, inform environmental management, and identify knowledge gaps and priorities for future research. Eco Evidence, and in particular its online re-usable bank of evidence, reduces the workload involved in systematic reviews. This is the first systematic review software for environmental science, and opens the way for increased uptake of this powerful approach.

automated aggregation of scientific literature

I am intrigued by this example from Stanford of computerized review and synthesis of scientific literature:

Over the last few years, we have built applications for both broad domains that read the Web and for specific domains like paleobiology. In collaboration with Shanan Peters (PaleobioDB), we built a system that reads documents with higher accuracy and from larger corpora than expert human volunteers. We find this very exciting as it demonstrates that trained systems may have the ability to change the way science is conducted.

In a number of research papers we demonstrated the power of DeepDive on NMR data and financial, oil, and gas documents. For example, we showed that DeepDive can understand tabular data. We are using DeepDive to support our own research, exploring how knowledge can be used to build the next generation of data processing systems.

Examples of DeepDive applications include:

  • PaleoDeepDive – A knowledge base for Paleobiologists
  • GeoDeepDive – Extracting dark data from geology journal articles
  • Wisci – Enriching Wikipedia with structured data

The complete code for these examples is available with DeepDive.

Let’s just say an organization is trying to be more innovative. First it needs to understand where its standard operating procedures are in relation to the leading edge. To do that, it needs to understand where the leading edge is. That means research, which can be very tedious, and time consuming. It means the organization is paying people to spend time reviewing large amounts of information, some or even most of which will not turn out to be useful. So a change in mindset is often necessary. But tools that could jump start the process and provide short cuts would be great.

This is my own developing theory of how an organization can become more innovative: First, figure out where the leading edge is. Second, figure out how far the various parts of your organization are from the leading edge. Third, figure out how you are going to bring a critical mass of your organization up to the leading edge – this is as much a human resource problem as an innovation problem. Fourth, then and only then, you are ready to try to advance the leading edge. I think a lot of organizations have a few people that do #1, but then they skip right to #4. Then that small group is way outside the leading edge while the bulk of the organization is nowhere near it. That’s not a recipe for success.

visualization

Solomon Messing has a pretty good article on data visualization and communicating scientific information, focusing on the ideas of Tufte and Cleveland. I like the idea that there is a science of what our brains can most easily process, and not just a need to create visual infotainment because we have lost our ability to concentrate on anything else. I’m not quite ready to give up on stacked bar charts in all cases.

When most people think about visualization, they think first of Edward Tufte.  Tufte emphasizes integrity to the data, showing relationships between phenomena, and above all else aesthetic minimalism.  I appreciate his ruthless crusade against chart junk and pie charts (nice quote from Data without Borders). We share an affinity for multipanel plotting approaches, which he calls “small multiples,” (thanks to Rebecca Weiss for pointing this out) though I think people give Tufte too much credit for their invention—both juiceanalytics and infovis-wiki write that Cleveland introduced the concept/principle. However, both Cleveland and Tufte published books in 1983 discussing the use of multipanel displays; David Smith over at Revolutions writes that “the “small-multiples” principle of data visualization [was] pioneered by Cleveland and popularized in Tufte’s first book”; and the earliest reference to a work containing multipanel displays I could find was published *long* before Tufte’s 1983 work–Seder, Leonard (1950), “Diagnosis with Diagrams—Part I”, Industrial Quality Control (New York, New York: American Society for Quality Control) 7 (1): 11–19.

I’m less sure about Tufte’s advice to always show axes starting at zero, which can make comparison between two groups difficult, and to “show causality,” which can end up misleading your readers.  Of course, the visualizations on display in the glossy pages of Tufte’s books are beautiful–they belong  in a museum.  But while his books are full of general advice that we should all keep in mind when creating plots, he does not put forth a theory of what works and what doesn’t when trying to visualize data.

Cleveland (with Robert McGill) develops such a theory and subjects it to rigorous scientific testing. In my last post I linked to one of Cleveland’s studies showing that dots (or bars) aligned on the same scale are indeed the best visualization to convey a series of numerical estimates.  In this work, Cleveland examined how accurately our visual system can process visual elements or “perceptual units” representing underlying data.  These elements include markers aligned on the same scale (e.g., dot plots, scatterplots, ordinary bar charts), the length of lines that are not aligned on the same scale (e.g., stacked bar plots), area (pie charts and mosaic plots), angles (also pie charts), shading/color, volume, curvature, and direction.

I’m slowly getting on board. I’ve given up pie charts in most cases. I’m not ready to give up stacked bar charts in all cases – I think they serve a purpose. Microscopic multi-panel charts still make my head spin sometimes, although if they were interactive and I could click on one panel to blow it up, that would be cool. There is one thing I am sure he is right about though, which is that the first step to serious analysis and visualization is to leave Excel behind.