Category Archives: Online Tools / Apps / Data Sources

R graph catalog

Here’s a nice catalog of graphs made with R, along with source code for each. Some of the images were broken or missing when I tried it, but hopefully they’ll get that fixed. (By they way, this is my personal experience with interactive “Shiny” apps so far – I love the idea and the look, but there always seems to be something wrong that needs to be fixed, and fixing it takes more time and requires more specialized training than just dealing with plain old code. At first, I thought it might be a productivity enhancer, but instead it’s a drag when your job is not to build cool-looking apps, but to produce useful data analysis results in a reasonable amount of time.)

open source street noise model

Here’s an open-source code for modeling street noise propagation. It’s written in R and open source database and GIS tools.

This paper describes the development of a model for assessing TRAffic Noise EXposure (TRANEX) in an open-source geographic information system. Instead of using proprietary software we developed our own model for two main reasons: 1) so that the treatment of source geometry, traffic information (flows/speeds/spatially varying diurnal traffic profiles) and receptors matched as closely as possible to that of the air pollution modelling being undertaken in the TRAFFIC project, and 2) to optimize model performance for practical reasons of needing to implement a noise model with detailed source geometry, over a large geographical area, to produce noise estimates at up to several million address locations, with limited computing resources. To evaluate TRANEX, noise estimates were compared with noise measurements made in the British cities of Leicester and Norwich. High correlation was seen between modelled and measured LAeq,1hr (Norwich: r = 0.85, p = .000; Leicester: r = 0.95, p = .000) with average model errors of 3.1 dB. TRANEX was used to estimate noise exposures (LAeq,1hr, LAeq,16hr, Lnight) for the resident population of London (2003–2010). Results suggest that 1.03 million (12%) people are exposed to daytime road traffic noise levels ≥ 65 dB(A) and 1.63 million (19%) people are exposed to night-time road traffic noise levels ≥ 55 dB(A). Differences in noise levels between 2010 and 2003 were on average relatively small: 0.25 dB (standard deviation: 0.89) and 0.26 dB (standard deviation: 0.87) for LAeq,16hr and Lnight.

 

automated aggregation of scientific literature

I am intrigued by this example from Stanford of computerized review and synthesis of scientific literature:

Over the last few years, we have built applications for both broad domains that read the Web and for specific domains like paleobiology. In collaboration with Shanan Peters (PaleobioDB), we built a system that reads documents with higher accuracy and from larger corpora than expert human volunteers. We find this very exciting as it demonstrates that trained systems may have the ability to change the way science is conducted.

In a number of research papers we demonstrated the power of DeepDive on NMR data and financial, oil, and gas documents. For example, we showed that DeepDive can understand tabular data. We are using DeepDive to support our own research, exploring how knowledge can be used to build the next generation of data processing systems.

Examples of DeepDive applications include:

  • PaleoDeepDive – A knowledge base for Paleobiologists
  • GeoDeepDive – Extracting dark data from geology journal articles
  • Wisci – Enriching Wikipedia with structured data

The complete code for these examples is available with DeepDive.

Let’s just say an organization is trying to be more innovative. First it needs to understand where its standard operating procedures are in relation to the leading edge. To do that, it needs to understand where the leading edge is. That means research, which can be very tedious, and time consuming. It means the organization is paying people to spend time reviewing large amounts of information, some or even most of which will not turn out to be useful. So a change in mindset is often necessary. But tools that could jump start the process and provide short cuts would be great.

This is my own developing theory of how an organization can become more innovative: First, figure out where the leading edge is. Second, figure out how far the various parts of your organization are from the leading edge. Third, figure out how you are going to bring a critical mass of your organization up to the leading edge – this is as much a human resource problem as an innovation problem. Fourth, then and only then, you are ready to try to advance the leading edge. I think a lot of organizations have a few people that do #1, but then they skip right to #4. Then that small group is way outside the leading edge while the bulk of the organization is nowhere near it. That’s not a recipe for success.

Scratch

Scratch” is another programming language supposedly aimed at children.

Scratch Overview from ScratchEd on Vimeo.

If you watch the TED talk in the first link, there is an analogy I like – just because you use technology created by others (web browsing, texting, etc.) doesn’t make you fully literate in that technology. It is akin to being able to read but not able to write.

online productivity and creativity apps

This article from Civicly lists useful online apps for planners – actually, I think they are useful for anybody whose job involves trying to solve problems with a little creative latitude. I especially like the free tools for infographics – it looks like you can pick a template and customize it for your data.

habitat fragmentation and connectivity

Did you ever wonder how to quantitatively analyze the quality, shape, and degree of connectivity of natural habitats? Well, there’s an open source app for that, called FRAGSTATS, and good documentation that describes the theory behind it. To summarize, it looks at area and edge, shape, core area, contrast, aggregation, and diversity. Here are just a few quotes describing some of the metrics.

“Core area is defined as the area within a patch beyond some specified depth-of-edge influence (i.e., edge distance) or buffer width.”

“Contrast refers to the magnitude of difference between adjacent patch types with respect to one or more ecological attributes at a given scale that are relevant to the organism or process under consideration.”

“Aggregation refers to the tendency of patch types to be spatially aggregated; that is,
to occur in large, aggregated or “contagious” distributions.”

“FRAGSTATS computes 3 diversity indices. These diversity measures are influenced by 2 components- richness and evenness. Richness refers to the number of patch types present; evenness refers to the distribution of area among different types.”

protected bike lanes

Continuing on my recent transportation theme, this article on Alternet has some really good statistics on protected bike lanes. I am convinced that biking (a.k.a. cycling) is just a more practical way to get around urban areas than cars – it gets more people from point A to point B with less infrastructure, less cost, less wasted space, and no pollution. Plus, it promotes a more healthful, active lifestyle and urban design that supports that.

But for all this to happen, we have to build cycling infrastructure that is truly safe, and the U.S. just hasn’t fully committed to that. There are signs of hope, however – here are some of the statistics I’m talking about:

  • 27% of all trips in the Netherlands are made on bicycles. The Dutch designs are not secret but are available here (although their manual costs 90 Euros and it is not clear to me whether an English version is available).
  • The “pioneering” American city in protected bike lanes is…Montreal with over 30 miles (I just remembered, Canada shares our North American continent). But New York City has caught up and surpassed them with 43 miles. Other cities are Chicago (23 miles), San Francisco (12 miles), Austin (9 miles), and D.C. (7 miles). (Here in my native Philadelphia, we have not built protected bike lanes but have closed some lanes to traffic and painted some new stripes on the streets that would allow us to eventually separate them. Philadelphia has a burgeoning cycling culture and I think eventually it will happen. We don’t like to do anything first, we always sit back and watch what New York is doing for a few years before we build up the courage to try something new.)
  • Studies are finding that bike infrastructure boosts retail sales – 49% for a street in New York, 24% in Portland – and 65% of merchants surveyed reporting positive effects in San Francisco. (I’m not surprised by this – there is less space wasted on car travel lanes and parking, less time wasted circling around looking for parking, less money spent on parking, more room for trees and fountains and sidewalk cafes – you have more people in a given space, yet less crowding, with more time and money on their hands and a nicer environment where they want to hang around.)
  • And…duh…protected bike lanes are safer for everyone, and add more capacity to move more people at much lower cost compared to new traffic lanes.

The article also links to this fantastic collection of articles and data on protected bike lanes from “peopleforbikes“.