The Power of Alternative Data for Academic Research
A sales receipt, a web search, a satellite image, a job posting that expired two years ago; the seemingly mundane multitudes of data being created by daily activities hold unique and valuable insights for researchers to parse.
What is alternative data, and what makes it so useful for research? We're not talking about fringe science here, but the creative application of non-traditional sources of data to answering research questions. I spend a lot of time working with a catalog of datasets that captures a vast and varied array of human behavior and world phenomena, and I am constantly inspired by the innovative problem solving researchers employ to extract answers from new sources and diverse perspectives.
What makes data "alternative"?
Alternative data, with its many alternate names (organic [data], novel, found, digital trace, unstructured, observational, derived, etc.), is essentially any data that was generated for some purpose other than research – typically as a byproduct of some administrative, commercial, or technical process – and then utilized for research. Common sources of “alt data” include satellite imagery, mobile-phone records, usage data, administrative records, transaction records, job postings, social media and internet activity, sensor data, and much more.
This differs from traditional datasets that are designed, built, generated or measured, such as structured surveys, experimental data collection, or official government reporting. Alt data offers novel "proxy" measurements of real-time, real-life behavior and phenomena, giving researchers unique perspectives to new and old questions. While economics has paved the way in alt data usage, adopting large-scale administrative and private sector data in the study of market trends (Einav & Levin, 2014), many fields are also implementing new data streams into their research, from global health to climate change, public policy and more.
Take an age old question like: does spending time in nature improve mental health? There are many ways researchers go about answering this, the most traditional method typically being a survey or maybe even a random assignment experiment. But these researchers (Chandra Mouli et al., 2026) combined a traditional data source (CDC health measures) with real-life measures of behavior (cell phone mobility data) and neighborhood green space (NASA satellite imagery). Looking at nine US metro areas, they found that neighborhoods where residents actually visited nearby green space had lower rates of poor mental health; the presence of parks alone did not explain it. No single one of these three sources could have provided this nuanced insight; the combination of disparate data sources provides researchers with opportunities to explore old questions in creative new ways.
The science of measurement: how computational advances open the door for alternative data
As new data sources become available, how do researchers approach applying them to a question, or implementing them in a field? What methods and data standards are being developed to meet the growing adoption of alternative datasets?
We study what we can measure, and the history of science is largely a history of reinventing what counts as a measurement. Take tree rings, once only a mesmerizing pattern of nature until an astronomer needed a measurement of sunspot cycles (Douglass, 1909). Now the field of dendrochronology is the foundation of climate archives and dating research, from tracking drought cycles to archaeological studies of Viking settlements. The source was always there, it just took a creative scientist to see the answers it held.
But science is more than noticing, and the meticulous methods that make an alternative dataset mainstream can take decades of measurement, collaboration, and calibration. As with tree rings, Douglass was not the first person to notice their potential: Leonardo Da Vinci noted their link to moisture centuries earlier (da Vinci, c. 1500; published 1817). But Douglass was among the first to formalize a technical method for scientific study, replicable across regions and time (Douglass, 1919).
A modern example of this can be found in the journey of “nighttime lights” being established as a standard proxy for GDP. It took decades from noticing the potential data source (Croft, 1978), to establishing a correlation (Elvidge et al., 1997), through testing and evaluating (Chen & Nordhaus, 2011) and finally establishing a replicable statistical framework (Henderson et al., 2012). So why go through such effort, when you could just use the standard GDP? Because researchers were now able to gain insights into regions where national accounts and standard economic activity were not consistently available. Nighttime lights shed light on the economics of new places, people, and questions with improved resolution.
By 2016, a review paper discussed not just lights, but many kinds of satellite and remote sensing data that have been implemented in economic studies; they concluded with the sentiment that “the real excitement lies ahead” due to the technological and computational advances that are emerging along with new data sources (Donaldson & Storeygard, 2016, p.194). And today, I couldn’t agree more. Researchers understand the true work and the true reward buried in these troves of data. Whether it is validating and standardizing an obscure or seemingly unrelated measurement, or digitizing, cleaning, and modeling vast quantities of near unreadable and unmanageable data, researchers are expanding their horizons as they find and implement more alternative data sources. And the process is spreading across disciplines and speeding up as computational power improves.
What alternative data unlocks for researchers
Alternative data often allows researchers to gain insights into new populations, time frames, untapped markets, or under-represented communities. The sources are wide reaching, and so are their applications. For instance, Dewey currently hosts hundreds of datasets from dozens of providers, which tend to be grouped as data about people, places, companies/government, or finances (check out our Market Map of commercial data sources). And Dewey users across hundreds of universities are applying these datasets to thousands of different research questions (browse our publications database). As these researchers have discovered, alternative datasets have widely variable applications, perhaps because they were not originally built for a specific research question.
Measure what was never measured
Places and populations with no survey infrastructure now become legible, surfacing rural areas, small populations, and under-studied issues. The absence of these communities or issues in traditional data sources often has significant downstream consequences. Places that go uncounted go unstudied, unfunded, and unrepresented in the evidence that shapes policy. Many alternative data sources are able to capture these regions and communities. For example, applying neural networks and machine learning models to alternative data sources like phone usage patterns and satellite imagery allowed researchers to map wealth distribution in low- and middle-income countries with high resolution (Blumenstock et al., 2015; Jean et al., 2016; Chi et al., 2022). These researchers reach more people and places by combining a variety of data sources and computational methods.
Observe at population scale
Big data becomes useful not just because of the scale, but also for the ability of that wide scope to capture smaller portions of the whole that typically get overlooked. Rare events, small subgroups, and trace effects become studiable at scale.
Showcasing this are several Nature papers that investigate how social networks influence behavioral outcomes by utilizing Facebook friendship data as a measure of personal connections. In 2012, an experimental analysis found that the "I Voted" icon measurably increased real-world voter turnout among friends in the 61-million-person sample (Bond et al., 2012). And in 2022, Raj Chetty and colleagues studied 21 billion Facebook friendships across 72 million users to show that "economic connectedness" is among the strongest known predictors of upward socioeconomic mobility (Chetty et al., 2022). Effects like these are often only visible with scale, with the ability to ask structural questions about a whole network.
Watch continuously instead of annually
Frequency stops being a constraint with a constant stream of data, as many digital trace measures offer real-time measurements of real-life behavior and events. The COVID-19 pandemic accelerated the need for providing near-real-time measurement of human mobility, and getting this data into the hands of researchers became Dewey's own origin story. To date, hundreds of studies have used mobility data to study vastly varied topics, from mental health (Chandra Mouli et al., 2026) to criminology (Moritz et al., 2026) to economic resilience (Yabe et al., 2025).
Continuous measurements and fine time resolution are useful to understand cause and effect, event impacts, and nuances that annual reporting can't capture. Whether it’s contesting official government inflation statistics with daily pricing data (Cavallo & Rigobon, 2016) or quantifying the business boom around major sporting venues on game days (Abbiasov & Sedov, 2023), real-time data with granular frequency provided answers that official or aggregated data could not.
Reach backwards in time
Longitudinal studies are notoriously difficult to implement and maintain, partly because they require someone to keep consistent infrastructure, measurement and methodology for years. Studies can often run longer than a grant cycle, or sometimes longer than a career or lifetime (depending on the research question), which makes consistency even more challenging. Digital and archival traces, on the other hand, provide a record that has been accumulating regardless, unmaintained and unfunded, simply waiting for a question.
Researchers may spend their due diligence creating such archives from available sources, but technological advances are improving the preservation and processing capabilities of historical sources. For example, a Science paper (Michel et al., 2011) that constructed a corpus of digitized books (representing 4% of all books ever printed) used it to quantitatively measure cultural trends all the way from 1800-2000. Or something typically thrown away, like expired newspaper job ads, was compiled as a text corpus tracing the evolution of work in the US from 1950-2000 (Atalay et al., 2020). By applying modern quantitative methods to historical data sources, these researchers are reaching far back in time and tracking trends across centuries.
And alternative datasets also offer valuable insights into more recent history in the study of emerging trends or unanticipated events, like the unfolding impacts of AI. By providing a baseline and historical data unbiased by certain events or measurements, data sources like scraped job ads can measure how the release of ChatGPT influenced the labor market in different global economies (Khosravi & Liu, 2026 working paper). Much of the digital trace data and digital records accumulated every day can be utilized as a retrospective record of behavior, market trends, or event impacts.
Dewey’s goal: make alternative data more accessible to researchers
We are a small team here at Dewey, fueled to empower researchers with the tools they need to answer the important questions. We are excited about alternative data because we see first hand the impact and insights that become possible when the right data reaches the right researchers. My own background being a mix of academia and data science, I love how alt data provides opportunities to apply computational methods towards meaningful questions.
But alternative data still comes with many challenges; Dewey aims to support and solve what we can, and more importantly to listen and learn along with researchers as these data sources and tools grow. The nature of alternative data is inherently unpredictable: creative, yet volatile. As a data platform serving academic researchers, we need to be a nimble, tech-forward team that is able to adapt quickly as data constraints change and research priorities shift. While alternative data sources offer incredible potential for new perspectives, deeper coverage, and novel insights, issues like access and reproducibility are common barriers, and researchers raise real concerns about ethics, representativeness, and validity. And we are no strangers to the messy world of navigating data infrastructure, joinability, and maintenance. These issues are not new or unique to alternative data, but deserve equal due diligence and potentially new solutions as the data landscape evolves.
Dewey exists to ease these friction points in data access for researchers, from the origin of opening up Safegraph's mobility data to researchers during the pandemic, to licensing hundreds of alternative datasets to thousands of researchers today. We actively work to improve data workflows, reproducibility, and transparency, with things like dataset DOIs, robust documentation, collaborative platform integration features, and helpful resources. Got ideas for other ways to improve data workflows? We always want to hear them!
I am constantly inspired by the researchers I get to support and celebrate with Dewey. Yes, there is power in alternative data, and that power is in the researchers that wield it for societal growth, advances in knowledge, and the gritty work of making the uncounted count, the unseen seen, and the once unusable data, usable. We are seeing these approaches applied to questions that provide a more nuanced understanding of the whole spectrum, a clearer view of the people and issues typical sources miss, and add new perspectives to old questions as well as find new questions nobody has asked yet. There is so much “green space” for researchers to explore, and I look forward to watching this research grow, along with the methods, computing power, and sources that make it possible. As computational methods improve, the method of the moment is to create a moment of method. The excitement and caution with alternative data is to be creative in application and rigorous in implementation towards impactful science.
Happy researching!
Resources
- Browse the Dewey Catalog
- Seminars - Upcoming: How Alternative Data is Shaping Research Today
- Unique Data Attributes - read my article series exploring some of the alternative datasets on Dewey with unique research value
- Dewey blog
- The publications database
- 2026 Market Map of commercial data providers
- Meet the Dewey Team!
References
Abbiasov, T., & Sedov, D. (2023). Do local businesses benefit from sports facilities? The case of major league sports stadiums and arenas. Regional Science and Urban Economics, 98, 103853. https://doi.org/10.1016/j.regsciurbeco.2022.103853
Atalay, E., Phongthiengtham, P., Sotelo, S., & Tannenbaum, D. (2020). The Evolution of Work in the United States. American Economic Journal: Applied Economics, 12(2), 1–34. https://www.aeaweb.org/articles?id=10.1257/app.20190070
Blumenstock, J., Cadamuro, G., & On, R. (2015). Predicting poverty and wealth from mobile phone metadata. Science, 350(6264), 1073–1076. https://www.science.org/doi/10.1126/science.aac4420
Bond, R. M., Fariss, C. J., Jones, J. J., Kramer, A. D. I., Marlow, C., Settle, J. E., & Fowler, J. H. (2012). A 61-million-person experiment in social influence and political mobilization. Nature, 489, 295–298. https://www.nature.com/articles/nature11421
Cavallo, A., & Rigobon, R. (2016). The Billion Prices Project: Using Online Prices for Measurement and Research. Journal of Economic Perspectives, 30(2), 151–178. https://www.aeaweb.org/articles?id=10.1257/jep.30.2.151
Chandra Mouli, S., Zhang, T., Chen, Z., Rajagopalan, S., Maddock, J. E., Nasir, K., Dong, W., & Al-Kindi, S. (2026). Neighborhood greenspace visits and mental health: insights from mobility data across nine U.S. metropolitan areas. Frontiers in Public Health, 14, 1731243. https://doi.org/10.3389/fpubh.2026.1731243
Chen, X., & Nordhaus, W. D. (2011). Using luminosity data as a proxy for economic statistics. PNAS, 108(21), 8589–8594. https://www.pnas.org/doi/10.1073/pnas.1017031108
Chetty, R., Jackson, M. O., Kuchler, T., Stroebel, J., et al. (2022). Social capital I: measurement and associations with economic mobility. Nature, 608, 108–121. https://www.nature.com/articles/s41586-022-04996-4
Chi, G., Fang, H., Chatterjee, S., & Blumenstock, J. E. (2022). Microestimates of wealth for all low- and middle-income countries. PNAS, 119(3), e2113658119. https://www.pnas.org/doi/10.1073/pnas.2113658119
Croft, T. A. (1978). Nighttime Images of the Earth from Space. Scientific American, 239(1), 86–101. https://www.scientificamerican.com/article/nighttime-images-of-the-earth-from/
da Vinci, L. (1817 edition). Trattato della Pittura. Rome: Stamperia de Romanis. https://books.google.com/books?id=FJI7-YOd8PAC
Donaldson, D., & Storeygard, A. (2016). The View from Above: Applications of Satellite Data in Economics. Journal of Economic Perspectives, 30(4), 171–198. https://www.aeaweb.org/articles?id=10.1257/jep.30.4.171
Douglass, A. E. (1909). Weather Cycles in the Growth of Big Trees. Monthly Weather Review, 37(6), 225–237. https://journals.ametsoc.org/view/journals/mwre/37/6/1520-0493_1909_37_225d_wcitgo_2_0_co_2.xml
Douglass, A. E. (1919). Climatic Cycles and Tree-Growth: A Study of the Annual Rings of Trees in Relation to Climate and Solar Activity. Carnegie Institution of Washington. https://books.google.com/books?id=z8sDAAAAIAAJ
Einav, L., & Levin, J. (2014). Economics in the age of big data. Science, 346(6210), 1243089. https://www.science.org/doi/10.1126/science.1243089
Elvidge, C. D., Baugh, K. E., Kihn, E. A., Kroehl, H. W., Davis, E. R., & Davis, C. W. (1997). Relation between satellite observed visible-near infrared emissions, population, economic activity and electric power consumption. International Journal of Remote Sensing, 18(6), 1373–1379. https://doi.org/10.1080/014311697218485
Henderson, J. V., Storeygard, A., & Weil, D. N. (2012). Measuring Economic Growth from Outer Space. American Economic Review, 102(2), 994–1028. https://www.aeaweb.org/articles?id=10.1257/aer.102.2.994
Jean, N., Burke, M., Xie, M., Davis, W. M., Lobell, D. B., & Ermon, S. (2016). Combining satellite imagery and machine learning to predict poverty. Science, 353(6301), 790–794. https://www.science.org/doi/abs/10.1126/science.aaf7894
Khosravi, F., & Liu, E. (2026). Generative AI and Firm Hiring Demand: Evidence from Advanced and Developing Economies. SSRN working paper. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6537038
Michel, J.-B., Shen, Y. K., Aiden, A. P., Veres, A., Gray, M. K., et al. (2011). Quantitative Analysis of Culture Using Millions of Digitized Books. Science, 331(6014), 176–182. https://www.science.org/doi/abs/10.1126/science.1199644
Moritz, M., et al. (2026). The Effect of Citywide Improvement in Street Lighting on the Use of Public Space by Community Residents. Justice Quarterly. https://www.tandfonline.com/doi/full/10.1080/07418825.2026.2712831
Yabe, T., García Bulle Bueno, B., Frank, M. R., Pentland, A., & Moro, E. (2025). Behaviour-based dependency networks between places shape urban economic resilience. Nature Human Behaviour, 9, 496–506. https://www.nature.com/articles/s41562-024-02072-7