Fundraising is a critical activity for non-profit organizations, social enterprises, and mission-driven startups, yet it remains one of the least data-informed processes in the impact ecosystem. Despite the growing availability of data on foundations, CSR programs, and donor behavior, most organizations still rely on manual research and intuition to identify potential funders. This results in low response rates, limited visibility into what drives engagement, and little opportunity to learn systematically from previous campaigns. The process is not only time-consuming but also unpredictable, as outreach efforts often depend on subjective judgment rather than quantitative evidence.

This project explored how data science and machine learning can be applied to optimize the fundraising process. By transforming unstructured funder information into analyzable datasets, we can identify latent patterns, cluster similar funders, and predict which entities are most likely to respond positively to outreach. The analytical framework combines unsupervised and supervised learning, natural language feature extraction, and experimental design to create an adaptive decision system for fundraising. The goal is to shift the process from static and reactive to dynamic and evidence-driven, where data, not intuition, guides targeting, messaging, and prioritization decisions.

To accomplish this, we created an end-to-end pipeline that can gather contextual data, segment the various organizations, organize outreach campaigns, track responses, and make predictions on successful interactions. The pipeline was put into a single application for ease of use. This reduces the time spent gathering data and creating capacity to do larger campaigns.

By applying data science to an area traditionally dominated by human intuition, this project demonstrates how machine learning can enhance decision-making and efficiency in the social impact domain. The broader objective is to build a repeatable, scalable framework that enables organizations to fundraise more intelligently using analytics to learn from every interaction, reduce wasted effort, and ultimately achieve greater impact with the same resources.

Watch the team present this project at 02:15:56 in the session recording here.

Faculty Advisor

Fan Yang is an experienced professional in the financial industry, blending over a decade of expertise in data science and financial modeling with a background in consumer banking. Specializing in Deposits, Residential Loans, Auto Loans, and Credit Cards. His journey has been marked by contributions to loss forecasting, marketing analytics, card acquisition model development, and credit data management. Currently serving as a Senior Manager at Discover Financial Service in Chicago, where he leads the Card Acquisition Risk Modeling team, overseeing both onshore and offshore operations. His role involves building machine learning models and providing strategic data-driven support for various business needs.

His professional narrative also includes tenures at BMO, where he led the Marketing Advanced Data Analytics team, and KeyBank/Northern Trust, focusing on Comprehensive Capital Analysis and Review (CCAR) and Current Expected Credit Loss (CECL) Stress Tests. In these roles, he used advanced quantitative techniques XGBoost, and Convolutional neural network (CNN), helping companies drive business growth and optimize risk management strategies effectively. Fan’s educational background includes a Ph.D. in Statistics and an MBA from the University of Iowa. He currently teaches Time Series Analysis and Forecasting at the University of Chicago and has three years of prior experience teaching at the University of Iowa.

arrow-left-smallarrow-right-large-greyarrow-right-large-yellowarrow-right-largearrow-right-long-yellowarrow-right-smallclosefacet-arrow-down-whitefacet-arrow-downCheckedCheckedlink-outmag-glass