please write the research paper on this Predicting Financial Inclusion Gaps Using Alternative Data (Mobile & Satellite) for Underbanked Populations Why this works well for a scholarship committee: Clear social impact — over 1.4 billion adults globally remain unbanked; this ties directly into UN SDG 8 and 10, which most scholarship boards (Fulbright, Chevening, DAAD, Erasmus Mundus, etc.) explicitly favor. Strong stats/math fit — you'd use techniques you already have a foundation for: logistic regression, hierarchical/multilevel models, spatial statistics, and gradient boosting, plus causal inference methods (propensity score matching, difference-in-differences) if you want to go further. Feasible with public data — no need for proprietary bank records to get started. Core research question: Can mobile phone usage patterns and satellite-derived socioeconomic indicators predict financial inclusion gaps at a sub-national level, and which features matter most for targeting interventions? Data sources (all public/accessible): World Bank Global Findex Database (financial inclusion survey data) DHS (Demographic and Health Surveys) — household wealth indices Satellite night-light data (NOAA VIIRS) as a proxy for regional economic activity Mobile money penetration data (GSMA reports) Methodology outline: Merge geospatial + survey data at district/regional level Build predictive models (start simple: logistic regression baseline → random forest/XGBoost) for financial inclusion status Use SHAP values or similar for interpretability — important since reviewers like policy-relevant, explainable outputs Optionally, add a causal layer: does mobile money rollout cause increased inclusion (using a natural experiment or DiD design)?
Predicting Financial Inclusion Gaps Using Alternative Data: Mobile and Satellite Signals for Underbanked Populations
Abstract
Financial exclusion is spatially uneven, changes faster than conventional surveys can observe, and is concentrated among people who are least visible to formal administrative systems. The 2025 Global Findex reports that 79% of adults globally had an account, implying a derived residual of roughly one in five adults without an account; the same edition covers about 148,000 adults in 141 economies and adds globally comparable measures of mobile ownership, internet use, and digital safety [citation_1og1xr]. This paper develops a feasible research design for predicting sub-national financial-inclusion gaps from mobile-phone behavior, mobile-money activity, and satellite-derived socioeconomic indicators. The central proposal is a multimodal, hierarchical geospatial model trained against survey or administrative labels and evaluated under geographically separated validation. Logistic regression provides an interpretable baseline; random forests and gradient boosting test nonlinearities and interactions; spatial or multilevel models quantify residual geographic structure; and SHAP-style explanations translate predictive patterns into policy-relevant features without treating associations as causal effects.
The evidence supports feasibility but imposes important limits. Mobile call-detail records (CDRs) have explained 50–65% of national socioeconomic-status variance across Namibia, Nepal, and Bangladesh [citation_1og1ws]. A multimodal study reported a highest R² of 0.78 for traditional poverty measures [citation_11fvsx], while a Rwanda model combining sparse CDR features, night lights, and population density reported a cross-validated correlation of 0.88 for a multidimensional poverty index [citation_1og1wt]. These precedents do not show that financial inclusion can be inferred everywhere. Phone-based models exclude nonusers; survey and operator coverage may not overlap; nighttime lights perform better in urban than rural areas; and digital targeting can reproduce exclusion through privacy, fairness, and representativeness failures [citation_1og1wq]. A causal module using staggered rollout and difference-in-differences must remain separate from predictive importance.
Keywords: financial inclusion; mobile money; call-detail records; satellite night lights; VIIRS; Global Findex; DHS; poverty mapping; machine learning; spatial statistics; causal inference
1. Introduction
Financial inclusion should be treated as a multidimensional outcome rather than a binary label. Account ownership measures formal access, but meaningful inclusion also involves the ability to make and receive payments, save, borrow, manage shocks, and use services safely. The Global Findex was designed to measure how adults save, borrow, make payments, and manage financial risk, and its 2021 microdata contain approximately 120 variables on bank accounts, mobile money, digital payments, savings, credit, and financial resilience [citation_1og1yv]. The 2025 edition extends this agenda with mobile connectivity and digital-safety indicators [citation_1og1xs]. A district with high account ownership but weak active use, poor network quality, or limited agent liquidity may therefore have a different policy problem from a district with no nearby access points.
The core research question is: Can mobile-phone usage patterns and satellite-derived socioeconomic indicators predict sub-national financial-inclusion gaps among underbanked populations, and which features matter most for targeting interventions? The question is policy-relevant because conventional demand-side surveys are periodic and often designed for national or broad regional inference, whereas mobile and satellite signals can be observed repeatedly and spatially. Yet the research should not assume that a proxy for wealth is a proxy for inclusion. Wealth, connectivity, distance to an agent, trust, documentation, literacy, gender norms, and the quality of the service network may affect inclusion through different mechanisms.
The proposed framework has three outputs: sub-national inclusion prevalence; an inclusion gap, defined as predicted access potential minus observed use; and an estimate of whether mobile-money expansion changes inclusion or welfare. The first two are predictive and diagnostic; the third is causal. This distinction matters because a variable may predict inclusion without being an intervention lever, while a policy may have a causal effect without being the strongest cross-sectional predictor.
2. Conceptual framework and hypotheses
The conceptual model has four layers. The first is material capacity, represented by household wealth, assets, population density, built-up area, roads, electricity proxies, and local economic activity. The second is digital opportunity, represented by phone ownership, network coverage, device type where available, call and SMS activity, mobility, and connection quality. The third is financial-service supply, represented by mobile-money agents, cash-in and cash-out activity, transfers, geographic distance to branches or agents, and service availability. The fourth is individual and institutional frictions, including documentation, literacy, trust, affordability, gender, age, and distance to formal institutions. The Findex literature identifies cost, physical distance, and documentation among reported barriers to account use [citation_10ddf6], while individual-level evidence from ASEAN countries finds that age, income, education, employment, mobile ownership, distance to formal institutions, and trust are associated with financial inclusion, with relationships differing across countries and periods [citation_1og1zg].
The first hypothesis is that a multimodal model will outperform models using one data family at a time. Mobile data capture behavior and mobility at high temporal resolution; satellite data provide countrywide spatial coverage and avoid the phone-user selection problem; survey data provide the outcome and demographic structure. In poverty mapping, combined mobile and geospatial models produced the best predictive power in one multi-country study, with a highest reported R² of 0.78 [citation_11fvsx]. A Rwanda study similarly combined mobile ownership per capita, calls per phone, normalized night lights, and population density to estimate sector-level multidimensional poverty, reporting a cross-validated correlation of 0.88 [citation_1og1wt]. These findings motivate data fusion, but the proposed study will test whether the gain persists for financial-inclusion outcomes rather than assume transferability.
The second hypothesis is that mobile-network and mobile-money variables will matter most where service supply is the binding constraint, whereas wealth and built-environment variables will matter more where affordability or economic capacity is the binding constraint. Evidence from M-Pesa shows that mobile-money adoption could be predicted with AUC 0.691 and spending with AUC 0.619; the most predictive features were mobile-phone activity, M-Pesa users in the person’s social network, and mobility [citation_10c71s]. This suggests that social diffusion and network embeddedness may be more informative than simple phone counts. It does not establish that these variables cause adoption.
The third hypothesis is that predictive performance will be heterogeneous. CDR-based mobility and call behavior explained 50–65% of the variance in socioeconomic status across three countries, but the study emphasized the importance of ancillary data and local context [citation_1og1ws]. Satellite night lights are a strong sub-national predictor of wealth after controlling for population density, but their relationship with the DHS wealth index is less pronounced in rural than urban areas [citation_1og1x0]. In addition, night lights may saturate in bright urban areas and be weak in dispersed agricultural settlements [citation_1og1wo]. The model should therefore estimate urban-rural interactions and report subgroup performance rather than publish one national accuracy number.
3. Data architecture
3.1 Outcome and reference data
The primary outcome should be defined at the smallest unit for which the reference survey is representative. Candidate outcomes include account ownership; personal use of mobile money; recent digital payment use; saving or borrowing through a formal or mobile channel; and a composite inclusion index. The Global Findex is the principal cross-country benchmark because it is a demand-side survey with repeated waves and indicators on access, use, payments, savings, borrowing, and resilience [citation_1og1yr]. The 2025 release provides country-level files and individual-level microdata, with indicators available for 2011, 2014, 2017, 2021, and 2024 [citation_1og1yi].
A critical design limitation is geographic support. Public Findex reporting is organized by country, region, income group, and demographic strata, while the individual microdata are nationally representative [citation_1og1yi]. The researcher should not infer district-level labels from a nationally representative sample unless the survey documentation explicitly supports that level. A defensible implementation has two tracks. Track A uses Findex to benchmark country and demographic patterns and to test whether alternative signals predict broad inclusion differences. Track B selects countries with a public survey, census, financial-access survey, or administrative source that contains an inclusion outcome and supports sub-national estimation. Findex can then serve as an external benchmark rather than an artificially precise district label.
DHS contributes an asset-based wealth index and demographic or household covariates. The mobile-poverty literature has used the DHS wealth index as a survey-based socioeconomic-status measure [citation_1og1ws], and night-light research has validated local inequality estimates against DHS-derived inequality [citation_1og1xp]. DHS wealth should be used as a socioeconomic covariate, stratification variable, or auxiliary label—not silently relabeled as financial inclusion. Where a selected DHS questionnaire contains relevant technology or financial-use items, those variables can become outcomes; otherwise, a separate inclusion survey or administrative label is required.
3.2 Mobile data
The mobile feature set should prioritize aggregates that can be computed without exposing raw communications. Candidate variables include active subscribers per capita; outgoing and incoming call counts; call duration; SMS counts; data-session activity; number of unique contacts; network diversity; mobility radius; regularity of home and work locations; cell-tower handoff counts; connection quality; phone type; and distance to the nearest agent or branch. Mobile-money features include active accounts, transaction frequency, transfers, cash-in, cash-out, agent density, and the balance between incoming and outgoing flows. The Ghana-Uganda poverty-estimation work used anonymized CDR and mobile-money data, including incoming and outgoing calls, SMS, cash-in, cash-out, transfers, and cell-tower geolocation [citation_1og1wn].
The unit of analysis should be a district-month or grid-month when temporal data are available, with features aggregated using counts, rates, medians, quantiles, entropy, and changes from a baseline. Counts must be normalized by population, estimated subscriber coverage, or active-user denominators. Otherwise the model may learn operator market share rather than inclusion. The model should include missingness indicators: low phone activity may be a measurement problem, a signal of weak connectivity, or both. The distinction is policy-relevant. Satellite indicators generally cover the full country, whereas CDRs represent phone users and are harder to obtain for privacy reasons, although CDRs can provide richer information about behavior and connection quality [citation_1og1yp].
3.3 Satellite and contextual data
VIIRS nighttime-light intensity can represent local economic activity, settlement concentration, electrification, and temporal change. The signal should be summarized at multiple spatial scales, such as 1 km cells, district means, population-weighted means, lit-area share, upper quantiles, and within-district inequality. Grid-cell aggregation is preferable to arbitrary buffers around survey clusters; local night-light research found significant relationships down to approximately 1 km² and recommended grid-cell analysis [citation_1og1xo]. Other public geospatial covariates should include population density, built-up area, land cover, roads, travel time to settlements, vegetation, elevation, rainfall, and distance to financial-service points where available.
Night lights must not be treated as a universal welfare measure. A 2025 multi-country study found that harmonized night lights remained a strong predictor of the DHS wealth index after controlling for population density, but the association was notably weaker in rural areas [citation_1og1x0]. Research on rural economic activity argues that agriculture and dispersed settlements generate weak or diffuse night illumination, creating a structural limitation rather than simply a sensor problem [citation_1og1xw]. In rural districts, daytime imagery, roads, settlement footprints, travel time, and mobile coverage may be more informative than radiance alone. The study should pre-specify these interactions and avoid a single global ranking based on light intensity.
3.4 Mobile-money market data
GSMA and operator data can describe regional penetration, active accounts, agent networks, and transaction volumes when geographic resolution permits. They are supply-side context, not individual inclusion: registered accounts may be inactive, and transaction value may reflect remittance corridors rather than broad access. Use survey outcomes for inclusion, operator aggregates for exposure, and GSMA statistics for external context.
4. Modeling strategy
The baseline should be a weighted logistic regression for individual inclusion or a fractional logit model for district prevalence. A multilevel specification can include random intercepts for country and district and random slopes for mobile-money exposure or rurality. For a district-level prevalence outcome, a binomial model with survey sample size is preferable to treating every district percentage as equally precise. Spatially structured random effects can capture residual clustering, while a spatial error or conditional-autoregressive component can test whether the covariates leave systematic geographic patterns unexplained.
The machine-learning comparison should include random forest and gradient boosting, with XGBoost as the principal nonlinear model if the sample size supports it. Models should be trained in nested cross-validation, with hyperparameter tuning performed only inside the training folds. Randomly splitting neighboring grid cells is not sufficient because spatially adjacent observations can share the same night-light pixels, mobile towers, roads, and survey clusters. The primary test should hold out complete districts or regions; a stronger test should hold out one country or one survey wave. Temporal validation should train on an earlier period and test on a later period to assess drift.
Report metrics matched to the outcome: AUROC and precision-recall area for binary status; RMSE, MAE, rank correlation, and calibration error for district prevalence; and prediction intervals for maps. Report performance separately for urban and rural areas, women and men, phone users and—where observable—nonusers, and wealth strata. High aggregate accuracy with systematic underprediction of remote rural districts is not equitable targeting.
Feature attribution should use SHAP or a comparable local-and-global explanation method to display how each variable changes predictions. Explanations should be grouped into interpretable families—economic activity, connectivity, mobility, social-network structure, and financial-service supply—to reduce false precision from correlated variables. The paper should explicitly state that feature importance is associative. For example, mobile-money activity may identify places already better connected, while night lights may proxy both wealth and electricity. Policy recommendations should therefore target a feature only after a causal analysis or a field intervention supports its modifiability.
The inclusion-gap map can be operationalized as:
[G_i = \widehat{P}(I_i=1\mid X_i^{capacity}], [X_i^{connectivity}) - P^{observed}(I_i=1)], []
where the first term is a model-based expected inclusion level and the second is the survey or administrative estimate. A positive value indicates lower observed inclusion than the model predicts, potentially pointing to barriers such as trust, documentation, agent liquidity, affordability, or gendered access. Because the two terms are estimated with error, the paper should publish uncertainty intervals and avoid ranking districts whose intervals overlap substantially. The exact gap definition must also be stress-tested by changing the covariate families included in the expected-inclusion model.
5. Causal extension: mobile-money rollout
The predictive model can identify where intervention may be useful, but it cannot establish that rollout caused inclusion. The causal extension should exploit staggered expansion of mobile-money agents, network coverage, or service availability across districts. An event-study specification can compare treated and not-yet-treated districts before and after rollout, with district and time fixed effects, country-specific trends where defensible, and pre-trend diagnostics. If treatment timing is staggered, the analysis should avoid interpreting a conventional two-way fixed-effects coefficient without checking heterogeneous-treatment-timing bias.
The treatment should be defined carefully: first agent opening, density threshold, network coverage threshold, or mobile-money availability. Outcomes should include account ownership, recent use, digital payments, savings, resilience, and—only as secondary outcomes—economic activity or poverty. A Kenya study combined early mobile-agent expansion with high-resolution night lights and found that access increased local economic activity, with larger effects in initially more affluent, urban, and better-connected areas [citation_1og1x1]. This heterogeneity is directly relevant to a targeting design: rollout may increase activity while leaving the most excluded districts behind.
A Tanzania study provides a stronger causal precedent for welfare mechanisms. It exploited rapid agent-network expansion between 2010 and 2012 together with rainfall shocks in an instrumented difference-in-differences design and found that adopters smoothed consumption during shocks and maintained human-capital investments [citation_1og1yx]. The proposed project should treat this as evidence that mobile money can affect resilience under particular institutional and market conditions, not as a universal estimate. A separate six-country study found a robust association between network accessibility and mobile-money use only in Pakistan and Tanzania; in those settings, moving 10 km closer to a multiple-network area was associated with a 10% increase in the probability of use, while the relationship was not robust in the other countries [citation_1og1yw]. This mixed result argues for country-specific causal identification rather than a pooled claim that connectivity automatically produces inclusion.
The causal and predictive modules can be linked through heterogeneous-treatment-effect analysis. Use the pre-rollout model to estimate baseline vulnerability, then test whether rollout effects differ by predicted gap, rurality, wealth, and agent distance. This answers a policy question more useful than “which feature is most important?”: where does an additional agent or network upgrade produce the greatest increase in inclusion? That estimate still requires credible treatment variation and should be reported with uncertainty.
6. Threats to validity and ethical safeguards
The first threat is selection into mobile use. Phone-based models predict the welfare or behavior of subscribers, not necessarily the full population. Recent multi-country evidence reports better performance in nationally representative and heterogeneous samples than in homogeneous urban-only, rural-only, or otherwise restricted samples, and notes that phone-based models exclude people without phones [citation_1og1yj]. In Uganda, only one-third of respondents in a rural-focused survey had phones, limiting matches to operator data; survey areas also overlapped poorly with the operator’s meaningful market share [citation_1og1ym]. The study should therefore estimate phone ownership or coverage explicitly, report a “coverage-adjusted” scenario, and avoid using a phone-derived score as a direct eligibility rule for people who cannot generate the data.
The second threat is label and time mismatch. Survey interviews, satellite composites, and mobile aggregates may refer to different periods. A current mobile signature paired with an old DHS wealth index may capture a changed settlement or temporary shock. The model should use only features preceding the outcome window in causal analyses, record exact dates, include lagged and change features, and test temporal transfer. It should distinguish stable wealth from volatile income or food security: CDR-based predictions have been more accurate for wealth than for short-term outcomes, and an Haiti impact-evaluation application failed to recover statistically significant expenditure impacts from CDR predictions even when conventional survey estimates were positive and significant [citation_1og1wu].
The third threat is spatial and construct validity. Night lights can measure electrification and commercial concentration as much as welfare; mobile-money flows can be driven by firms or remittance corridors; and districts may not match tower or agent service areas. Use population-weighted aggregation, grid-size sensitivity checks, boundary-robust features, and uncertainty propagation. Compare light-only, mobile-only, contextual-only, and multimodal specifications so the gain from fusion is transparent.
The fourth threat is algorithmic exclusion. A prediction map may be used to deny service, price credit, or allocate public benefits. Evidence from digital targeting emphasizes that machine-learning approaches can outperform broad categorical rules when traditional data are missing, yet generally remain less accurate than traditional survey-based poverty measurement and raise risks involving privacy, transparency, fairness, and digital exclusion [citation_1og1wq]. The appropriate policy use is triage and resource prioritization, not automated denial. Every operational score should have an appeal route, a human review protocol, a documented error budget, and a prohibition on using inferred vulnerability for punitive or commercial purposes without separate legal and ethical authorization.
Privacy protection should be designed into the data architecture. Raw CDRs should remain with the operator; researchers should receive only pre-specified, aggregated, de-identified features at a spatial and temporal resolution that limits re-identification. A recent Bangladesh targeting study used consent, operator-side feature extraction, removal of phone identifiers, anonymized survey merges, and secure storage so that the research team did not access raw CDRs [citation_1og1wx]. A similar operational recommendation is to run standard feature generation within the operator and share only geographic poverty values, preserving individual privacy [citation_1og1yn]. The proposed study should publish a data dictionary and model card, minimize feature granularity, suppress small cells, audit subgroup error, and disclose who can access each data layer.
7. Expected contribution and research gaps
The project’s scholarly contribution is a validated bridge between financial-inclusion measurement and multimodal small-area estimation. Existing studies establish that mobile behavior and geospatial signals can predict socioeconomic status, but they do not justify assuming that a poverty map is an inclusion map. The proposed design addresses this gap by using direct inclusion outcomes where available, treating DHS wealth as an auxiliary socioeconomic measure, and defining the inclusion gap as a measurable discrepancy between expected capacity and observed use.
The most important empirical gap is generalization beyond selected operators, countries, and phone users. Cross-setting studies are promising, but performance depends on population heterogeneity, coverage, market share, and the outcome’s stability [citation_1og1ys]. A second gap is temporal validity: much of the evidence is cross-sectional, so it remains uncertain whether mobile signals detect changes in livelihoods or inclusion reliably over time. A third gap is causal targeting: predictive scores rarely tell policymakers whether investment in agents, coverage, literacy, affordability, or consumer protection will close a measured gap. A fourth is rural validity, because night lights and phone data can both be least informative where exclusion is deepest.
The project should therefore preregister a transportability plan. It should select at least one urban and one rural setting, conduct leave-one-region-out validation, test a later period, and report failure cases. It should compare model performance against a simple geographic baseline and against survey-only models. It should also conduct a small qualitative validation with local financial-service providers or community organizations—not to replace statistical validation, but to determine whether highly ranked features correspond to actionable barriers such as distance, liquidity, trust, documentation, or affordability.
8. Conclusion
Mobile and satellite data can plausibly support sub-national financial-inclusion diagnostics, especially when they are fused with survey labels and evaluated under geographic and temporal holdouts. The strongest precedent is not a single algorithm but the combination of complementary signals: CDRs capture behavior and mobility, mobile-money data capture service use and networked transactions, and satellite indicators provide broad spatial coverage. Prior work reports substantial predictive performance for poverty and socioeconomic status, including a multimodal R² of 0.78 and a Rwanda MPI correlation of 0.88, but these results should be treated as feasibility evidence rather than guarantees for financial-inclusion prediction [citation_11fvsx] [citation_1og1wt].
A scholarship-quality study should make three disciplined claims. It can claim that alternative data improve measurement where conventional data are sparse, if multimodal performance beats transparent baselines under honest validation. It can claim that particular features are associated with observed gaps, if subgroup and robustness analyses support that interpretation. It can claim that mobile-money rollout changes inclusion or welfare only when a credible quasi-experimental design demonstrates it. The recommended output is therefore an interpretable, uncertainty-aware inclusion-gap map accompanied by a causal analysis of rollout where the institutional setting permits it—not an opaque ranking of supposedly deserving populations. That framing preserves the project’s mathematical strength while keeping its social-impact argument aligned with the people the model is intended to serve.
Generated Outputs
[SANDBOX_OUTPUT:e6a7eb578b392f8e:financial_inclusion_alternative_data_paper.md]