Work

Duke Global Health Innovation CenterData / research2026

Financial Codebook

AI & Global Health Financing Analysis

Extending an inherited research dataset into a cleaner, more explicit analytical layer for studying how AI and global-health organizations are financed.

Data cleaning · standardization · analytical definitions · Python-supported analysis

Context
Inherited research dataset on AI and global-health financing
Role
Data cleaning, standardization, analysis, documentation
Focus
Financing patterns, classification, data quality
Status
Preliminary analysis · handed off

Where I started

The Financial Codebook existed before I worked on it. An earlier three-person student team built it in Summer 2025 to track publicly announced financing for AI and global health, structured across four main funder types: government, philanthropy, corporate, and venture capital.

That version already included analyses of funding, geography, AI type, health area, recipient characteristics, and funding mechanisms. My work began from that foundation, not from an empty file.

Inherited project state (earlier version)

~393entries in the earlier version
31data fields
4main funder categories

These describe the version I inherited. They are not a measure of my work.

The data problem

The difficult part was not getting every cell filled. It was deciding when the available evidence supported a classification, and when it did not.

The records came from public announcements, so the underlying information was uneven. Some entries carried detailed financing information; others had a name and very little else.

The issues I kept working through were missing values, entity types that were inconsistent or unclear, categories that applied to some organizations but not others, incomplete financing information, ambiguous stage evidence, and records that should be marked unknown or not applicable rather than pushed into a category.

Different working versions of the dataset also contained different record counts and units of analysis. Rather than treating those counts as directly comparable, I focused on standardizing the records needed for the next analytical stage.

Working analysis flow

How I structured the next analytical stage

  1. Inherited codebookthe earlier dataset and its fields
  2. Clean and standardizeconsistent types, values, and naming across versions
  3. Entity typefor-profit, non-profit or research, unclear
  4. Record sufficiencywhether there is enough here to analyze at all
  5. Derived financing signalsgrant dependence, funding diversity, financing stage
  6. Applicability checkswhether this category applies to this entity
  7. Preliminary cross-analysispatterns across the applicable records
  8. Document and hand offdefinitions, limitations, open items

This was not a fully automated pipeline. Records moved through this structure with manual review and revision at several points.

Data decisions

Three decisions shaped how useful the dataset could be. None of them are impressive on their own, which is part of the point.

01

Missing is different from not applicable

One of the recurring problems was that not every classification applied to every organization. A financing-stage label can make sense for a startup, but not for a university, a government institution, or many non-profits. Instead of forcing every row into the same taxonomy, I separated applicability from missingness, so that not applicable and unknown could mean different things.

02

A derived category needs a basis

For applicable entities, financing stage was often inferred rather than stated. The processing kept a basis alongside the stage itself, distinguishing whether an estimate came from an equity round, an amount proxy, a description proxy, no evidence at all, or non-applicability. Without that field, a stage column looks more certain than the evidence behind it. The proxies are inference bases, not ground truth.

03

Analysis populations need explicit boundaries

The later analysis file contained 899 company-level records. Treating all of them as comparable companies would have been wrong: 702 were for-profit, 119 were non-profit or research organizations, and 78 were unclear. Financing-stage analysis was therefore limited to the applicable for-profit subset, while other signals used the wider set.

Analytical fields

The later working layer carried these fields. This is a register of names and meanings, not a sample of data.

Entity
the organization the record describes
Entity type
for-profit, non-profit or research, or unclear
Funding type
the kind of financing recorded
Funding source
who provided it, where disclosed
Record sufficiency
whether the record holds enough to support downstream analysis
Grant dependence
how strongly the financing evidence depends on grants
Funding diversity
whether the record shows one financing source or several
Financing stage
for applicable entities, the stage the evidence can reasonably support
Stage basis
why that stage was assigned: equity round, amount proxy, description proxy, none, or not applicable

What the preliminary analysis could support

The signals describe aspects of an organization's financing profile as recorded in this dataset. They are not validated measures of business-model sustainability, and this page does not present them that way.

A preliminary cross-analysis sat on top of them, scoped to the applicable populations rather than treating every record as comparable. It stayed preliminary: a meaningful share of classifications still required review, and the dataset is a working research file, not a census of the sector.

What the data could not tell me

The limitations are part of the result, not a footnote to it.

  • Public financing information is incomplete: many rounds and grants are never announced, and amounts are often undisclosed.
  • Source quality varies across records, countries, and organization types.
  • Entity types differ, so categories such as financing stage apply to some organizations and not others.
  • Inferred stages may rest on proxies, which are weaker evidence than a disclosed round.
  • Working versions of the dataset used different units and record counts, so counts across versions are not directly comparable.
  • The financing signals are working definitions. They are not validated sustainability measures.
  • Manual review remains necessary for a substantial share of records.

Handoff

The project was still in progress when I handed it off. It was not finished, and this page should not pretend otherwise.

What I left behind was the cleaned working data, the signal definitions, the processing logic used to derive and check classifications, the preliminary analysis, and documentation of which items still needed review.

What I took from it

This project changed how I think about clean data. The hardest decisions were usually not technical ones. They were about whether two records were truly comparable, whether a category applied at all, and how much confidence the available evidence justified.

Making those decisions explicit made the analysis more useful, even when the answer was not applicable or needs review. A dataset becomes more useful when its categories, assumptions, and limits are visible, not when every blank is filled.

Sources and notes

This page is based on the project's working files and methods documents: the inherited codebook and its earlier presentation, the working workbooks and cleaned datasets, the processing and validation scripts, the signal definition notes, the preliminary cross-analysis, and the handoff documentation. Those files are not published here.

  1. Earlier Financial Codebook presentation and dataset (Summer 2025 student team).
  2. Working workbooks and cleaned datasets, later versions.
  3. Processing and validation scripts for financing-stage classification, including the stage-basis field.
  4. Signal definition notes and preliminary cross-analysis documents, v3 series.
  5. Handoff documentation for the next team.

The original Financial Codebook was created by an earlier student team. My work was to extend, clean, standardize, and document it, and to build the first version of the analytical layer described here. The project remained unfinished at handoff. Record counts on this page are scoped to specific versions and units of analysis, and are not directly comparable across versions.

← All work