Day 13: Why Causal Claims Are Hard in Networks

Four Published Studies and the Problems They Had to Solve

Shahryar Minhas

How I Want You to Use This Lecture

This is not a training session for reproducing four expensive studies.

Use the cases to learn how to:

  1. Recognize why a causal claim is difficult.
  2. Decide what comparison your own data can support.
  3. Improve the design before reaching for a complicated model.
  4. Make the strongest claim the evidence can honestly carry.

Most of us will not randomize a national election-monitoring program or a school-wide intervention. That is fine. The practical goal is to become better at separating a credible causal comparison from an adjusted association, and to know what would make your own study more convincing.

You do not need a perfect experiment to do useful research. You do need to be clear about what your design can and cannot establish.

What Are We Trying to Learn?

Causal inference asks a simple question: what would change if we took one action rather than another? The hard part is that we never see the same person, school, or community experience both versions at the same time.

For every application, we will ask:

  1. What action changed?
  2. Whose outcome could change?
  3. What two situations are being compared?
  4. Why should we trust that comparison?

The goal is to recognize the problem and read the evidence carefully. You are not expected to reproduce these studies.

A Causal Effect Compares Two Possible Outcomes

Imagine that we want to know whether a voting message raises turnout.

Possible Outcome 1

What this voter would do after hearing the voting message.

Possible Outcome 2

What this same voter would do after hearing the comparison message.

The causal effect is the difference between those two possible outcomes. We observe only one, so we need another group that gives us a credible picture of the missing outcome.

Random assignment is useful because it can create groups that were comparable before the intervention.

Why Networks Make This Harder

One Route

My outcome may change because I receive the intervention.

A Second Route

My outcome may change because someone connected to me receives the intervention.

In a network, an untreated person with a treated friend has not had the same experience as an untreated person with no treated friends. We need to describe both people’s own assignments and what happened around them.

Other people’s assignments can matter, and people often choose who they are connected to.

Four Applications, Four Versions of the Problem

Application Setting Main Difficulty
Nickerson (2008) Voting messages in two-voter households One intervention may affect two people
Paluck, Shepherd, and Aronow (2016) An anti-conflict program in middle schools Students can be reached through assigned friends
Ichino and Schündeln (2012) Election observers during voter registration in Ghana Monitoring may move behavior to nearby places
Egami and Tchetgen Tchetgen (2024) Friends’ GPA and students’ later GPA Friends may resemble one another before influence occurs

The cases use different methods, but the same questions keep returning: what changed, who could respond, what is the comparison, and what could still make that comparison misleading?

Case 1: A Voting Message Enters a Household

Nickerson (2008) studied households with two registered voters in Denver and Minneapolis during the 2002 congressional primary elections.

A canvasser knocked on the door and spoke with the person who answered. Successfully contacted households heard either a message encouraging voting or a message encouraging recycling. Administrative voting records then showed whether the person who heard the script voted and whether the other registered voter in the household voted.

Action

Voting message instead of recycling message.

Outcomes

Turnout for the person contacted and turnout for the other voter.

The Household Creates Two Comparisons

Person Who Heard the Script

Was turnout higher under the voting message than under the recycling message?

Other Household Voter

Was turnout higher when the person at the door heard the voting message than when that person heard the recycling message?

The second comparison is the network part. The other voter did not hear the canvasser’s script, but the message may have traveled through conversation or another household interaction.

The experiment changes the message. It does not directly change whether the first person votes.

Turnout Was Higher for Both People

Compared with the recycling-message group, turnout was about 9.7 percentage points higher among the people who heard the voting message and about 5.9 points higher among the other voters in their households.

What the Household Study Teaches Us

The study gives us evidence that assigning the voting message changed turnout for both people in a successfully contacted household.

It does not show that:

  • The first person’s act of voting caused the other person’s vote.
  • Every attempted household would respond in the same way.
  • Household conversation was the exact pathway.

The randomized intervention was the message. The partner result is a spillover from that message. The study did not separately randomize one household member’s actual decision to vote.

Always name the action that was actually assigned and the outcome that was actually measured.

“Untreated” Can Hide Different Experiences

Suppose a school invites some students to join a program.

Student Invited? Invited Friend? Experience
Alex No No No direct or friend exposure
Ben No Yes Possible exposure through a friend
Cara Yes No Direct invitation
Dani Yes Yes Direct invitation and friend exposure

Calling Alex and Ben both untreated throws away an important difference. Ben has an invited friend and Alex does not. A network study must decide whether that difference matters for the question.

An Exposure Condition Is Just a Useful Description

An exposure condition summarizes the parts of the assignment that we think could matter for one person’s outcome.

Examples include:

  • Was the person assigned?
  • Did at least one friend receive an assignment?
  • How many close neighbors received an assignment?
  • Was a nearby location assigned monitoring?

The full assignment pattern can be enormous. The exposure condition turns it into a small set of experiences that we can compare.

The network does not choose the exposure definition for us. The intervention and the substantive story should determine it.

The Exposure Definition Can Change the Answer

For a student with assigned friends, we might record:

Any Assigned Friend

Zero versus at least one.

Number Assigned

Zero, one, two, or more.

Close Friends Only

Count only stronger reported friendships.

These definitions answer different questions. “At least one” treats one assigned friend and six assigned friends as the same experience. That may be sensible for a message that spreads once, but not for a process that grows with repeated contact.

A clean calculation cannot rescue an exposure definition that does not match how the intervention could travel.

Some Comparisons Are Not Available

Before comparing two exposure conditions, ask:

Could the people in this analysis realistically have experienced either condition under the study design?

Examples:

  • A student with no recorded friends cannot have an assigned friend.
  • A student who was not eligible for invitation cannot experience direct assignment.
  • A school-level design may make some combinations impossible.

If a person could never enter one side of the comparison, no statistical adjustment can create that missing possibility. Researchers may need to narrow the group they are describing.

Why Weighting Sometimes Appears

Some people are more likely than others to end up in a particular exposure condition.

A student with six eligible friends has more ways to have an assigned friend than a student with one eligible friend. Design-based weighting gives less influence to observations that were especially likely to appear in their condition and more influence to observations that were unlikely to appear there.

The weighting tries to recover a fair comparison under the known randomization. The detailed design-based calculation is in the walkthrough, not something you need to reproduce here.

Case 2: An Anti-Conflict Program in Middle Schools

Paluck, Shepherd, and Aronow (2016) studied 56 public middle schools in New Jersey during the 2012 to 2013 school year. The larger project involved more than 24,000 students between ages 11 and 15.

Half of the schools were assigned to host the Roots program. In those schools, some eligible students were randomly invited to become “seed” students. The students met regularly to identify common sources of conflict in their own schools, then designed slogans, posters, public events, and other activities meant to make handling conflict constructively more visible.

The broader study reported fewer disciplinary incidents involving peer conflict in program schools. Our network question is how the program reached students through their friendships.

What the Orange Wristbands Meant

The orange wristbands, printed with the Roots tree logo, were a visible way to endorse the program’s anti-conflict message. Seed students handed them out for friendly or conflict-reducing behavior. Wearing one therefore gave researchers a concrete sign that the program’s message had reached a student, although it was not itself a measure of whether conflict declined.

The network analysis used friendship information collected before the program to ask whether reported wristband wearing was higher among:

  • Students who received invitations.
  • Students with invited friends.
  • Other students in program schools.

The wristband is useful because it makes program reach visible. It is not a direct measure of whether a student’s behavior changed or whether school conflict declined.

Assignment and Participation Are Different

What Was Randomized

Whether an eligible student received an invitation to join the program.

What Happened Later

Whether an invited student accepted and participated.

About 24% of invited seed students did not accept the invitation.

The cleanest results therefore compare invitation assignments and exposure to friends’ assignments. They are not automatically the effect of actual participation, because participation was a choice made after invitation.

The School Design Creates Several Experiences

Student’s Experience Invited? Invited Friend? Program School?
No program exposure No No No
Program-school context only No No Yes
Friend invited No Yes Yes
Student invited Yes No Yes
Student and friend invited Yes Yes Yes

This table is the heart of the network analysis. It separates direct assignment, possible spread through friends, and the broader experience of attending a program school.

These groups are defined from randomized assignments and the friendship network measured before the program.

The Program Reached Students Through More Than One Route

Compared with eligible students in control schools, reported wristband wearing was 15.4 percentage points higher among students who were not invited but had an invited friend. The largest differences appeared among students who were invited themselves.

What the School Study Does and Does Not Tell Us

The study supports a network-based conclusion:

Assignment to the program reached students who were not invited themselves but had an invited friend.

Keep three limits in view:

  • The measured outcome here is reported wristband wearing, not conflict itself.
  • “At least one invited friend” combines students with very different numbers and kinds of friends.
  • The study identifies effects of assignment-based experiences, not the effect of choosing to participate.

Randomization strengthens the comparison, but we still have to define exposure and interpret the measured outcome carefully.

Case 3: Election Monitoring May Move Behavior Nearby

Ichino and Schündeln (2012) studied Ghana’s 2008 voter-registration period, before a closely contested national election. The number of newly registered voters was approaching 2 million, far above the roughly 800,000 people expected to have become newly eligible by age. Some of the gap had innocent explanations, but it also intensified accusations that parties were moving supporters or registering people improperly.

Why do we care? Registration lists determine who can cast a ballot. If monitoring only moves questionable registration down the road, looking only at the monitored center would make the program appear more successful than it was.

This was a national election. The intervention took place at local registration centers connected by geography and political strategy.

How the Observer Study Worked

Ghana’s Coalition of Domestic Election Observers, or CODEO, sent observers to registration centers during the 13-day registration period. The observers could discourage questionable registration where they were stationed. But political organizers could respond strategically by shifting registration efforts to nearby centers without observers.

The study examined growth in the number of registered voters from 2004 to 2008 at:

  • Centers assigned an observer.
  • Unmonitored centers near observer-assigned centers.
  • More distant centers.

This Spillover Does Not Require Social Influence

Deterrence

Registration growth may fall where an observer is present.

Displacement

Registration growth may rise at nearby unmonitored centers if activity moves away from monitoring.

Nothing here requires friends persuading one another. The connection is geographic: monitoring in one place may change behavior in another place. Network and spillover ideas apply whenever units are connected through a route that can carry a response.

Spillovers can involve information, strategy, movement, competition, or shared resources.

Registration Growth Fell Locally and Rose Nearby

In the fitted comparisons, registration growth was about 3.5 percentage points lower at an observer-assigned center. For an unmonitored center, one additional observer-assigned center within 5 kilometers was associated with about 2.7 points more registration growth.

The Outcome Is Registration Growth

The local decline and nearby increase are consistent with deterrence and displacement.

They do not directly show:

  • Who moved from one registration center to another.
  • Whether every additional registration was improper.
  • Whether the same pattern would appear throughout Ghana.

The experiment gives strong evidence about how observer assignment changed registration growth in the studied areas. Calling all excess growth fraudulent would go beyond the outcome that was measured.

An intervention can move an outcome across connected units instead of simply reducing it.

Case 4: Friends’ GPA and Later Academic Performance

Egami and Tchetgen Tchetgen (2024) study a familiar observational question using Add Health, a large US study that surveyed adolescents in grades 7 through 12 in 1994 and 1995. Students named up to five male and five female friends, and the researchers linked those friendship nominations to later academic outcomes.

Would a student’s later GPA change if their friends had a higher average GPA at baseline?

The application follows more than 10,000 students. GPA summarizes reported grades in English, mathematics, history or social studies, and science. The appeal of the question is clear: friends study together, exchange information, shape expectations, and spend time in the same settings. But students do not receive friends through random assignment. They choose friends, share courses, and often enter friendships because they already have similar backgrounds and interests.

Similar Friends Do Not Automatically Show Influence

Suppose Maya and Jordan are friends and both earn high grades.

Influence Story

Jordan’s study habits help Maya improve.

Selection Story

Maya and Jordan became friends because both were already academically motivated.

Both stories can produce the same observed pattern: friends with similar grades. A regression with many measured controls can still confuse the stories if motivation, family resources, course placement, or other shared causes were not measured well.

This is hidden homophily: an unmeasured reason for friendship also helps explain the outcome.

The Paper Uses Two Negative Controls Together

A basic placebo check asks whether a variable with no plausible effect still predicts the outcome. Friends’ headaches provide that intuition here. Egami and Tchetgen Tchetgen (2024) go farther by pairing an exposure control with an outcome control.

Role Application variable
Main exposure Friends’ baseline GPA
Outcome Student’s later GPA
Negative-control exposure Friends’ headaches or peers-of-peers’ GPA
Negative-control outcome Student’s own baseline GPA

The two controls contain indirect information about hidden friendship selection and shared context. A specialized bridge estimated with generalized method of moments uses that information to revise the peer-effect estimate. This is not the same as simply adding headaches to an ordinary regression.

The Peer-GPA Estimate Became Much Smaller

The conventional adjusted regression estimated a 0.176-point increase in later GPA for a one-point increase in friends’ average baseline GPA. The two negative-control estimates were 0.033 and 0.078, and both were too uncertain to separate from zero.

How to Read the Peer-GPA Result

The result does not prove that friends have no influence.

It shows that:

  1. The ordinary adjusted regression produced a large and precise positive association.
  2. The estimate became much smaller with either alternative negative-control exposure, each paired with the student’s baseline GPA as the negative-control outcome.
  3. The causal conclusion therefore depends heavily on whether the negative controls and their assumptions are credible.

The lesson is not that negative controls always solve the problem. The lesson is that an apparently convincing peer association can change substantially when hidden friendship selection is treated more seriously.

Can the Network Help Reveal Hidden Selection?

Latent-variable network models use the larger pattern of ties to estimate unobserved social positions or groups.

If students become friends partly because of an unmeasured trait such as academic orientation, that trait may leave a footprint in the wider friendship network. Researchers can estimate latent positions from the network and compare students who are similar in both observed characteristics and inferred social structure.

McFowland and Shalizi (2023) show how latent community or latent-space structure can be used to adjust peer-influence estimates under specific assumptions.

This connects the causal problem to the blockmodels and latent-factor models we studied earlier.

Latent Adjustment Can Help, but It Is Not Magic

Latent adjustment is most useful when:

  • The hidden cause of friendship leaves a strong pattern in the network.
  • The latent model recovers the part of that pattern that also affects the outcome.
  • The outcome model uses the recovered structure appropriately.
  • Uncertainty from estimating latent positions is carried into the final result.

A latent position is not a direct measurement of ambition, resources, or any other missing variable. It is a model-based summary of network structure. It can reduce hidden-homophily bias when its assumptions are plausible, but it can also improve prediction without making a causal comparison valid.

The network may contain clues about missing confounders. Whether those clues are enough is an identification assumption.

Four Difficulties Keep Returning

Difficulty Plain-Language Question
Spillovers Can another person’s or place’s assignment affect my outcome?
Exposure definition What exactly counts as being reached by the intervention?
Missing comparisons Could the people studied realistically have experienced both situations?
Hidden selection Did connected people resemble one another before any influence occurred?

The applications use different tools because they face different versions of these problems. There is no single “network causal model” that solves all four.

Design, Modeling, and Uncertainty Do Different Jobs

Research Design

Explains why the groups being compared are credible stand-ins for one another.

Statistical Model

Summarizes relationships, adjusts for measured differences, or represents a proposed data-generating process.

Uncertainty Calculation

Describes how much the estimate might vary, including dependence among connected observations.

A flexible model cannot quietly replace a missing design. A better standard error cannot remove hidden confounding. Each tool has a particular job.

Read the Claim at the Scale of the Evidence

Application Careful Conclusion
Household voting The voting message increased turnout for the contacted voter and the other voter in the household comparison
School program Invitation assignments reached uninvited students through invited friends
Ghana observers Monitoring reduced registration growth locally and coincided with greater growth nearby
Friends’ GPA The large adjusted association became much smaller under proposed corrections for hidden selection

Notice that every conclusion names the assigned action, the measured outcome, and the group being described. None replaces the measured outcome with a larger claim the study did not directly test.

Lesson 1: Ask Where the Effect Could Travel

Nickerson’s turnout study and the school program both show that an intervention can reach people who were not directly assigned to receive it.

In IR or comparative research, the same issue appears when:

  • Sanctions on one country change trade through its partners.
  • Aid in one municipality affects nearby municipalities through migration or shared services.
  • Training one military unit changes practices in units connected through personnel or command.

Before estimating anything, name the routes that could carry an effect. Then define direct exposure, indirect exposure through those routes, and little or no exposure. If we only compare treated and untreated units, we may place indirectly affected units in the control group.

The useful question is not only “Was this unit treated?” It is also “What treatment could reach this unit through its connections?”

Lesson 2: Look Beyond the Treated Place

The Ghana study shows that monitoring can move behavior nearby rather than eliminate it.

The same concern appears when:

  • Border enforcement moves crossings to another district.
  • Election monitoring moves manipulation to nearby polling places.
  • Anti-corruption audits move activity to agencies receiving less scrutiny.

Measure outcomes beyond the treated location. Compare nearby and more distant places, or several meaningful network neighborhoods. A local improvement can overstate success when the unwanted behavior simply moved.

Look beside the treated unit for displaced behavior.

Lesson 3: Ask Why Connected Units Look Alike

The GPA study shows that connected units may look alike because they selected one another or already shared important causes.

In IR or comparative research, allies and trading partners may have similar outcomes because they:

  • Face common security threats.
  • Belong to the same region or institutions.
  • Had similar political or economic trajectories before the relationship.

Use pre-treatment histories, compare prior trends, add credible negative-control outcomes or exposures, and ask whether the result survives alternative definitions of the network.

Look before the exposure for evidence that connected units were already different.

Choose the Strongest Design You Can Defend

What you actually have A sensible goal
A randomized rollout or encouragement Define direct and network exposure before assignment
A lottery, threshold, boundary, or policy change Ask whether the rule plausibly shifts exposure and whether effects cross the boundary
Longitudinal observational data Measure the network and outcomes before exposure, then probe hidden selection
One observational snapshot Describe patterns or conditional associations; do not infer influence from similarity

Start with the source of variation, not the estimator. A rollout or administrative rule may offer a useful comparison without a purpose-built experiment. Longitudinal data improve temporal ordering, but time alone does not remove confounding.

The design determines the kind of claim. The model helps analyze that design.

Build the Comparison Before Fitting the Model

  1. Name the action or exposure. What could differ across people, places, or periods?
  2. Name the outcome and timing. What is measured, and when could it respond?
  3. Draw the plausible routes. Can treatment reach friends, households, neighboring places, or shared organizations?
  4. Define the comparison groups. Who represents each exposure condition?
  5. Check support. Do the data contain enough comparable cases in every group?
  6. Write the claim in advance. What sentence would the design justify if the estimate were positive?

This workflow is useful even when the final analysis is a regression. It forces us to decide what the regression is supposed to compare and which connected units may have experienced a different version of treatment.

Make Observational Network Research More Credible

Do as Much of This as Possible

  • Measure the network, outcome, and major confounders before the proposed exposure.
  • Use repeated observations to establish sequence.
  • Compare units with similar observed histories.
  • Try negative controls, placebo outcomes, alternative exposure definitions, or sensitivity analysis.
  • Use latent network structure as a possible proxy for hidden selection, while carrying its uncertainty forward.

Do Not Treat These as Automatic Solutions

  • Actor or pair fixed effects do not remove changing unmeasured causes.
  • Lagged outcomes do not make exposure random.
  • A latent position is not a measured confounder.
  • Dyadic-robust standard errors do not repair bias.
  • A more flexible model does not create the missing comparison.

Observational tools can make a claim more credible and show how fragile it is. They cannot make the identification problem disappear.

A Useful Paper Does Not Need the Biggest Causal Claim

If causal identification is limited, the project can still make a strong contribution by:

  • Describing who is exposed through the network and how exposure varies.
  • Estimating carefully labeled conditional associations.
  • Showing whether conclusions change across plausible exposure definitions.
  • Identifying where a proposed comparison lacks support.
  • Demonstrating how much hidden selection would be needed to change the result.
  • Explaining what future data or design would distinguish the competing stories.

The goal is not to add a ritual disclaimer at the end. Let the design shape the question from the beginning. A precise descriptive or associational result is more useful than a causal claim that the data cannot defend.

Takeaways for Your Own Research

  1. What exactly changed?
  2. Who could be affected, including through connections?
  3. What two situations are being compared?
  4. What assumptions make that comparison credible?
  5. What evidence could reveal that those assumptions are wrong?
  6. What is the strongest sentence the design actually supports?

If you can answer these questions, you can design and evaluate an applied network study without reproducing the specialized estimators in the published examples. If an answer is vague, narrow the claim or improve the design before adding technical complexity.

Start with the action, the routes of exposure, and the comparison. Then choose a model that serves that design, and write the conclusion at the scale of the evidence.