Four Published Studies and the Problems They Had to Solve
This is not a training session for reproducing four expensive studies.
Use the cases to learn how to:
Most of us will not randomize a national election-monitoring program or a school-wide intervention. That is fine. The practical goal is to become better at separating a credible causal comparison from an adjusted association, and to know what would make your own study more convincing.
You do not need a perfect experiment to do useful research. You do need to be clear about what your design can and cannot establish.
Causal inference asks a simple question: what would change if we took one action rather than another? The hard part is that we never see the same person, school, or community experience both versions at the same time.
For every application, we will ask:
The goal is to recognize the problem and read the evidence carefully. You are not expected to reproduce these studies.
Imagine that we want to know whether a voting message raises turnout.
Possible Outcome 1
What this voter would do after hearing the voting message.
Possible Outcome 2
What this same voter would do after hearing the comparison message.
The causal effect is the difference between those two possible outcomes. We observe only one, so we need another group that gives us a credible picture of the missing outcome.
Random assignment is useful because it can create groups that were comparable before the intervention.
One Route
My outcome may change because I receive the intervention.
A Second Route
My outcome may change because someone connected to me receives the intervention.
In a network, an untreated person with a treated friend has not had the same experience as an untreated person with no treated friends. We need to describe both people’s own assignments and what happened around them.
Other people’s assignments can matter, and people often choose who they are connected to.
| Application | Setting | Main Difficulty |
|---|---|---|
| Nickerson (2008) | Voting messages in two-voter households | One intervention may affect two people |
| Paluck, Shepherd, and Aronow (2016) | An anti-conflict program in middle schools | Students can be reached through assigned friends |
| Ichino and Schündeln (2012) | Election observers during voter registration in Ghana | Monitoring may move behavior to nearby places |
| Egami and Tchetgen Tchetgen (2024) | Friends’ GPA and students’ later GPA | Friends may resemble one another before influence occurs |
The cases use different methods, but the same questions keep returning: what changed, who could respond, what is the comparison, and what could still make that comparison misleading?
Nickerson (2008) studied households with two registered voters in Denver and Minneapolis during the 2002 congressional primary elections.
A canvasser knocked on the door and spoke with the person who answered. Successfully contacted households heard either a message encouraging voting or a message encouraging recycling. Administrative voting records then showed whether the person who heard the script voted and whether the other registered voter in the household voted.
Action
Voting message instead of recycling message.
Outcomes
Turnout for the person contacted and turnout for the other voter.
Person Who Heard the Script
Was turnout higher under the voting message than under the recycling message?
Other Household Voter
Was turnout higher when the person at the door heard the voting message than when that person heard the recycling message?
The second comparison is the network part. The other voter did not hear the canvasser’s script, but the message may have traveled through conversation or another household interaction.
The experiment changes the message. It does not directly change whether the first person votes.
Compared with the recycling-message group, turnout was about 9.7 percentage points higher among the people who heard the voting message and about 5.9 points higher among the other voters in their households.
The study gives us evidence that assigning the voting message changed turnout for both people in a successfully contacted household.
It does not show that:
The randomized intervention was the message. The partner result is a spillover from that message. The study did not separately randomize one household member’s actual decision to vote.
Always name the action that was actually assigned and the outcome that was actually measured.
Suppose a school invites some students to join a program.
| Student | Invited? | Invited Friend? | Experience |
|---|---|---|---|
| Alex | No | No | No direct or friend exposure |
| Ben | No | Yes | Possible exposure through a friend |
| Cara | Yes | No | Direct invitation |
| Dani | Yes | Yes | Direct invitation and friend exposure |
Calling Alex and Ben both untreated throws away an important difference. Ben has an invited friend and Alex does not. A network study must decide whether that difference matters for the question.
An exposure condition summarizes the parts of the assignment that we think could matter for one person’s outcome.
Examples include:
The full assignment pattern can be enormous. The exposure condition turns it into a small set of experiences that we can compare.
The network does not choose the exposure definition for us. The intervention and the substantive story should determine it.
For a student with assigned friends, we might record:
Any Assigned Friend
Zero versus at least one.
Number Assigned
Zero, one, two, or more.
Close Friends Only
Count only stronger reported friendships.
These definitions answer different questions. “At least one” treats one assigned friend and six assigned friends as the same experience. That may be sensible for a message that spreads once, but not for a process that grows with repeated contact.
A clean calculation cannot rescue an exposure definition that does not match how the intervention could travel.
Before comparing two exposure conditions, ask:
Could the people in this analysis realistically have experienced either condition under the study design?
Examples:
If a person could never enter one side of the comparison, no statistical adjustment can create that missing possibility. Researchers may need to narrow the group they are describing.
Some people are more likely than others to end up in a particular exposure condition.
A student with six eligible friends has more ways to have an assigned friend than a student with one eligible friend. Design-based weighting gives less influence to observations that were especially likely to appear in their condition and more influence to observations that were unlikely to appear there.
The weighting tries to recover a fair comparison under the known randomization. The detailed design-based calculation is in the walkthrough, not something you need to reproduce here.
Paluck, Shepherd, and Aronow (2016) studied 56 public middle schools in New Jersey during the 2012 to 2013 school year. The larger project involved more than 24,000 students between ages 11 and 15.
Half of the schools were assigned to host the Roots program. In those schools, some eligible students were randomly invited to become “seed” students. The students met regularly to identify common sources of conflict in their own schools, then designed slogans, posters, public events, and other activities meant to make handling conflict constructively more visible.
The broader study reported fewer disciplinary incidents involving peer conflict in program schools. Our network question is how the program reached students through their friendships.
The orange wristbands, printed with the Roots tree logo, were a visible way to endorse the program’s anti-conflict message. Seed students handed them out for friendly or conflict-reducing behavior. Wearing one therefore gave researchers a concrete sign that the program’s message had reached a student, although it was not itself a measure of whether conflict declined.
The network analysis used friendship information collected before the program to ask whether reported wristband wearing was higher among:
The wristband is useful because it makes program reach visible. It is not a direct measure of whether a student’s behavior changed or whether school conflict declined.
What Was Randomized
Whether an eligible student received an invitation to join the program.
What Happened Later
Whether an invited student accepted and participated.
About 24% of invited seed students did not accept the invitation.
The cleanest results therefore compare invitation assignments and exposure to friends’ assignments. They are not automatically the effect of actual participation, because participation was a choice made after invitation.
| Student’s Experience | Invited? | Invited Friend? | Program School? |
|---|---|---|---|
| No program exposure | No | No | No |
| Program-school context only | No | No | Yes |
| Friend invited | No | Yes | Yes |
| Student invited | Yes | No | Yes |
| Student and friend invited | Yes | Yes | Yes |
This table is the heart of the network analysis. It separates direct assignment, possible spread through friends, and the broader experience of attending a program school.
These groups are defined from randomized assignments and the friendship network measured before the program.
Compared with eligible students in control schools, reported wristband wearing was 15.4 percentage points higher among students who were not invited but had an invited friend. The largest differences appeared among students who were invited themselves.
The study supports a network-based conclusion:
Assignment to the program reached students who were not invited themselves but had an invited friend.
Keep three limits in view:
Randomization strengthens the comparison, but we still have to define exposure and interpret the measured outcome carefully.
Ichino and Schündeln (2012) studied Ghana’s 2008 voter-registration period, before a closely contested national election. The number of newly registered voters was approaching 2 million, far above the roughly 800,000 people expected to have become newly eligible by age. Some of the gap had innocent explanations, but it also intensified accusations that parties were moving supporters or registering people improperly.
Why do we care? Registration lists determine who can cast a ballot. If monitoring only moves questionable registration down the road, looking only at the monitored center would make the program appear more successful than it was.
This was a national election. The intervention took place at local registration centers connected by geography and political strategy.
Ghana’s Coalition of Domestic Election Observers, or CODEO, sent observers to registration centers during the 13-day registration period. The observers could discourage questionable registration where they were stationed. But political organizers could respond strategically by shifting registration efforts to nearby centers without observers.
The study examined growth in the number of registered voters from 2004 to 2008 at:
Deterrence
Registration growth may fall where an observer is present.
Displacement
Registration growth may rise at nearby unmonitored centers if activity moves away from monitoring.
Nothing here requires friends persuading one another. The connection is geographic: monitoring in one place may change behavior in another place. Network and spillover ideas apply whenever units are connected through a route that can carry a response.
Spillovers can involve information, strategy, movement, competition, or shared resources.
In the fitted comparisons, registration growth was about 3.5 percentage points lower at an observer-assigned center. For an unmonitored center, one additional observer-assigned center within 5 kilometers was associated with about 2.7 points more registration growth.
The local decline and nearby increase are consistent with deterrence and displacement.
They do not directly show:
The experiment gives strong evidence about how observer assignment changed registration growth in the studied areas. Calling all excess growth fraudulent would go beyond the outcome that was measured.
An intervention can move an outcome across connected units instead of simply reducing it.
Egami and Tchetgen Tchetgen (2024) study a familiar observational question using Add Health, a large US study that surveyed adolescents in grades 7 through 12 in 1994 and 1995. Students named up to five male and five female friends, and the researchers linked those friendship nominations to later academic outcomes.
Would a student’s later GPA change if their friends had a higher average GPA at baseline?
The application follows more than 10,000 students. GPA summarizes reported grades in English, mathematics, history or social studies, and science. The appeal of the question is clear: friends study together, exchange information, shape expectations, and spend time in the same settings. But students do not receive friends through random assignment. They choose friends, share courses, and often enter friendships because they already have similar backgrounds and interests.
Suppose Maya and Jordan are friends and both earn high grades.
Influence Story
Jordan’s study habits help Maya improve.
Selection Story
Maya and Jordan became friends because both were already academically motivated.
Both stories can produce the same observed pattern: friends with similar grades. A regression with many measured controls can still confuse the stories if motivation, family resources, course placement, or other shared causes were not measured well.
This is hidden homophily: an unmeasured reason for friendship also helps explain the outcome.
A basic placebo check asks whether a variable with no plausible effect still predicts the outcome. Friends’ headaches provide that intuition here. Egami and Tchetgen Tchetgen (2024) go farther by pairing an exposure control with an outcome control.
| Role | Application variable |
|---|---|
| Main exposure | Friends’ baseline GPA |
| Outcome | Student’s later GPA |
| Negative-control exposure | Friends’ headaches or peers-of-peers’ GPA |
| Negative-control outcome | Student’s own baseline GPA |
The two controls contain indirect information about hidden friendship selection and shared context. A specialized bridge estimated with generalized method of moments uses that information to revise the peer-effect estimate. This is not the same as simply adding headaches to an ordinary regression.
The conventional adjusted regression estimated a 0.176-point increase in later GPA for a one-point increase in friends’ average baseline GPA. The two negative-control estimates were 0.033 and 0.078, and both were too uncertain to separate from zero.
The result does not prove that friends have no influence.
It shows that:
The lesson is not that negative controls always solve the problem. The lesson is that an apparently convincing peer association can change substantially when hidden friendship selection is treated more seriously.
Latent-variable network models use the larger pattern of ties to estimate unobserved social positions or groups.
If students become friends partly because of an unmeasured trait such as academic orientation, that trait may leave a footprint in the wider friendship network. Researchers can estimate latent positions from the network and compare students who are similar in both observed characteristics and inferred social structure.
McFowland and Shalizi (2023) show how latent community or latent-space structure can be used to adjust peer-influence estimates under specific assumptions.
This connects the causal problem to the blockmodels and latent-factor models we studied earlier.
Latent adjustment is most useful when:
A latent position is not a direct measurement of ambition, resources, or any other missing variable. It is a model-based summary of network structure. It can reduce hidden-homophily bias when its assumptions are plausible, but it can also improve prediction without making a causal comparison valid.
The network may contain clues about missing confounders. Whether those clues are enough is an identification assumption.
| Difficulty | Plain-Language Question |
|---|---|
| Spillovers | Can another person’s or place’s assignment affect my outcome? |
| Exposure definition | What exactly counts as being reached by the intervention? |
| Missing comparisons | Could the people studied realistically have experienced both situations? |
| Hidden selection | Did connected people resemble one another before any influence occurred? |
The applications use different tools because they face different versions of these problems. There is no single “network causal model” that solves all four.
Research Design
Explains why the groups being compared are credible stand-ins for one another.
Statistical Model
Summarizes relationships, adjusts for measured differences, or represents a proposed data-generating process.
Uncertainty Calculation
Describes how much the estimate might vary, including dependence among connected observations.
A flexible model cannot quietly replace a missing design. A better standard error cannot remove hidden confounding. Each tool has a particular job.
| Application | Careful Conclusion |
|---|---|
| Household voting | The voting message increased turnout for the contacted voter and the other voter in the household comparison |
| School program | Invitation assignments reached uninvited students through invited friends |
| Ghana observers | Monitoring reduced registration growth locally and coincided with greater growth nearby |
| Friends’ GPA | The large adjusted association became much smaller under proposed corrections for hidden selection |
Notice that every conclusion names the assigned action, the measured outcome, and the group being described. None replaces the measured outcome with a larger claim the study did not directly test.
Nickerson’s turnout study and the school program both show that an intervention can reach people who were not directly assigned to receive it.
In IR or comparative research, the same issue appears when:
Before estimating anything, name the routes that could carry an effect. Then define direct exposure, indirect exposure through those routes, and little or no exposure. If we only compare treated and untreated units, we may place indirectly affected units in the control group.
The useful question is not only “Was this unit treated?” It is also “What treatment could reach this unit through its connections?”
The Ghana study shows that monitoring can move behavior nearby rather than eliminate it.
The same concern appears when:
Measure outcomes beyond the treated location. Compare nearby and more distant places, or several meaningful network neighborhoods. A local improvement can overstate success when the unwanted behavior simply moved.
Look beside the treated unit for displaced behavior.
The GPA study shows that connected units may look alike because they selected one another or already shared important causes.
In IR or comparative research, allies and trading partners may have similar outcomes because they:
Use pre-treatment histories, compare prior trends, add credible negative-control outcomes or exposures, and ask whether the result survives alternative definitions of the network.
Look before the exposure for evidence that connected units were already different.
| What you actually have | A sensible goal |
|---|---|
| A randomized rollout or encouragement | Define direct and network exposure before assignment |
| A lottery, threshold, boundary, or policy change | Ask whether the rule plausibly shifts exposure and whether effects cross the boundary |
| Longitudinal observational data | Measure the network and outcomes before exposure, then probe hidden selection |
| One observational snapshot | Describe patterns or conditional associations; do not infer influence from similarity |
Start with the source of variation, not the estimator. A rollout or administrative rule may offer a useful comparison without a purpose-built experiment. Longitudinal data improve temporal ordering, but time alone does not remove confounding.
The design determines the kind of claim. The model helps analyze that design.
This workflow is useful even when the final analysis is a regression. It forces us to decide what the regression is supposed to compare and which connected units may have experienced a different version of treatment.
Do as Much of This as Possible
Do Not Treat These as Automatic Solutions
Observational tools can make a claim more credible and show how fragile it is. They cannot make the identification problem disappear.
If causal identification is limited, the project can still make a strong contribution by:
The goal is not to add a ritual disclaimer at the end. Let the design shape the question from the beginning. A precise descriptive or associational result is more useful than a causal claim that the data cannot defend.
If you can answer these questions, you can design and evaluate an applied network study without reproducing the specialized estimators in the published examples. If an answer is vague, narrow the claim or improve the design before adding technical complexity.
Start with the action, the routes of exposure, and the comparison. Then choose a model that serves that design, and write the conclusion at the scale of the evidence.