The India That Never Was, and How to Argue About It
A working paper says India underperformed after 2014. Before you use it or attack it, understand what synthetic control can and cannot establish.
GOVERNANCE MEASUREMENTDATA AND ELECTIONSV-DEMSYNTHETIC CONTROLCOUNTERFACTUAL ANALYSISPOLITICAL METHODOLOGYPOLITICAL COMMUNICATIONECONOMYBJPACCOUNTABILITYELECTIONSPOLITICSPOLITICSPOLITICAL ECONOMYDEMOCRACYPOLITICAL STRATEGYSOCIO-ECONOMICSINDIAGOVERNANCENARRATIVE


On Synthetic India, and the instrument nobody explained
A working paper appeared on SSRN on 12 August. Two economists from Texas Tech, Kevin Grier and Robin Grier, built a statistical stand-in for India, compared it to the country's actual performance, and found that India performed worse on 10 governance measures and on per capita income after 2014. A national news website covered the story on 25 August. Kanchan Gupta, Senior Adviser at the Information and Broadcasting Ministry, responded on 26 August. The authors replied on 28 August. Both sides said they were satisfied.
It is worth noting that the paper has not been peer-reviewed. It is a preprint on a working paper server, not a published article, so it should be seen as a work in progress rather than a final conclusion.
During two weeks of public debate, hardly anyone explained what the method actually does.
This is a gap that needs to be filled. The reason is not whether the finding is right or wrong, but because this method will be used in Indian politics many times over the next three years. People affected by it should understand it before it is used, not after. The following is an attempt to explain synthetic control to a practitioner, using this study as an example.
The problem the method solves
Every political argument about performance runs into the same wall. You cannot re-run the decade.
Governments highlight what has improved, while the opposition points out what has not. Both sides talk about levels, but levels are not very useful evidence because things usually improve over time in a growing economy. Roads are built, phones become cheaper, and poverty goes down. The real question is whether these things improved more than they would have otherwise. We cannot observe this directly because we have no version of India without the last twelve years to compare against.
Synthetic control is one way to answer this problem. The idea, developed by Alberto Abadie and his team and first used to study California's tobacco control programme, is to build the missing comparison using data from other countries. You start with a group of countries that are similar to the one you want to study. An algorithm finds a weighted mix of these countries whose combined history matches your country's record before the event you are interested in. This mix acts as the stand-in. After the event, you see where the real country and the stand-in diverge, and that difference is called the effect.
Keep two important things in mind when interpreting this method.
First, the weights are always non-negative and sum to 1. The stand-in is always a mix of real countries and cannot go beyond their range. If the real alternative history would have been better than any country in the pool, this method cannot show it.
Second, the model is designed to be sparse. Most countries in the pool get a weight of zero, and only a few carry most of the weight. This is intentional, not a mistake, and it is the main reason for much of the public debate about this paper.
What this study did
The authors used two different sets of models because India is unique. It has stronger political institutions than most countries with similar income, so countries that are good economic comparisons are not good political comparisons.
For income, the pool included 14 large developing economies. The resulting mix was 38 per cent Ethiopia, 28 per cent China, 25 per cent Bangladesh, 7 per cent Pakistan, and 2 per cent Philippines. For governance, the pool included 12 developing democracies, with the largest weights assigned to Argentina, Malaysia, Sri Lanka, and Chile.
The pre-treatment period covers 30 years, from 1984 to 2013. For income, the model fits very closely, with an average prediction error of about $59 for an income level above $5,000.
Here are the results in brief. The study used 10 governance indicators from the V-Dem database, each over 10 years, giving 100 yearly estimates. Ninety-six of these show India falling below its stand-in. For income, the gap grows over the decade to about $1,000 per person, or roughly 10 per cent.
What the method can establish
It can detect a break. If two series track each other for 30 years and then separate sharply at a specific date, something happened at that date.
It can also rank results. The method does this by pretending to assign the treatment to each donor country in turn, creating a set of fake effects, and then seeing where the real effect falls among them.
It can also check if the model breaks down. The authors tested this by assuming the treatment started in 2001 rather than 2014. The result showed no effect. This is important because a common criticism of long-term models is that they might not work outside the sample. In this case, the model held up.
The paper's strongest point is something most commentators missed. The 10 governance models are not just the same result repeated. Each is built from a different mix of countries. For example, the corruption model relies mostly on Malaysia at about 55 per cent, while the executive-constraints models use Chile and Colombia, and the polyarchy model relies on Argentina at 61 per cent. These are very different counterfactuals, but they all point in the same direction. Getting the same result from models with very different structures is much harder to dismiss than a single finding.
This is different from the income finding, which everyone quoted. The income result comes from just one model based on a single mix of countries.
What it cannot
It cannot show what caused the gap. The method only measures the difference. It does not explain why the gap exists, and the paper's idea that institutional decline caused the income shortfall is just an interpretation, not a direct finding.
It cannot separate the events that occurred after 2014. Everything such as demonetisation, the goods and services tax, the pandemic, the lockdown response, global interest rate changes, and the Ukraine shock is grouped together into one switch. The authors say domestic policy choices should be included in the treatment, and that global shocks affecting all donor countries cancel out. Both points are partly correct, but neither tells the whole story.
It also cannot show an alternative history that goes beyond its pool of countries. If the real counterfactual would have been better than any donor country achieved, the method cannot capture that.
Finally, it cannot turn a ranking into a probability.
Read the p-values properly
With 13 usable donor countries, the possible values are approximately 0, 0.08, 0.15, and so on. Here, a p-value of 0.00 does not mean the same thing as in a regression. It just means India was the most extreme of the 13. A p-value of 0.08 indicates it was the second-most extreme.
For governance, 92 out of 100 estimates are 0.00, indicating a strong and consistent pattern. For income, seven out of 10 years are at 0.08, meaning India is second, not first.
The authors themselves make this clear. Their main point is that the government promised to do better than business as usual but did not. This is a weaker claim than what the headlines suggest, but it is honest. Anyone who repeats the income result as highly significant is treating a rank as if it were a probability.
Check the fit before you quote the number
The biggest effect in the paper is for freedom of religion, which is reported as a 168.7 per cent drop compared to the 2013 level. This number is likely to be quoted the most because it is the largest.
However, this result comes from the weakest model in the paper.
The authors share their pre-treatment prediction errors, and most are very low. Judicial constraints: 0.004. Legislative constraints and equality before the law: 0.008. Equal protection: 0.011. Polyarchy and corruption: 0.017. For freedom of religion, the error is 0.120, which is seven times worse than the typical model and thirty times worse than the best one. Their own table clearly shows this mismatch, with India at 0.652 in 2001 and the stand-in at 0.802.
This does not mean the result is incorrect. The gap after 2014 is much bigger than the pre-treatment error, so there is still something to notice. However, it means the study's biggest headline is based on its weakest model. Careful readers should instead quote the freedom of expression or polyarchy results, where the model fits better, and the effect is almost as large.
To their credit, the authors published the number that reveals this weakness. Not every paper does that.
The Ethiopia problem
This is where the method and the argument diverge, and the lesson here matters more than the paper itself.
Methodologically, the donor weights are defensible. Abadie has written explicitly that stuffing a donor pool with unsuitable units invites bias rather than reducing it. Sparse weights are the design working as intended. The authors also report that dropping Ethiopia considerably worsens the fit and makes their estimated effects larger, not smaller. On the technical merits, they have an answer.
From a communication standpoint, they have no answer. The line, "Your India without Modi is thirty-eight per cent Ethiopia and twenty-five per cent Bangladesh," is the kind of statement that ends a TV segment. It takes no statistical training to deliver and about four minutes to answer. Every critic will spot that line quickly, and the technical reply will never catch up.
For people in political communication, this gap is the real lesson. A research finding is not the same as a message. A defensible method that cannot be defended in one sentence will lose the public argument, even if it is correct.
The missing five
An even stronger criticism has not yet been raised.
The biggest criticism of any synthetic control study is that the pre-treatment trend might not have continued. In 2013, India was not on a smooth path. It was part of the Fragile Five, a group Morgan Stanley identified in 2013 as most at risk from a reversal of capital flows. India faced high inflation, a large current account deficit, and hidden bad loans. If 2014 was a turning point rather than a stable period, then the projected alternative history is showing something that would not have happened.
The obvious defence is that this stress was not unique to India and that the donor pool already includes it. That defence is available, because the other four members of the Fragile Five, Brazil, Indonesia, South Africa and Turkey, are all in the paper's economic donor pool.
However, all four of these countries receive a weight of zero.
The income counterfactual is based on Ethiopia, China, Bangladesh, Pakistan, and the Philippines. None of the countries that shared India's specific 2013 problems contributes to the projected path. The algorithm did not exclude them on purpose. It just found that other countries' 30-year income trends matched India's more closely, as expected. As a result, the comparison group did not go through the same stress at the start of the projection.
You can check this directly in the paper's own table. It is the strongest argument against the income result, more specific than anything discussed publicly so far, and it can be tested. You could weight the stressed countries more heavily, or add 2013 macroeconomic indicators as matching factors, to see whether the gap remains. Until someone does this, the objection is just an argument, not a finding.
The exchange, read for method rather than politics
The ministry's response made four main points. It said the donor pool was cherry-picked and asked for a complete one. It claimed the result changes with other datasets. It criticised V-Dem as subjective and unclear. Finally, it listed a decade of achievements, such as Jan Dhan accounts, digital payments, highways, reserves, and reduced extreme poverty.
If we judge only by the method, three out of the four points do not hold up.
The call for a complete donor pool misunderstands how the method works. It applies the regression logic, where more controls are generally better, to a matching method in which including too many countries can actually introduce bias.
The criticism of V-Dem is the weakest, because the authors have a clear answer. All donor countries are scored using the same process, so any global changes in expert opinion cancel out. There is a real academic debate exists about how to measure democratic backsliding, with research on both sides. But calling an index unclear, even when it publishes detailed data and uncertainty estimates, is not part of that debate.
The list of achievements is an argument about levels, used to answer a counterfactual claim. No one denies that Indian incomes increased or that poverty went down. Some of these numbers also have their own caveats, since the World Bank's poverty comparison starts from 2011-12 and includes changes to both the threshold and the survey method.
But there is another way to look at that list, and it should be mentioned. Mr Gupta may not be giving a bad answer to the counterfactual question. He may be rejecting the question itself. In this view, a counterfactual is not just uncertain; it is unknowable, so the actual record is the only evidence we have. This is a reasonable position, held by serious people, and it should not be dismissed as a mistake. My own view is that if we accept this, we lose the ability to judge any government, since every leader in a growing economy can make a list. The right answer to a weak counterfactual is a better one, not to stop asking the question.
What a counterfactual cannot tell you
There is a limit here that is not about statistics.
A gap compared to a model is not the same as real experience. No voter in Rae Bareli or Ratnagiri lived through a 10 per cent shortfall against a mix of Ethiopia, China, and Bangladesh. They experienced changes in prices, jobs, whether a road was built, whether a bank account was working, or whether a family member found work elsewhere. Elections are decided by these lived experiences, compared to memory and to neighbours, not to a synthetic model.
I do not think any study like this, whether for or against, has much influence in a political debate. Voters do not make decisions based on reports, whether they are factual or counterfactual. This works both ways. It is why this paper will not change any election results, and why a long list of achievements does not prove whether the decade was well spent.
What these studies really do is shape how writers, advisers, and teachers discuss the record. This is a smaller reward than winning an election, but it lasts longer.
Why this matters past this one paper
The method can be used elsewhere, and it is actually stronger at the state level.
The main weakness of this paper is that it groups many events together. Demonetisation, GST, and the pandemic are all included in the same category as everything else. If you use the same method across Indian states, these national events affect both the treated state and the donor states, so their effects cancel out. This gives a clearer result than what is possible at the country level, using state income data and surveys that are already available.
The work itself is not difficult. The hardest part is assembling the data, since long state income series undergo several base-year changes and must be combined carefully. But someone will do this, and the best time to do so is before the next round of assembly elections.
Any Chief Minister who has served a long time should expect this method to test their record before they face voters. The best response is not to wait and then claim the result is unfair. Instead, they should learn how the method works now, know its real weaknesses, and be ready with a clear answer.
A counterfactual is just a model, not a memory. It deserves neither too much reverence nor outright dismissal. The right response is another model, run openly, with the data and code published for anyone to check. Three tests would address most of the debates here: force the Fragile Five countries into the income mix, rerun the income model using World Bank and Maddison data, and publish the freedom-of-religion model with a fit good enough to support its result.
None of this needs a press conference. It just takes a weekend and a willingness to publish whatever comes out.
Sources
Grier and Grier, Promises, Promises: Governance and Growth in India under Modi and the BJP, SSRN, August 2026: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7248338
Authors' data companion: https://rgrier88.github.io/modi-promises/
The Wire, rejoinder by Kanchan Gupta and the authors' reply: https://thewire.in/political-economy/debate-modi-govt-official-questions-methodology-data-of-broken-promises-study-authors-respond
Abadie, Using Synthetic Controls: Feasibility, Data Requirements, and Methodological Aspects, Journal of Economic Literature, 2021: https://www.aeaweb.org/articles?id=10.1257/jel.20191450
