Showing posts with label Rogoff. Show all posts
Showing posts with label Rogoff. Show all posts

Thursday, October 24, 2013

Why a lot of published scientific research could be wrong

We've talked about this before, but this infographic presents it pretty cleverly.

----

First I would love to see a source on where he gets his numbers of false positives and false negatives.

He also oversimplifies the way experiments work.  For example with the higgs boson.  The experiments were run continuously until statistical significance  was seen.  That means the hypothesis was testing hundreds and hundreds of times.  The false negative/positive issue is resolved by virtue of the experimental setup.  

And on top of that, published work is not set on a pedestal because it is published.  Once it is published there is generally rigorous work to review and replicate the findings.  
So in short, he is wrong, and when he is right it doesn't matter because that is part of the process.  

-----

Those #s were just hypothetical based on a probability distrib, just to make his point. But sadly few scientists even consider type 1 and 2 errors, and just project confidence that their findings are bulletproof. In bio, when ppl try to reproduce expts, they can't about half the time, according to Science mag.
Yeah I think they are referring to distinct hypotheses, so a series of expts relating to the same hypoth. would probably be 1 data point only. There is also a risk associated with repeated "testing until significance", cuz the more measurements you make, the higher the chance you'll get some random signal and mistakenly use it to prove your hypoth. 

Actually you would be surprised how little rigor goes into the review process, I am sure S and A could comment. If the author is a big name with political clout, free pass. No one has the time or resources to verify and replicate work, it is the honor system. If another group wants to build on the research and don't get the same results, they may publish that, but may not to avoid controversy and be embarrassed if others don't buy it. 

But considering the incentive structure and human nature and the higher stakes of published results, it's a bad combo. Also the journal wants to publish big sexy results to make more money so there is COI.
----
re: "mistakenly use it to prove your hypoth" I think you don't understand what I'm describing.  The measurements means more averaging as I'm describing it.  And more averaging means LESS chance of random signal causing an error.  It is just a method for eliminating "noise" whatever that noise may be.  For the higgs boson it was a very specific energy level measured from the decay.  Every time they took another measurement, if the higgs existed the energy signature would be reinforced and other sources of energy negated, or the opposite if it did not exist.  And with some fancy math they get a confidence level based on number of trials and blah blah.  So not all experiments are susceptible in the same way.
Re COI and big names.  Yea.  Reinheart and Roghoff (sp?) is an example where more weight was given to big names.  But it is also an example of the fact that people do look at this stuff.  Though you could argue the damage was done by the time we figured out their conclusions weren't great.

In any case I'm not arguing that false positivies don't exist in academic journals.  I am arguing that this particular guy's math that proves it is false at worst or misleading at best.
----
I see, thanks M. In that case, the higgs boson expt was doing its own internal QA and replicating the experiment many times (something very lacking in bio and business, where many expts are underpowered). So if the variance is small and there is a strong signal from most of the trials, then that is a very convincing result.
At work we were struggling with a related issue. We usually don't approve changes to the website unless they pass a lot of statistical and sanity checks. But due to probability we know that we are likely rejecting some features that were truly positive (but we couldn't detect it, by chance or by poor design), and we are approving some features that are in fact neutral or even negative. That is the scary possibility. Sometime our approval criterion is "do no harm", as in if all the key metrics look to be within noise, and we can't reject the null hypothesis, then it's OK. But there is a slim chance that we are approving stuff that is actually very harmful. And if the company is approving 100's of features a year, it's likely some of them are harmful but mislabeled as positive or neutral. Though for mature businesses, where you are squeezing out basis points on the margin, the risk is probably not terrible. I guess there is a balancing act between scientific rigor and business expediency/strategy. But at least in e-commerce, an error isn't likely going to cost lives (unless it's Apple maps LOL). But for consumables, finance, or public works, it could be really bad. And they say we don't need regulation? :)
----
Yeah, it's too bad that knowledge of this sort of thing seems to only reinforce the worst of carpetbagging instincts: get in early, publish as much as you can whether of integrity or not, and then do everything in your power to ensure that no one who comes after you could possibly do it as its supposed to be done in the first place.

Maybe it's a cultural thing - like, perpetuating a hierarchical ponzi scheme requires the smartest of people to buy in to the stupidest of ideas (like blinders, narrowness, and petty infighting)...

PS:  Technology will (ultimately) solve all these problems as democratized access to information will lead to market forces being brought to bear on the most indefensible of regressive attitudes and mentalities (most of which are sustained on the basis of petty economics).
----
We were talking about this, and here were our general follow-up thoughts:

- The sci. method and the peer-review process are hundreds of years old and developed during times when methods were not so empirical & specialized, and there wasn't such $$ at stake for discoveries (for people like Euler and Carnot, I think reputation, knowledge, and prestige might have trumped the incentive to cheat, be negligent, and cut corners)

- Obviously we live in a new age now, and as C said the democratizing effects of technology could mitigate the problem (i.e. a grad student found Rogoff's errors when big time editorial staff and top peers didn't), but there are limits to that in hard science because of the cost and specialization of certain expts (though if you need a certain highly skilled postdoc to do a test a particular nuanced way to get a positive result, probably your result is not that robust)

- At least for bio, we thought to develop an independent, confidential auditing lab within the NIH to verify all published results. Journals and authors have to pay a fee to support it, and an article that passes the audit gets a certification that gives it more clout. The author's lab has to send all the materials over to the NIH (or let them use their equipment), with instructions, and the expert NIH staff have to reproduce the result within reasonable variance (they sign NDAs and no-competes so the author has no fear of being scooped). If they can't reproduce, then the paper can still be published at the journal's discretion, but without certification.
----
I would think an algorithm could be developed (or maybe already has?) that can blindly take in the statistical data from an experiment and verify the conclusions.  Rogoff was a case of bad math not bad data right?  That removes the requirement to reproduce experiments and leaves you with only the first two cases of lies, damned lies, and statistics.
The NIH thing is interesting but there is still money involved which is always problematic.  The NIH is not free from politics since that is its major source of funding.  And there are many many public institutions that don't share the data found through public dollar funded experiments already.  The democratization of technology doesn't help when all the research is behind a pay wall.
----
My comments are enclosed (in-line)...  Also, in case anyone didn't see the full article:  http://www.economist.com/news/briefing/21588057-scientists-think-science-self-correcting-alarming-degree-it-not-trouble

I would think an algorithm could be developed (or maybe already has?) that can blindly take in the statistical data from an experiment and verify the conclusions.  Rogoff was a case of bad math not bad data right?  That removes the requirement to reproduce experiments and leaves you with only the first two cases of lies, damned lies, and statistics. 

It's a good idea; one would think some of that would be built-in to the tools the scientists and researchers are using but clearly there is progress to be made...
The NIH thing is interesting but there is still money involved which is always problematic.  The NIH is not free from politics since that is its major source of funding.  
True... but federal institutions tend to be bulwarks of trust (certainly to a greater extent than a tobacco company or The Koch Brothers)...
And there are many many public institutions that don't share the data found through public dollar funded experiments already.  

This is also true... but changing.  Even institutions which are known for being benefactors of the public (e.g. U.C. Berkeley) are getting more formal about sharing research and ensuring access... so the change is bound to ripple through to other public institutions too...
The democratization of technology doesn't help when all the research is behind a pay wall.

Yeah, that is slow to change - but perhaps the most hopeful front.  The economics of information mean is such that librarians and libraries are being confronted by these questions so the clock is already ticking: if our engineering library at U.C. Berkeley is getting rid of its books for the value the physical space has to the college then its probably a matter of time until someone rationalizes our giving research to private journals so that they can rip off the campus with subscription fees that (over time) do not seem to be worth more than the promise of a new student.
-----
For Rogoff, I think it was a variety of issues, but some of it was excel formulas pointing to the wrong cells, as well as "improper" suppression of some data points. So that is mostly human error and poor judgment, which is hard for an algo to rectify.

An algo would be cool (like how the IRS algos randomly scan over people's returns and flag errors/possible issues), but then the scientific data has to be formatted and structured according to some standards so the algo can read it properly. And since each article kind of measures different things and employs different significance tests, it could be tricky. Hell, they can't even get healthcare.gov right using 50 contractors (maybe that is their problem!).
There are some free or lower cost internet journals out there... hopefully they can give the heavies a run for their money some day. But due to the prestige factor, no big names want to "slum it" with online journals. But I agree that those journals are ripping off institutions to give then "novel research" that is at least 30% useless and another 30% incorrect.

Jokes aside, we piss and moan about how bad public institutions are (and they are of course not free of corruption either... see the MMS), but as C said, they are our last resort against a totally for-profit world. We need to strengthen the public institutions so that they offer a legit alternative to the for-profits, and then with customer choice it will compel the for-profits to clean up and stop shafting us so much. I guess that's why the health industrial complex fought so hard against single payer, which is BY FAR the best feasible health system in the Western world, warts and all. But with all the dysfunction in Washington, furloughs, and a reduction in public worker compensation/respect, it makes it less likely that our best and brightest will want to go into public service and stay there long enough to make an impact. 

Thursday, April 25, 2013

Elite Harvard debt-hawk econ profs sunk by a public school grad student down the street

A seminal academic paper by Reinhardt & Rogoff has been used by conservative central bankers, politicians, and pundits from here to the EU to justify austerity cuts (because their analysis showed that higher sovereign debt levels result in low or even negative GDP growth). So we have to cut in order to have growth. But it turns out that they were wrong (either deliberately or not).
A UMass Amherst econ grad student was trying to reproduce their results, and R&R were kind enough to give him their original spreadsheet. Well it turned out that there were "coding errors, selective exclusion of available data, and unconventional weighting of summary statistics", and after those were rectified, then their original conclusions were invalidated. Now it looks like countries with even >90% debt/GDP can still have 2.2% GDP growth (wouldn't we love to have that much growth, and our carried debt is about 100% GDP). 
So just like trickle-down "voodoo" supply-side Reaganomics, and other BS conservative theories of how they want the real world to behave, this further shows that the "science" behind GOP economics is about as scientific as Scientology. It's just a shame that millions of people have lost their jobs and suffered in other ways due to R&R's errors, and I wish the "expert" peer-review community would have caught it sooner (especially since so much consequential policy was based on a single source).

---------

Yeah, this whole episode has been pretty illuminating.

The paper was pretty clearly shenanigans from day one. The question is the link between slower growth and higher debt. It's fairly obvious that slower growth can cause higher debt: if you make less money than you expect, your debt will generally be higher. R-R were attempting to prove causality in the opposite direction, that having more debt causes slower growth. That would be an interesting result. But you can't just demonstrate it by showing there's a correlation between debt and growth, because correlation doesn't tell us in which direction the causation runs, and we've got a plausible theoretical story for why it should run from slow growth to higher debt. So even if their math were tip top, they still wouldn't have proven what the austerians wanted them to have proven.

The so-called "coding error" actually isn't a big factor. Of the overall error, from R-R's -0.1 to the corrected 2.2, roughly 0.1 of that is the "coding error." The rest is their selective picking of the data and their weird weightings (they took the average growth rate for each "episode" and averaged those all together, ignoring duration, so a 10-year span of 2% growth in one country and a 1-year span of -4% growth in another country averaged to -1%). But the "coding error" is so easy to explain and so asinine that it makes for great TV. Honestly, if they hadn't made that error, this probably wouldn't have been nearly as big a story, even though that was a tiny error, and it's clearly an honest mistake where the others smack of cherry-picking your data and methodology to fit pre-selected conclusions.

Also, I think it's sad that this gets called a "coding error." The issue is that they had an Excel formula which should have been "AVG(E30:E49)", and instead they put "AVG(E30:E44)" (44 instead of 49). That's not "coding." It's data entry. A "formula error" at best. But it's not math or computer science that we're talking about here. They just typed the wrong number into the cell.

------

I didn't read the original R&R paper and haven't followed their story, but as you said - it seems pretty ridiculous that they could make a counter-intuitive causation argument while only armed with heavily massaged archival data that happened to show a correlation. Yeah, I've seen plenty of that data cherry-picking and massaging until the result is pleasing (I have unfortunately participated in it too).

LOL, "coding" sounds cooler than "typo", and technically Excel is a programming language - albeit a very graphical, limited one. :) Most Americans can't perform a square root without Google. Frankly I have rarely seen old folks (and esp. profs) who are competent in Excel - I wonder if R&R actually made the goof themselves, or one of their student slaves instead and it wasn't detected?

I guess the lesson is: don't base sweeping policy on a single controversial source, even if it's from famous authors. Try to get a 2nd or 3rd validation, and even better -  research all the counter-arguments to see if they have merit. But I guess that would be too rigorous and "fair and balanced" for ideologues.

Well another lesson is - NEVER give out your data analysis files except under subpoena! And then you might want to wipe your hard drive first and say it was "user error". :)

---------

Agreed. I think for all peer-reviewed academic publications (and for-the-public gov't studies too), all the raw data and analyses should be made available. Let enthusiasts pour over it at their leisure, and if the authors did their work properly, they should have nothing to hide. To err is human, to leave errors undetected to cause harm day after day is American.
But I think the debt-GDP growth connection is clearly not open-and-shut (even when that paper was published, otherwise all nations should have adopted austerity), and case-specific factors can affect things. Japan has been at or near the top in terms of debt/GDP for some time. They also had a lost generation and a decade-long "recession", arguably because the gov't didn't spend ENOUGH on Keynesian stimuli. But Japan's sovereign debt was mostly internal (Japan owed its people, not China like us) and at very low interest rates, so the "burden of debt" was not as risky and crippling as say Greece, where Germany is charging blood money rates (maybe for good reason since they are a higher default risk, but I think their repayment terms are more punitive than prudent).

I guess using a company as an analogy, leveraging to the hilt is not a problem as long as you use the money on "smart" projects that give you returns well in excess of your borrowing costs. But as we saw, leverage can blow up at times when interest rates go up or income declines (then you need to take out new, worse loans to pay off your old loans that are coming due). But that is the case that J described: lower growth/revenue leads to more debt, not the other way around.

---------

First, china owns a very small amount of our debt.  Strong majority is govt or us public owned.
Also, low revenue causes debt just as much as it causes austerity.  Lower revenue means you either spend less to compensate or dont but it need not cause debt.

Of course the reality is the us is functionally incapable of significantly reducing spending so sort of a moot point.

---------

Well if you look at new borrowing over the last 10 years, I think China and the Middle East hold more share than domestic buyers (recession caused flight to perceived safety and higher demand for US Treasuries, which drove down yields). But overall, yes the biggest investor of US debt is Social Security.
I agree, if we lack the political will to cut spending (meaning the big drivers of spending, like Medicare and defense), it is a challenge - but then there is always the revenue side of things. Though there seems to be just as little will to tackle tax reforms too.

I forget which journalist/historian it was, but he was saying that Athens and Rome's declines exhibited some of the same features: heavy military spending with little strategic gains, political corruption and gridlock, and social apathy to hold leaders to account.
It might have been this guy: http://www.kqed.org/a/forum/R201304221000



----------