Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Friday, March 1, 2013

Lead in chicken broth?

After the lead and crime article in Mother Jones last month, I've been more aware of lead in the environment.  Perhaps that is why a paper finding lead in organic chicken stock in the journal Medical Hypotheses alarms me, at least slightly and hypothetically.  The researchers found concentrations of 9.5 mcg/L and 7 mcg/L, still below the EPA's "action level" of 15 mcg/L for drinking water, but the EPA in the same publication says that no concentration of lead is acceptable:  i.e., the goal level concentration is zero. 

As is the nature for a hypothesis-generating publication, the study used only one batch of chicken stock of each type:  one with bones, one with skin and cartilage, and one with meat.  Perhaps they picked a bad chicken, or perhaps it is only chicken from the UK.  It's impossible to know without more tests.  People always ask us statisticians what would have happened if the sample size would have been larger, and it's important to remember that we're statisticians, not fortune-tellers:  we can't know what results would have been without having the actual results.  It's strange and frustrating that with 3 authors on the paper, they couldn't be bothered to get a few more chickens to test.  At least they did a control of plain water boiled for the same period of time, which found near-zero concentrations, under 1 mcg/L.  (Chicken meat broth by contrast was just over 2 mcg/L.)

The authors neglected to mention that the technique of saving bones for stock from every piece of meat you eat is almost universally recommended by cookbook authors including Melissa Clark and Nigella Lawson, to ensure that you always have stock available when it's called for.  Cookbooks say that there is no substitute for stock, and if you don't have stock around, use water rather than canned.  It's disappointing that there might be dangers from what is considered a best practice for cooking, not to mention one that I've just started following myself!  The lentil soup made from homemade beef stock was amazing, and one guest told us that he'd never tasted anything like it. 

I will still use the chicken stock in my freezer, but I'd be reluctant to continue this habit if pregnant or if I had children, given how dangerous lead is during critical formative periods. I do wish that the authors had gotten their act together to get at least an n=10 from geographically dispersed chickens.  How hard would that have been, seriously? 

Friday, September 21, 2012

Beating a dead fish: more reasons to correct for multiple comparisons

Last night, the Ignobel prizes rewarded research using an fMRI on a dead salmon as an extreme case of a null effect.  fMRI machines identify over a hundred thousand of voxels (the tiny pieces they divide the brain into), and by chance some voxels show activity even where there isn't activity simply because of the sheer number of opportunities for a false positive: similar to how a broken clock is right twice a day, and if you looked at the clock 100,000 times, you would see several times when it said the correct time.  The researchers had a dead salmon complete a standard task that would be used on live human subjects and showed that --- without correcting for multiple comparisons --- the salmon appeared to be thinking.  With correction for multiple comparisons, the scientists saw no effects.  This research was published in the Journal of Serendipitous and Unexpected Results.

Tuesday, November 29, 2011

Publication bias and the result that got away

An economist writes about his would-have-been dissertation: he gathered the data, did a regression, and found no relationship. So his dissertation was on another topic. 8 years later, he publishes the null relationship. It's refreshing to see someone talking about how publication bias affected their own research, as opposed to the studies published in epidemiology journals that suggest publication bias in macro.

Wednesday, August 31, 2011

The Chicago Tribune is unprepared for college

The Chicago Tribune published today an article with a plausible headline --- public high school students are unprepared for college --- but their logic is suspect. A state law required public universities to publish the average freshman year GPAs for students from each high school, and the Tribune has published their interpretation of the data. For each university, the Tribune plotted the high school grade point averages (GPAs) of incoming students versus their average GPA in colleges. They conclude that if students have a lower GPA in freshman year of college than high school, that implies that students were unprepared.

It seems obvious to me that college GPAs are lower than high school GPAs. About 50% of Harvard students were ranked 1 or 2 in their high school graduating class, and probably over 80% had a GPA of 4.0, and yet college GPAs are way lower. Further, college GPAs come from a wide variety of courses, so they're even harder to interpret than high school GPAs. Engineering students have lower GPAs than humanities majors, so a really good high school that sends most of its students to major in engineering might look worse than a mediocre high school that sends most of its students to major in humanities. If college GPAs weren't lower than high school GPA, I would think that the students weren't challenging themselves enough. Finally, high schools with disadvantaged students can be expected to have lower college GPAs not because the students are poorly prepared by their high schools but because the students were disadvantaged.


The report itself is more useful, and it provides a good comparison of the students from each high school that tend to attend each school. Not a comparison of the high schools themselves, of course: enormously confounded by the which students from each school tend to attend the state universities. It does break down GPA by subject. A high school that sends its best students to the University of Illinois at Urbana-Champaign will look like a better high school who sends its best students out of state. Still interesting.

The Tribune's assumption that having a lower GPA in freshman year of college than high school implies that students weren't prepared conflates many issues in an illusion that a few numbers can summarize a high school without any statistical assistance. Is this a reminder that a weak attempt at evidence-based decision-making is worse than none at all?

Tuesday, June 28, 2011

Medical system biases

A respectably sized randomized trial finds transcendental meditation has enormous effects on heart attack mortality, decreasing by half. Among the most vulnerable or the most involved in meditation, mortality decreased by 2/3rds. That would be huge even for a drug. As one of the study authors said, "The effect is as large or larger than major categories of drug treatment for cardiovascular disease."

Nonetheless, in spite of the much larger effect from meditation than any drug therapy, the article paraphrases the researchers as saying that the meditation should complement rather than replace drug treatment. If we believe that randomized trials yield correct information that can be used for treatment, why limit the results in this way?

Yes, the results need to be replicated a few times in different populations, etc., but my suspicion is that even after they are replicated (possibly with smaller effect sizes), the "don't stop taking drugs" message will remain.

Given recent results about increased risk of type 2 diabetes from statins and indications of memory problems from statins, those who advocate drugs need to defend their choice more. It seems that there's an implicit bias that treatment within the medical system must be healthier. Similar to the implicit bias against fat that caused Ancel Keys's views to prevail; now new studies that low-fat diets contribute to unhealthy weight gain are rarely publicized.

UPDATE: Now the article has been held back from publication due to last minute data, and the Telegraph took down the article. Wonder why.

Wednesday, June 22, 2011

When statistics doesn't cooperate: ratios of regression coefficients.

Everyone knows that there's error in all statistical estimates, but we don't always think about all the implications of that error. Gelman makes the important point about ratios of (among other things) regression coefficients, with implications for instrumental variables (and a cringe-inducing published example). He notes that the ratio of two normal-like variables is Cauchy-like; my recollection from Don Rubin's causal inference class is that he criticized instrumental variables on the grounds that Cauchy distributions have infinite variance, which fills in why Gelman notes that in theory ratios of regression coefficients can take on absurdly large and useless values like 100,000.

The published example comparing regression coefficients is something that people do all the time in conversation without even thinking about it.

Gelman says that the post was really time-consuming to write, so perhaps it counts as a few posts. I agree that it should. In fact, someone could review papers in top journals in the past 5 years and document how frequently this error is made, and that would be a good project.

Sunday, June 19, 2011

Making statistics adorable: plushy edition

I got a wonderful gift of a plushie Beta distribution, definitely a way to make statistics more adorable. Very sweet, and the beta is wearing a dapper monocle.

They need comedy help in the poster department for catchier/funnier slogans. They also have a poster of relationships between statistical distributions (the third page of this document,) which could be extremely helpful but takes some cognitive effort to figure out which arrows are labeled with what.

Monday, June 6, 2011

Tables are better than plots (April Fools) and other reading


  • A several participant discussion of plots vs. tables begins with a 5 page piece by Andrew Gelman about why tables are better than plots, concluding with "I recommend using Excel, which has some really nice defaults as well as options such as those 3-D colored bar charts." April Fools! The remaining pieces knock down the straw man.
  • At Society for Research on Child Development this March, I asked NICHD head Alan Guttmacher for the current scope of the "child health and development" that his institute covers since some say that adolescence extends until 25 because the brain is not fully developed until then. He said that age 25 is a fine age to use as the end of child development. The National Campaign has issued areport on young adulthood describing this changing stage of life.
  • AP grading and what it's like to grade AP exams. I was surprised that 85-90% of these history exams were written by 9th and 10th graders.
  • Humor in romantic relationships: women have more romantic interest in online dating profiles that are funny (in their opinion) than not funny. No difference for men's assessment of women. Practical implications aside (Do any men like me for my sense of humor?), this study is a great example of how gender could be used as a counterfactual, even though they didn't use it here. Now that people are frequently represented not just by resumes (as in studies of racial discrimination in choosing interview candidates) but by electronic profiles on facebook or dating sites, we can alter gender or other immutable characteristics on the electronic profiles and infer causal effects.
  • Seating location and voting behavior on FDA advisory committees: those who speak later may have less influence on the vote, perhaps attributable to seating location.

Wednesday, May 25, 2011

Commitment devices versus treatments

Andrew Gelman has an interesting point in a blog entry about religious commitment devices, such as a WWJD bracelet. He says that the bracelet is not a treatment, but rather it goes along with a set of behaviors and commitments, and it's a sign of that. This issue of commitment devices applies as well to my virginity pledge work: like the WWJD bracelent, the pledge may not be an intervention, but rather it could be a sign that someone has a certain identity, so the pledge itself isn't important.

Conversely, Gelman's argument could be critiqued on the grounds that there are folks who have the set of attitudes and behaviors, some of whom have the bracelet and others of them who don't, and conceivably, it would be possible to balance the two groups and get a treatment effect. Whether the treatment effect of wearing a WWJD bracelet could possibly be meaningful is a good question.

More thought needed.

Friday, July 30, 2010

Could quantitative students bypass premed requirements?


Mt Sinai medical school has a program that admits non-science majors who have not taken the traditional premed curriculum, but instead have majored in the humanities and social sciences. These students do just as well, are more likely to spend time doing research, and are twice as likely to become psychiatrists than regular track students.

It's interesting that they limit the exceptions to humanities and social science majors. What about quantitative majors? Is there a med school who will take math, physics, and computer science majors who have not taken all the premed courses? Many science-smart students like learning hard material and are turned off from the grade-conscious nature of premed courses. If you are required to be grade conscious, you don't take risks by taking hard classes: the course that I learned the most from during my freshman year was a full-year math class where I got one of the bottom 2 grades in the 55 person class during the first semester and my advisor urged me to drop the class, but the second semester I was in the top third of the course. It was a trial by fire that helped me in theoretical graduate school statistics courses, and I wouldn't have taken it if I had been afraid of that first semester B-. And if I'd moved up from the 2nd percentile to "only" the 40th percentile, that would have also been a victory. The happy ending is that I learned how to write proofs, or in a pinch fake them reasonably well.

By contrast, I took a premed biology course during my first semester of college, and I loved the material. It was totally fascinating, as were the dissections (though the calves heart smelled), and I think my A- was good enough that I could have been premed if I'd wanted to. The difficult part was keeping my lunch down when my classmates continually asked what was going to be on the test. After half a dozen memory-based quizzes, I was so pleased when a quiz actually asked us to figure something out for ourselves. Finally: thinking! Some of my classmates protested this question because "it wasn't in the textbook!" Someone who spoke that way in a math or physics course would be ridiculed, and an instructor wouldn't even see the need to defend the question. Of course students are required to use critical thinking. In this premed class, memorizing the textbook seemed to be an accepted part of the culture.

From the other side of the lectern, at least in a biostatistics course, premed students were a complete joy to teach. In graduate school, I was the only TF for a 53 student biostatistics class that was 90% premed, and they were some of the nicest and most interested students that I taught. Perhaps they were fun students to teach because only the most intellectually curious premed students self-selected into biostatistics, or because all of the material was highly related to actual medicine rather than abstract science, so they saw the immediate relevance, or because of the way we taught the class with lots of handouts and everything in the textbook and being very clear about expectations.

So I don't mean to stereotype premed students. But as a math and physics undergraduate, the prospect of taking orgo and p-chem was not appealing just because of the culture that I perceived there.

A bigger argument for a quantitative track to med school is that the quantitative fields of study absorb just as much time as the intensive humanities majors: it takes (at least) as much time and coursework to learn to do proofs, derivations, program computer operating systems, and analyze complex datasets, as it does to read, write, and speak foreign languages and write a 200 page senior thesis.

Sunday, May 23, 2010

Limits of randomized experiments

I just got back from the Mid-Atlantic Causal Inference Conference, the leading meeting for statisticians who look not just for associations, but for causality. Randomized experiments are the gold standard for causality because randomization ensures that on average, the treatment and comparison groups are similar. Experiments do have limitations, however, that come primarily from their great expense: experiments may need to be small and short duration, weakening the chance that experimenters can see an effect. The study described in this article is a perfect example: 22 autistic children were randomized to a gluten-free, casein-free (GFCF) diet for 18 weeks and then given a "challenge" of these foods about 4 weeks into the trial; by the end of the trial 8 of the subjects had dropped out.

Seemingly, there are hundreds of parents on internet mailing lists and websites putting their children on a GFCF diet. GFCF diet is hard to implement, and it takes weeks or months or practice to get right, and even then an errant crumb can disrupt the progress, and it's unclear how long kids need to be on the diet to see an improvement because determining the starting point is so inexact. A parent can probably remove >90% of gluten and casein from their child's diet starting on day 1, but hunting down the remaining 10% to reach 100% adherence takes a long time. And 99.9998% adherence may be exactly what's required: the FDA definition of gluten-free is 20 ppm. Once the GFCF diet is in place, many parents say that it improves their children. Now a randomized trial that started out with 22 participants and lost 8 of them comes into the news with the headline, "Eliminating Wheat, Milk From Diet Doesn't Help Autistic Kids."

An experiment doesn't have the luxury of trying to refine the diet to make sure that it's being done correctly, or to figure out the length of time the diet needs to continue until there's improvement. An experiment generally determines the treatment in advance rather than trial and error, since trying to get the best result is, to a certain extent, cheating (i.e., risking a spuriously significant result that occurred simply by chance).

A good experiment is an invaluable tool for understanding reality, but a so-so experiment is no better than a qualitative study of people on internet websites.

Tuesday, May 18, 2010

Chemicals and comparison groups

Current legislation is trying to ban a plastic that has been used for 50 years to line cans, so I decided to look into how much evidence there is that this plastic is dangerous, and whether the potential substitutes for this plastic are safe. The American Council on Science and Health finds little evidence that this plastic is dangerous; they find no proposals for what plastics might substitute, much less any evidence on the alternatives' safety profiles. Their analysis raises very good points and is worth reading.

Just as in statistics, the important question for any risk analysis is "compared to what?" Nothing is dangerous on an absolute level: risks always have to be weighed against their alternatives. When we banned DDT decades ago, it may or may not have had beneficial effects for the eagle population, but malaria has rebounded: going from millions of cases in Sri Lanka to a couple dozen, and then back up to a million cases after the DDT ban. Malaria still affects hundreds of millions of people around the world, many cases that might be prevented if DDT spraying were allowed. The developed world has not had malaria since the 1940s --- perhaps if malaria had rebounded, perhaps we would see pesticides as the life-saving tools that they are --- but we do have the resurgence of bed bugs, even on the Upper East Side, after they had been almost completely eliminated 50 years ago. Maybe the good effects of the DDT ban are worth hundreds of millions of cases of malaria in the developing world and bedbugs in the developed world, but alternatives always need to be considered. The WHO has backed bringing back DDT because it was so useful.

With the current talk of banning BPA, the comparison group is completely missing. By banning a plastic without discussing alternatives and their risks, we risk having worse alternatives or no alternatives.

Sunday, May 2, 2010

Why doctors need to know Bayes theorem

In graduate school, I was the head TF for several general audience statistics courses, and my favorite subject was Bayes theorem because it implies that many "common sense" policies are, in fact, dangerous. Given a dreaded disease, drug use among ship captains or pilots, or anything else, it's so easy to say, "Just test everyone." But in fact, that's not good policy.

Social psychologist Gird Gigerenzer's new book covers some instances of asking doctors to give probabilities to their patients, and they do a horrible job. The question presents the information exactly as doctors are taught: prevalence, sensitivity, and specificity (false positives).

The probability that one of these women has breast cancer is 0.8 percent. If a woman has breast cancer, the probability is 90 percent that she will have a positive mammogram. If a woman does not have breast cancer, the probability is 7 percent that she will still have a positive mammogram. Imagine a woman who has a positive mammogram. What is the probability that she actually has breast cancer?


A prestigious doctor, department chief with 30 years of experience "was visibly nervous while trying to figure out what he would tell the woman. After mulling the numbers over, he finally estimated the woman’s probability of having breast cancer, given that she has a positive mammogram, to be 90 percent. Nervously, he added, ‘Oh, what nonsense. I can’t do this. You should test my daughter; she is studying medicine.’ He knew that his estimate was wrong, but he did not know how to reason better. Despite the fact that he had spent 10 minutes wringing his mind for an answer, he could not figure out how to draw a sound inference from the probabilities."

And he was typical: more than 90% of the doctors were wrong, mostly very wrong.

When the question was posed in terms that are easier for people to understand, nearly all of the doctors got the question right.

Wednesday, February 17, 2010

Positive predictive values and gun control

The US has a number of workplace and school shootings, and each incident goes the same way. Media dig up events from the perpetrator's past (putting the perpetrator's name everywhere when the perpetrator's name really should forgotten, and the victims the ones who are remembered), including criminal record, complaints from co-workers, qualitative assessments of the shooter's mental health, and random anecdotes, and all of these events are assembled to answer the question of whether someone could have known that the perpetrator was so crazy.

While some of these incidents from a perpetrator's past are indeed crazy, they have poor positive predictive value: the most recent shooter apparently punched a woman over a booster seat at an IHOP. But such ironic and strange incidents happen all the time, which I know as an avid reader of News of the Weird. Attempting to predict who will go bezerk is like looking for a needle in a haystack: many people do minor strange, crazy things, and even do so multiple times, but most people who do minor strange things will not commit homicide. Even looking at more severe incidents, such as shooting a brother at age 20, may not necessarily predict future violence. Though allegedly sending a bomb to a dissertation committee member seems more likely to predict future violence, she was cleared of that charge.

Attempting to predict who will snap begs the question. We can't. Even if we had Big Brother comprehensive databases of every person's past, and were willing to disregard rules of evidence and dropped charges, we couldn't predict severe violence. The real issue is not which people are likely to snap, but rather why our gun control policy allows people such easy access to firearms that when they do snap they can do so much damage. Human psychology is fallible, but in countries without easy access to firearms, the damage comes in the form of broken objects, bruises and broken bones, and perhaps even a stabbing. Firearms create more damage, and damage that is most likely to be deadly.

. . .

Speaking of News of the Weird, here's one from last year about a Yale PhD and professor:

Love Can Mess You Up: Before Arthur David Horn met his future bride Lynette (a "metaphysical healer") in 1988, he was a tenured professor at Colorado State, with a Ph.D. in anthropology from Yale, teaching a mainstream course in human evolution. With Lynette's guidance (after a revelatory week with her in California's Trinity Mountains, searching for Bigfoot), Horn evolved, himself, resigning from Colorado State and seeking to remedy his inadequate Ivy League education. At a conference in Denver in September, Horn said he now realizes that humans come from an alien race of shape-shifting reptilians that continue to control civilization through the secretive leaders known as the Illuminati. Other panelists in Denver included enthusiasts describing their own experiences with various alien races. [Rocky Mountain Collegian, 9-28-09]

Wednesday, January 13, 2010

Non-useful graphical display of data: in the comics

Today's xkcd has a fantastic illustration of exactly what non-useful data display is.



Some displays of data are no more informative than the pie chart.

Tuesday, December 29, 2009

Dead salmon CAN think! Or an argument for multiple comparison corrections

Just over five years ago, a New Square fish store owner and his employee claimed to have had a talking carp:

Mr. Rosen said that when he approached the fish he heard it uttering warnings and commands in Hebrew.

"It said `Tzaruch shemirah' and `Hasof bah,' " he said, "which essentially means that everyone needs to account for themselves because the end is near."

The fish commanded Mr. Rosen to pray and to study the Torah and identified itself as the soul of a local Hasidic man who died last year, childless. The man often bought carp at the shop for the Sabbath meals of poorer village residents.


Few believed them, though many jokes were made such as the gefillte fish manufacturer who considered taking on the slogan "Our fish speaks for itself".

Now neuroscientists have documented brain activity, not just in a live carp of the story, but in a dead salmon, in the paper, Neural correlates of interspecies perspective taking in the post-mortem Atlantic Salmon: An argument for multiple comparisons correction. As they put it in their Methods section:

Subject. One mature Atlantic Salmon (Salmo salar) participated in the fMRI study. The salmon was approximately 18 inches long, weighed 3.8 lbs, and was not alive at the time of scanning.

Task. The task administered to the salmon involved completing an open-ended mentalizing task. The salmon was shown a series of photographs depicting human individuals in social situations with a specified emotional valence. The salmon was asked to determine what emotion the individual in the photo must have been experiencing.


Just picture the scene for a moment. I would love to talk to the research assistant who had to talk to the salmon and show it pictures. What kind of "mentalizing" does a person have while speaking to a dead fish?

The conclusion:

Can we conclude from this data that the salmon is engaging in the perspective-taking task? Certainly not. What we can determine is that random noise in the EPI timeseries may yield spurious results if multiple comparisons are not controlled for. Adaptive methods for controlling the false discovery rate and familywise error rate are excellent options and are widely available in all major fMRI analysis packages. We argue that relying on standard statistical thresholds (p < 0.001) and low minimum cluster sizes (k > 8) is an ineffective control for multiple comparisons. We further argue that the vast majority of fMRI studies should be utilizing multiple comparisons correction as standard practice in the computation of their statistics.


The study was also covered by Science News in an article on lack of replicability of fMRI experiments. The Science News story includes a quote that is a good idea for everyone, whether or not they do fMRI experiments: “Statistics should support common sense. If the math is so complicated that you don’t understand it, do something else.”

When this study wins the Ignobel Prize, you can say you saw that prediction here first.

Friday, October 2, 2009

Why the placebo effect is an effect

There was an interesting article in Wired recently that spoke about the placebo effect getting stronger: that the pre-post difference from a placebo drug is greater than it was a decade or two ago and that it differs between countries. That is, if you are looking at antidepressants and your outcome measure is a score on the Beck Depression Inventory that measures how depressed someone is, the score before the drug minus the score after the drug is different now than it was 10 years ago.

One criticism of the article is that the placebo effect cannot be considered an effect unless it is compared with another experimental condition. Since drug trials don't include both patients who receive a placebo and patients who receive nothing, there is no such thing as a placebo effect unless we know what the pre-post difference would have been in the absence of the placebo. Without a nothing arm to compare with, the writer contends that the pre-post difference in the placebo arm of a trial is just by definition the background noise in the trial.

I think that he's making a semantic point because a true placebo effect is impossible to measure.

To break the problem down further:

We do not know what the pre-post difference in a nothing arm of a trial would be. In some trials and for some diseases, there would be spontaneous improvement in the patient's condition: in that case, the pre-post difference in the placebo might just be that spontaneous improvement that would have happened if nothing were done.

In some trials and for some diseases, there would not be much change in the patient's condition, so the nothing arm would have no difference: in that case, the pre-post difference in the placebo arm would represent an "effect" and we could say that we have a placebo effect.

The question is which diseases have spontaneous improvement and which don't. There are three ways I can think of to figure this out.

1. A randomized clinical trial with patients that actually have some disease in which half the patients get a sugar pill and half the patients get nothing. No human subjects board would authorize this trial. Second, the study would not measure what we want it to. Ethically patients have to be told that the two possibilities are sugar pill and nothing. The Wired article contends that the placebo "effect" is based on a patient's prior beliefs about a drug's effectiveness, so it's specific to the drug, rather than being just the effect of a plain sugar pill.

2. The placebo effect could in theory be measured with matching, were there any subjects to match them to. The placebo pre-post difference can be defined in two ways: the pre-post difference of the sugar pill plus the pre-post difference of enrolling in the trial, or just the pre-post difference of the sugar pill alone. I would say it's the former. In that case, where we want to measure the effect of enrolling in a trial and taking a sugar pill, we could match normal patients with placebo patients based on their records and compare their pre-post differences. Except for the fact that medical records of normal patients with a disease are there because the patients are getting some treatment from their doctors. So there's no group to compare the placebo patients to.

3. The one remaining possibility is for each drug trial to divide their control group into two unequal groups: one receiving a sugar pill would be the larger group and one being put on a waiting list for the drug would be a smaller group. The problem is that placebos serve two purposes: one is for the statistical purpose and one is to keep the participants in the study and encourage them against taking other treatments. Depending on the condition, a control participant put on a waiting list might leave the trial or take another treatment in addition to the waiting list. So you might lose a good portion of the nothing arm of the trial.

Given the impossibility of rigorous measurement of what would happen under no treatment, the best we can do is guess which are the diseases where symptoms spontaneously resolve and which are the diseases where they don't. And that's what we already do when we talk about a placebo effect. We compare the pre-post difference in the placebo arm of a trial with our beliefs about what the pre-post difference would be with no treatment. In that sense, the placebo effect is really an effect. It's just imprecise.

Further, it's reasonable to assume that whatever the pre-post difference under nothing is, it's not going to change with time in any systematic way. If we could put all the placebo arms of, say, antidepressant trials together and find a trend with time, that's not sampling error. And that's exactly what the Wired article is talking about.

Statisticians protest at G-20 conference for safer data mining



Dating miners protest alongside United Steelworkers at the G-20 conference. My favorites: "Repeal Power Laws," "Our Sets. Our Axiom of Choice."

They even got John Oliver from the Daily Show to join in. I can't read the sign he is holding.

Sunday, September 13, 2009

Overly conservative statistics and yogurt



Are overly conservative statistics preventing the adoption of low-risk potentially beneficial health care?

It seems like probiotics are being talked about everywhere. We know that "good bacteria" are vital in many cases: babies delivered vaginally versus via c-section, for instance, have better immune function due in part to the bacterial colonization they get on their way out. (Of course, if the mother has chlamydia or other bad bacteria, the babies can get colonized by those too and develop eye infections.) Now that flu is in the air, people are citing studies that certain probiotics can help prevent and shorten flu infection. Probiotics are inexpensive and reasonably harmless: the worst side-effects I've seen attributed to them are the same as placebos such as mild GI distress. Probiotics seem like the canonical case of "can't hurt, could help." Kefir and yogurt are tasty, too.

Recently I ran across an immunologist's summary of the report of a 2005 Yale medical school conference about probiotics, mentioning among other things that probiotics might be able to help a disease a friend has. The hypothesized mechanism makes sense that it would help, so I looked at the Cochrane reviews, a formalized method for summarizing medical literature, and they say there's no evidence. The only studies were so hopelessly small, though, that there's no way to know at this point. So I looked up "probiotics" in Cochrane and got these results showing that there are about 82 abstracts relevant to probiotics. Of the 10 or so that I read, the only ones where Cochrane said there was conclusive evidence was for acute infectious diarrhea.

An interesting case: pediatric antibiotic-induced diarrhea. They noted the effects of missing data: if all the study drop-outs were treatment failure, which seems unlikely, the treatment doesn't work. Immediately after that, they acknowledge that there is almost no downside to the treatment: "Probiotics were generally well tolerated and side effects occurred infrequently." and yet they conclude, "Although current data are promising, there is insufficient evidence to routinely recommend the use of probiotics for the prevention of pediatric AAD."

In other words, there's no downside to using probiotics, but because the overly conservative statistical analysis that counts all treatment drop-outs as failures finds that they don't work, they can't recommend them. There are many reasons why subjects might have dropped out of this study, primarily boiling down to the studies being almost certainly poorly funded and unable to adequately compensate busy parents of sick children needing to catch up on their lives after their children recovered. That caution in counting drop-outs as failures is reasonable in some cases: for instance, if the proposed treatment is invasive or risky. Or in the case of the female condom hearings the commercial sex workers who dropped out of the study could have been the ones for whom the condoms didn't work as well. In this case of probiotics, they're virtually risk free and there's a good reason why parents may have dropped out of the study.

In medical statistics (biostatistics), the methods most commonly used are straight out of a textbook, rules of thumb that apply in general. Obviously context counts and we should be more conservative when there's a risk and less conservative when there's little risk. Biostatistics is not my primary area, but I have helped doctors out with the occasional clinical trial, using the textbook methods because that's what they wanted. There are many better methods that could be used to analyze this data, such as decision theory that accounts for risks, or missing data methods that model the potential outcomes of the study drop-outs. Biostatisticians have no malpractice risks, so there's no reason they couldn't be less conservative in their choice of data analysis methods to account for risk. Somehow the conservatism that US doctors practice under has spread to biostatisticians, though. Until statistics becomes less conservative in their analysis methods, patients may end up missing out on low-risk treatments still being studied.

Thursday, September 3, 2009

R Statistics flash mob for Tuesday




Wow, this is too funny.

> From: The R Flashmob Project
>Subject: R Flashmob #2
>
>You are invited to take part in R Flashmob, the project that makes the
>world a better place by posting helpful questions and answers about the
>R statistical language to the programmer’s Q & A site stackoverflow.com
>
>Please forward this to other people you know who might like to join.
>
>FAQ
>
>Q. Why would I want to join an inexplicable R mob?
>
>A. Tons of other people are doing it.
>
>Q. Why else?
>
>A. Stackoverflow was built specifically for handling programming questions.
>It’s a better mousetrap. It offers search (and is well indexed by search engines),
>tagging, voting, the ability to choose the “best” answer to a question, and the ability to
>edit questions and answers as technology progresses. It has a karma system to
>reward people who are happy to help and discourage MLJs (mailing list jerks).
>
>Q. Do the organizers of this MOB have any commercial interest in stackoverflow?
>
A. None at all. We’re just convinced it is the best way to help and promote R. All
>the content submitted to stackoverflow is protected by a Creative Commons
>CC-Wiki License, meaning anyone is free to copy, distribute, transmit, and
>remix the information on stackoverflow. All the content on stackoverflow is
>regularly made available for download by the public.
>
>INSTRUCTIONS – R MOB #2
>Location: stackoverflow.com
>Start Date: Tuesday, September 8th, 2009
>Start Time:
>10:04 AM – US Pacific
>11:04 AM – US Mountiain
>12:04 PM – US Central
>1:04 PM – US Eastern
>6:04 PM – UK
>7:04 PM – Continental W. Europe
>5:04 AM (Weds) – New Zealand (birthplace of R)
>Duration: 50 minutes
>
>(1) At some point during the day on September 8th, synchronize your watch to
>http://timeanddate.com/worldclock/personal.html?cities=137,75,64,179,136,37,22
>
>(2) The mob should form at precisely 4 minutes past the hour and not beforehand.
>
>(3) At 4 minutes past the hour, you should arrive at stackoverflow.com, log in,
>and post 3 R questions. Be sure to tag the questions “R”. See the posting
>guidelines at http://stackoverflow.com/faq to understand what makes a good
>question.
>
>(4) Follow R Flashmob updates at http://twitter.com/rstatsmob
>
>(5) Post twitter messages tagged #rstats and #rstatsmob during the mob,
>providing links to your questions.
>
>(6) During the R MOB, you can chat with other participants on the #R channel
>on IRC (freenode). To do this, install the Chatzilla extension on Firefox.
>Click “freenode” on the main screen. Then type /join #R in the field at the
>bottom of the screen. Then chat.
>
>(7) If you finish posting your three questions within the 50 minutes, stick
> around to answer questions and give “up votes” to good questions and answers.
>
>(8) IMPORTANT: After posting, sign the R Flashmob guestbook at
>http://bit.ly/6F8B2
>
>(9) Return to what you would otherwise have been doing. Await
>instructions for R MOB #3.