Results 1 to 4 of 4

Thread: Aha-moment for Minxie and for you

  1. #1

    Default Aha-moment for Minxie and for you

    Today, for the first time, I realised that statistical tests for significance--and the practice of having p < 0.05 as a cutoff line for significance--should be treated like any other test I may use in medicine. In other words, I have to consider not only the merits of the test itself, but also the characteristics of the things to which I apply it.

    The usefulness of a diagnostic test for a disease depends to a great extent on the number of people who have the disease in the tested population: if I test 1000 people, of whom only 5 have the disease for which I'm looking, then even a good test will give me some false alarms.

    This point is hammered home in med-schools here, and is part of the rationale behind the restrictive use of tests in many situations. However, for some reason, it's rarely (if ever) brought up in discussions of how to evaluate medical studies--people conducting (or making use of) medical research frequently behave as if p-values under 0.05 are assurances of true and significant relationships. Which is dangerous if you're applying your significance tests to 1000 relationships,only 100 of which are significant.





    I'm quite thrilled by this I've read articles about this before, and I've heard people talk about the problems with conducting "fishing expeditions" in research, but it's never fallen into place like this there's something about maths and stat that just makes my brain shut down.



    What's YOUR most recent aha-moment?? Ziggy, don't post "Take On Me"
    "One day, we shall die. All the other days, we shall live."

  2. #2
    I did not have an 'aha' moment whilst reading the requirements for your significance testing.

  3. #3
    Quote Originally Posted by Aimless View Post
    The usefulness of a diagnostic test for a disease depends to a great extent on the number of people who have the disease in the tested population: if I test 1000 people, of whom only 5 have the disease for which I'm looking, then even a good test will give me some false alarms.

    This point is hammered home in med-schools here, and is part of the rationale behind the restrictive use of tests in many situations. However, for some reason, it's rarely (if ever) brought up in discussions of how to evaluate medical studies--people conducting (or making use of) medical research frequently behave as if p-values under 0.05 are assurances of true and significant relationships. Which is dangerous if you're applying your significance tests to 1000 relationships,only 100 of which are significant.
    No offense, but I never cease to be amazed by the lack of statistical sophistication in epidemiology. There are statistical tests out there that take into account the rarity of the dependent variable (rare events logit for instance).

    Not sure what you mean by applying the tests to 1000 people and finding that only 100 are significant. Either your sample is significant or it is not. I think the point you're trying to get at is that a test of significance isn't perfect. For example, is there really a difference between a variable that has a p-value of 0.049 and one that has a p-value of 0.051 (the latter is not significant if using the standard two-tailed test). Furthermore, even if something is significant, there's a decent chance you're getting a false positive. Some statistician actually proved that a result that's significant at the 95% CI has a far higher than a 5% chance of being "wrong". All it means is that if you took the same sample 100 times, your result would be as high (or more so) than your current one in 95 of those samples. What it doesn't say is if those 95% of samples will themselves be significant. In fact, something like 20% would not be (don't remember the exact number). The solution isn't to disregard the test, but to A) test different scenarios, B) run the tests using other data (if possible), and C) have a strong theory that explains the results, and D) test possible implications of the theory.
    Hope is the denial of reality

  4. #4
    Quote Originally Posted by Loki View Post
    No offense, but I never cease to be amazed by the lack of statistical sophistication in epidemiology. There are statistical tests out there that take into account the rarity of the dependent variable (rare events logit for instance).
    I think part of the problem is due to ignorance or a lack of understanding of the methods (so that problems don't get filtered out by the peer review process nor out in clinical practice) and part of it is due to the constant pressure of having to publish positive findings in order to secure funding I'm not sure how to account for the relative scarcity of true relationships in epidemiological research using statistical methods, given that we don't really know beforehand precisely how many of the relationships we investigate are true.

    Not sure what you mean by applying the tests to 1000 people and finding that only 100 are significant. Either your sample is significant or it is not. I think the point you're trying to get at is that a test of significance isn't perfect. For example, is there really a difference between a variable that has a p-value of 0.049 and one that has a p-value of 0.051 (the latter is not significant if using the standard two-tailed test). Furthermore, even if something is significant, there's a decent chance you're getting a false positive. Some statistician actually proved that a result that's significant at the 95% CI has a far higher than a 5% chance of being "wrong". All it means is that if you took the same sample 100 times, your result would be as high (or more so) than your current one in 95 of those samples. What it doesn't say is if those 95% of samples will themselves be significant. In fact, something like 20% would not be (don't remember the exact number).
    The comparison I was making in my head was (I think) to predictive value. With my numbers I was just comparing diagnostic tests in medicine to tests of significance in research: a disease prevalence of 10% (100 in 1000) or a true relationship prevalence of 10% (based on some estimates that only about 10% of all investigated relationships in medical epidemiological research are true). Not very different from what you outlined in your post, I think

    The solution isn't to disregard the test, but to A) test different scenarios, B) run the tests using other data (if possible), and C) have a strong theory that explains the results, and D) test possible implications of the theory.
    All great ways to keep making money off of epidemiological research but yeah, I wasn't implying that tests of significance should be dismissed entirely. The measures you describe are just good practice
    "One day, we shall die. All the other days, we shall live."

Posting Permissions

  • You may not post new threads
  • You may not post replies
  • You may not post attachments
  • You may not edit your posts
  •