Sunday, January 29, 2017

17 what is a traceback? run time environment: it could be the client side


I am posting this answer so the topmost answer (when sorted by activity) is not one that is just plain wrong.
http://stackoverflow.com/questions/3988788/what-is-a-stack-trace-and-how-can-i-use-it-to-debug-my-application-errors
What is a Stacktrace?
A stacktrace is a very helpful debugging tool. It shows the call stack (meaning, the stack of functions that were called up to that point) at the time an uncaught exception was thrown (or the time the stacktrace was generated manually). This is very useful because it doesn't only show you where the error happened, but also how the program ended up in that place of the code. This leads over to the next question:
What is an Exception?
An Exception is what the runtime environment uses to tell you that an error occurred. Popular examples are NullPointerException, IndexOutOfBoundsException or ArithmeticException. Each of these are caused when you try to do something that is not possible. For example, a NullPointerException will be thrown when you try to dereference a Null-object:
Object a = null;
a.toString();                 //this line throws a NullPointerException

Object[] b = new Object[5];
System.out.println(b[10]);    //this line throws an IndexOutOfBoundsException,
                              //because b is only 5 elements long
int ia = 5;
int ib = 0;
ia = ia/ib;                   //this line throws an  ArithmeticException with the 
                              //message "/ by 0", because you are trying to
                              //divide by 0, which is not possible.
How should I deal with Stacktraces/Exceptions?
At first, find out what is causing the Exception. Try googleing the name of the exception to find out, what is the cause of that exception. Most of the time it will be caused by incorrect code. In the given examples above, all of the exceptions are caused by incorrect code. So for the NullPointerException example you could make sure that a is never null at that time. You could, for example, initialise a or include a check like this one:
if (a!=null) {
    a.toString();
}
This way, the offending line is not executed if a==null. Same goes for the other examples.
Sometimes you can't make sure that you don't get an exception. For example, if you are using a network connection in your program, you cannot stop the computer from loosing it's internet connection (e.g. you can't stop the user from disconnecting the computer's network connection). In this case the network library will probably throw an exception. Now you should catch the exception and handle it. This means, in the example with the network connection, you should try to reopen the connection or notify the user or something like that. Also, whenever you use catch, always catch only the exception you want to catch, do not use broad catch statements like catch (Exception e)that would catch all exceptions. This is very important, because otherwise you might accidentally catch the wrong exception and react in the wrong way.
try {
    Socket x = new Socket("1.1.1.1", 6789);
    x.getInputStream().read()
} catch (IOException e) {
    System.err.println("Connection could not be established, please try again later!")
}
Why should I not use catch (Exception e)?
Let's use a small example to show why you should not just catch all exceptions:
int mult(Integer a,Integer b) {
    try {
        int result = a/b
        return result;
    } catch (Exception e) {
        System.err.println("Error: Division by zero!");
        return 0;
    }
}
What this code is trying to do is to catch the ArithmeticException caused by a possible division by 0. But it also catches a possible NullPointerException that is thrown if a or b are null. This means, you might get a NullPointerException but you'll treat it as an ArithmeticException and probably do the wrong thing. In the best case you still miss that there was a NullPointerException. Stuff like that makes debugging much harder, so don't do that.
TLDR
  1. Figure out what is the cause of the exception and fix it, so that it doesn't throw the exception at all.
  2. If 1. is not possible, catch the specific exception and handle it.
    • Never just add a try/catch and then just ignore the exception! Don't do that!
    • Never use catch (Exception e), always catch specific Exceptions. That will save you a lot of headaches.
shareedit
   
Thanks, this is a really helpful answer. Please can you fix the typo: "try to deference a Null-object". (You need fixes to edit a post unless you are the author) – Zach Smith Jan 19 '16 at 8:27
   
Thanks for the reply and the typo ;) – Dakkaron Jan 19 '16 at 9:30
   
@Dakkaron: Great answer, thank you! I suppose ia = ia/b; should be ia = ia/ib;? – Setily Sep 9 '16 at 12:53
   
@Setily: thanks, I fixed it – Dakkaron Sep 13 '16 at 14:17
1 
nice explanation for why we should avoid bug masking – Sudip Bhandari Dec 20 '16 at 10:07 

To add on to what Rob has mentioned. Setting break points in your application allows for the step-by-step processing of the stack. This enables the developer to use the debugger to see at what exact point the method is doing something that was unanticipated.
Since Rob has used the NullPointerException (NPE) to illustrate something common, we can help to remove this issue in the following manner:
if we have a method that takes parameters such as: void (String firstName)
In our code we would want to evaluate that firstName contains a value, we would do this like so: if(firstName == null || firstName.equals("")) return;
The above prevents us from using firstName as an unsafe parameter. Therefore by doing null checks before processing we can help to ensure that our code will run properly. To expand on an example that utilizes an object with methods we can look here:
if(dog == null || dog.firstName == null) return;
The above is the proper order to check for nulls, we start with the base object, dog in this case, and then begin walking down the tree of possibilities to make sure everything is valid before processing. If the order were reversed a NPE could potentially be thrown and our program would crash.
shareedit
   
Agreed. This approach could be used to find out which reference in a statement is null when a NullPointerException is being examined, for example. – Rob Hruska Oct 21 '10 at 15:07
13 
When dealing with String, if you want to use equals method I think it´s better to use the constant in the left side of the comparation, like this: Instead of: if(firstName == null || firstName.equals("")) return; I always use: if(("").equals(firstName)) This prevents the Nullpointer exception – Torres Oct 26 '10 at 6:23 


I maintain an old application written in VB6. In client's environment it raises runtime errors which I can't reproduce under debugger. Is there any way to get the stacktrace or location of error?
I mean, without putting trace statements all over the code like here or adding error handlers for logging to every procedure like here.
It seems to be a simple question. Sorry. I just don't know VB6 very well. And it is surprisingly hard to google out any information, considering how widely it is (or used to be) used.
shareedit
   
I asked the same question stackoverflow.com/questions/127645/… I'm convinced that it can't be done. – raven Jul 8 '09 at 18:35
   
Maybe I wasn't clear. I have an application in production in a remote location. I haven't got access to this system and I can't run debugger there. There is something in their environment that triggers runtime error. I can't expect their IT staff (not to mention regular users) to really help, beyond sending me whatever application's shown or dumped to log file. I need some tool, instrumentation or anything, that will help me get meaningful input from them. Is--as raven writes--writing a "On Error GoTo/Reraise/LogError" in every routine the only way? – Tomek Szpakowicz Jul 9 '09 at 15:28
   
Well, you could compile with debug symbols as I mentioned, then get them to do a memory dump when the error occurs. You'll then be able to load the memory dump and hopefully get the stack trace using Visual Studio. – Ant Jul 10 '09 at 8:57
   
The "get them to do" just about anything is the hardest part. – Tomek Szpakowicz Jul 10 '09 at 11:31
   
@Ant - will only work if these are unhandled exception errors rather than intrinsic Visual Basic runtime errors. It is not clear from the question which it is. – MarkJ Jul 12 '09 at 17:14

4 Answers

Try compiling to pcode and see if you still get the error. This is one common difference between the debug mode of VB6 and runtime. I used to compile to native and ran into errors that only occurred in runtime. When I switched to pcode I found either the error went away or more likely a new error that reflected the real problem cropped up and was more easily reproduced in debug mode.
If despite that you still getting the error then I really recommend starting at the top of your procedure stack and working you way down using Maero's suggestion of
On Error Goto Handler
<code>
Exit <routine>
Handler:
Err.Raise Err.Number, "(function_name)->" & Err.source, Err.Description
It is a pain but there is no real way around it.
shareedit
   
The app is compiled to P-code. The problem is not debugging native code. The problem is, runtime error happens only in environment, I don't have access to. I just expected that interpreted (P-code) code would be able to give me some more information about runtime error in production system than C/C++, without me putting trace statements/error handlers all over the code. – Tomek Szpakowicz Jul 9 '09 at 15:36
   
I see one big drawback here. Now, if I do this, all runtime errors are handled. Debugger will not stop application at error location. Instead it will stop inside error handler in some other procedure down the stack. So this method helps with a I-have-no-debugger-in-production-environment scenario but breaks normal work with VB6 IDE. – Tomek Szpakowicz Jul 10 '09 at 12:58
2 
@tomekszpakowicz: correct! A classic problem with a classic solution. Try this "If Not IsInIDE() Then On Error Goto Handler", using the IsInIDE function from here vbnet.mvps.org/index.html?code/helpers/isinide.htm – MarkJ Jul 12 '09 at 17:17
2 
I'm not really satisfied by any answer here. I've chosen this answer only because a) I actually used this technique to pinpoint the cause of failure, although I'm not willing to litter all the code with this stuff, and b) as we migrate our codebase to .NET, I simply stopped to care about finding some better solution. – Tomek Szpakowicz Jul 15 '10 at 8:43
The VB6 debugger is flaky sometimes. There are alternatives.
  • You could try Windbg, a free standalone debugger from Microsoft. Compile your VB6 with no optimisation and "create symbolic debug info" (i.e. create PDB files), and you will be able to debug. Here's a 2006 blog post by a Microsoft guy about using Windbg with VB6, and 2004 blog post by another Microsoft guy with a brief introduction to Windbg.
  • You could also use the Visual Studio 2008 debugger with VB6 and PDB files, e.g. with Visual C++ Express Edition (which is free). See this for more details.
  • Both Windbg and Visual Studio expect the source code to be in exactly the same path on the debug machine as it was on the build machine when the VB6 was built. The easiest way is to build and debug on the same machine. Otherwise you might need to fiddle with SUBST to create virtual drives - or I'm told the serious way is to use a Symbol Server.
shareedit
If you check the "Create Symbolic Debug Info" checkbox on the Project Properties/Compile tab, then you can debug in Visual Studio just like you would a native C++ application.
shareedit
   
+1. You can also use free debuggers like Windbg or Visual Studio 2008, see my answer. – MarkJ Jul 12 '09 at 17:21
   
I meant Visual Studio 2008 (or 2003 or 2005 or whatever), but yeah - good point about Windbg! – Ant Jul 13 '09 at 9:31
It's been a while, but I don't think there is a way to get a stack trace in a VB6 application without adding an error handler and outputting the appropriate message. There were some third party tools that would add error handling to an entire application but I believe it just added "On Error Goto" error handlers throughout the code.
Just as an aside, one of the more insidious runtime errors I ever encountered in a VB6 app was when I used a font that didn't exist on the client's PC in the property of a control. This generates a runtime error that cannot be trapped in code, so no amount of error handling that I added ever uncovered the error. I finally came across it by chance. Hope this helps.
shareedit
1 
If you invoke forms via the deprecated Form1.Show method, you won't be able to catch the error, but if you use the Dim form1Instance as Form1: Set form1Instance = new Form1(): form1Instance.Show syntax, an error will be thrown at the Set... line – rpetrich Jul 10 '09 at 18:13

python 16. what is a docstring?

16. what is a docstring?

http://www.pythonforbeginners.com/basics/python-docstrings

Python Docstrings

What is a Docstring?

Python documentation strings (or docstrings) provide a convenient way of
associating documentation with Python modules, functions, classes, and methods. 

An object's docsting is defined by including a string constant as the first
statement in the object's definition. 

It's specified in source code that is used, like a comment, to document a
specific segment of code.

Unlike conventional source code comments the docstring should describe what the
function does, not how.

All functions should have a docstring

This allows the program to inspect these comments at run time, for instance as
an interactive help system, or as metadata.

Docstrings can be accessed by the __doc__ attribute on objects.

How should a Docstring look like?

The doc string line should begin with a capital letter and end with a period. 

The first line should be a short description.

Don't write the name of the object. 

If there are more lines in the documentation string, the second line should be
blank, visually separating the summary from the rest of the description. 

The following lines should be one or more paragraphs describing the object’s
calling conventions, its side effects, etc.

Docstring Example

Let's show how an example of a multi-line docstring:
def my_function():
    """Do nothing, but document it.

    No, really, it doesn't do anything.
    """
    pass
Let's see how this would look like when we print it
>>> print my_function.__doc__
Do nothing, but document it.

    No, really, it doesn't do anything.

Declaration of docstrings

The following Python file shows the declaration of docstrings within a python
source file:
"""
Assuming this is file mymodule.py, then this string, being the
first statement in the file, will become the "mymodule" module's
docstring when the file is imported.
"""
 
class MyClass(object):
    """The class's docstring"""
 
    def my_method(self):
        """The method's docstring"""
 
def my_function():
    """The function's docstring"""

How to access the Docstring

The following is an interactive session showing how the docstrings may be accessed
>>> import mymodule
>>> help(mymodule)
Assuming this is file mymodule.py then this string, being the first statement in 
the file will become the mymodule modules docstring when the file is imported. 
>>> help(mymodule.MyClass)
The class's docstring

>>> help(mymodule.MyClass.my_method)
The method's docstring

>>> help(mymodule.my_function)
The function's docstring

Tuesday, January 24, 2017

What Happens If An Astronaut Floats Off In Space?





What Happens If An Astronaut Floats Off In Space?

In short: he's in trouble.

By Erik Sofge posted Sep 30th, 2013 at 10:00am
Courtesy Warner Bros. Pictures

Space Scare

Despite the risks, no mission has ever lost a space-walking astronaut.

In the film Gravity, which opens this month, two astronauts are on a spacewalk when an accident hurtles them into the void. So what would actually happen if you went, in NASA's terminology, "overboard"?

NASA requires spacewalking astronauts to use tethers (and sometimes additional anchors). But should those fail, you'd float off according to whatever forces were acting on you when you broke loose. You'd definitely be weightless. You'd possibly be spinning. In space, no kicking and flailing can change your fate. And your fate could be horrible. At the right angle and velocity, you might even fall back into Earth's atmosphere and burn up. That's why NASA has protocols that it drills into astronauts for such situations. You would be wearing your emergency jetpack, called SAFER, which would automatically counter any tumbling to stabilize you. Then NASA's plan dictates that you take manual control and fly back to safety.

However, if the pack's three pounds of fuel runs out, if another astronaut doesn't quickly grab you, or if the air lock is irreparably damaged, you're in big trouble. No protocols can save you now. (In fact, there aren't any.) At the moment, there's no spacecraft to pick you up. The only one with a rescue-ready air-locked compartment—the Space Shuttle—is in retirement. So your only choice is to orbit, waiting for your roughly 7.5 hours of breathable air to run out. It wouldn't be too terrible. You might get a little hungry, but there's up to a liter of water available via straw in your helmet. You'd simply sip and think of your family as you watched the sun rise and set—approximately five times, depending on your altitude.




As a start:

http://lipn.univ-paris13.fr/~duchamp/Books&more/Penrose/Road_to_Reality-CAPE_JONATHAN_(RAND)(2004).pdf

Page 391:

"Fig. 17.5 (a) Galileo’s (alleged) experiment. Two rocks, one large and one small, are dropped from the top of Leaning Tower of Pisa. Galileo’s insight was that if the eVects of air resistance can be ignored, each would fall at the same rate. (b) Oppositely charged pith balls (of equal small mass), in an electric Weld, directed towards the ground. One charge would ‘fall’ downwards, but the other would rise upwards.

Now the Wrst point to make here is that this is a particular property of the gravitational Weld, and it is not to be expected for any other force acting on bodies. The property of gravity that Galileo’s insight depends upon is the fact that the strength of the gravitational force on a body, exerted by some given gravitational Weld, is proportional to the mass of that body, whereas the resistance to motion (the quantity m appearing in Newton’s second law) is also the mass. It is useful to distinguish these two mass notions and call the Wrst the gravitational mass and the second, the inertial mass. (One might also choose to distinguish the passive from the active gravitational mass. The passive mass is the contribution m in Newton’s inverse square formula GmM/r2, when we consider the gravitational force on the m particle due to the M particle. When we consider the force on the M particle due to the m particle, then the mass m appears in its active role. But Newton’s third law decrees that passive and active masses be equal, so I am not going to distinguish between these two here.6 ) Thus, Galileo’s insight depends upon the equality (or, more correctly, the proportionality) of the gravitational and inertial mass"

Tuesday, January 3, 2017

David Colquhoun: The problem with p-values

https://aeon.co/essays/it-s-time-for-science-to-abandon-the-term-statistically-significant

Academic psychology and medical testing are both dogged by unreliability. The reason is clear: we got probability wrong

False positives; numerous microscopic cancerous and non-cancerous human tissue samples. Photo courtesy Wellcome Images
David Colquhoun
is a professor of pharmacology at University College London and a Fellow of the Royal Society. He is the author of Lectures on Biostatistics (1971) and blogs at DC’s Improbable Science.
2,400 words
Edited by Sally Davies
REPUBLISH - LICENCE ONLY
What should be done to improve statistical literacy?
 109Responses
The aim of science is to establish facts, as accurately as possible. It is therefore crucially important to determine whether an observed phenomenon is real, or whether it’s the result of pure chance. If you declare that you’ve discovered something when in fact it’s just random, that’s called a false discovery or a false positive. And false positives are alarmingly common in some areas of medical science.
In 2005, the epidemiologist John Ioannidis at Stanford caused a storm when he wrote the paper ‘Why Most Published Research Findings Are False’, focusing on results in certain areas of biomedicine. He’s been vindicated by subsequent investigations. For example, a recent article found that repeating 100 different results in experimental psychology confirmed the original conclusions in only 38 per cent of cases. It’s probably at least as bad for brain-imaging studies and cognitive neuroscience. How can this happen?
The problem of how to distinguish a genuine observation from random chance is a very old one. It’s been debated for centuries by philosophers and, more fruitfully, by statisticians. It turns on the distinction between induction and deduction. Science is an exercise in inductive reasoning: we are making observations and trying to infer general rules from them. Induction can never be certain. In contrast, deductive reasoning is easier: you deduce what you would expect to observe if some general rule were true and then compare it with what you actually see. The problem is that, for a scientist, deductive arguments don’t directly answer the question that you want to ask.
What matters to a scientific observer is how often you’ll be wrong if you claim that an effect is real, rather than being merely random. That’s a question of induction, so it’s hard. In the early 20th century, it became the custom to avoid induction, by changing the question into one that used only deductive reasoning. In the 1920s, the statistician Ronald Fisher did this by advocating tests of statistical significance. These are wholly deductive and so sidestep the philosophical problems of induction.
Tests of statistical significance proceed by calculating the probability of making our observations (or the more extreme ones) if there were no real effect. This isn’t an assertion that there is no real effect, but rather a calculation of what would be expected if there were no real effect. The postulate that there is no real effect is called the null hypothesis, and the probability is called the p-value. Clearly the smaller the p-value, the less plausible the null hypothesis, so the more likely it is that there is, in fact, a real effect. All you have to do is to decide how small the p-value must be before you declare that you’ve made a discovery. But that turns out to be very difficult.
The problem is that the p-value gives the right answer to the wrong question. What we really want to know is not the probability of the observations given a hypothesis about the existence of a real effect, but rather the probability that there is a real effect – that the hypothesis is true – given the observations. And that is a problem of induction.
Confusion between these two quite different probabilities lies at the heart of why p-values are so often misinterpreted. It’s called the error of the transposed conditional. Even quite respectable sources will tell you that the p-value is the probability that your observations occurred by chance. And that is plain wrong.
Get Aeon straight to your inbox
Suppose, for example, that you give a pill to each of 10 people. You measure some response (such as their blood pressure). Each person will give a different response. And you give a different pill to 10 other people, and again get 10 different responses. How do you tell whether the two pills are really different?
The conventional procedure would be to follow Fisher and calculate the probability of making the observations (or the more extreme ones) if there were no true difference between the two pills. That’s the p-value, based on deductive reasoning. P-values of less than 5 per cent have come to be called ‘statistically significant’, a term that’s ubiquitous in the biomedical literature, and is now used to suggest that an effect is real, not just chance.
But the dichotomy between ‘significant’ and ‘not significant’ is absurd. There’s obviously very little difference between the implication of a p-value of 4.7 per cent and of 5.3 per cent, yet the former has come to be regarded as success and the latter as failure. And ‘success’ will get your work published, even in the most prestigious journals. That’s bad enough, but the real killer is that, if you observe a ‘just significant’ result, say P = 0.047 (4.7 per cent) in a single test, and claim to have made a discovery, the chance that you are wrong is at least 26 per cent, and could easily be more than 80 per cent. How can this be so?Take the proposition that the Earth goes round the Sun. It either does or it doesn’t, so it’s hard to see how we could pick a probability for this statement
For one, it’s of little use to say that your observations would be rare if there were no real difference between the pills (which is what the p-value tells you), unless you can say whether or not the observations would also be rare when there is a true difference between the pills. Which brings us back to induction.
The problem of induction was solved, in principle, by the Reverend Thomas Bayes in the middle of the 18th century. He showed how to convert the probability of the observations given a hypothesis (the deductive problem) to what we actually want, the probability that the hypothesis is true given some observations (the inductive problem). But how to use his famous theorem in practice has been the subject of heated debate ever since.
Take the proposition that the Earth goes round the Sun. It either does or it doesn’t, so it’s hard to see how we could pick a probability for this statement. Furthermore, the Bayesian conversion involves assigning a value to the probability that your hypothesis is right beforeany observations have been made (the ‘prior probability’). Bayes’s theorem allows that prior probability to be converted to what we want, the probability that the hypothesis is true given some relevant observations, which is known as the ‘posterior probability’.
These intangible probabilities persuaded Fisher that Bayes’s approach wasn’t feasible. Instead, he proposed the wholly deductive process of null hypothesis significance testing. The realisation that this method, as it is commonly used, gives alarmingly large numbers of false positive results has spurred several recent attempts to bridge the gap.
There is one uncontroversial application of Bayes’s theorem: diagnostic screening, the tests that doctors give healthy people to detect warning signs of disease. They’re a good way to understand the perils of the deductive approach.
In theory, picking up on the early signs of illness is obviously good. But in practice there are usually so many false positive diagnoses that it just doesn’t work very well. Take dementia. Roughly 1 per cent of the population suffer from mild cognitive impairment, which might, but doesn’t always, lead to dementia. Suppose that the test is quite a good one, in the sense that 95 per cent of the time it gives the right (negative) answer for people who are free of the condition. That means that 5 per cent of the people who don’t have cognitive impairment will test, falsely, as positive. That doesn’t sound bad. It’s directly analogous to tests of significance which will give 5 per cent of false positives when there is no real effect, if we use a p-value of less than 5 per cent to mean ‘statistically significant’.
But in fact the screening test is not good – it’s actually appallingly bad, because 86 per cent, not 5 per cent, of all positive tests are false positives. So only 14 per cent of positive tests are correct. This happens because most people don’t have the condition, and so the false positives from these people (5 per cent of 99 per cent of the people), outweigh the number of true positives that arise from the much smaller number of people who have the condition (80 per cent of 1 per cent of the people, if we assume 80 per cent of people with the disease are detected successfully). There’s a YouTube video of my attempt to explain this principle, or you can read my recent paper on the subject.
The number of false positives in the tests where there is no real effect outweighs the number of true positives that arise from the cases in which there is a real effect
Notice, though, that it’s possible to calculate the disastrous false-positive rate for screening tests only because we have estimates for the prevalence of the condition in the whole population being tested. This is the prior probability that we need to use Bayes’s theorem. If we return to the problem of tests of significance, it’s not so easy. The analogue of the prevalence of disease in the population becomes, in the case of significance tests, the probability that there is a real difference between the pills before the experiment is done – the prior probability that there’s a real effect. And it’s usually impossible to make a good guess at the value of this figure.
An example should make the idea more concrete. Imagine testing 1,000 different drugs, one at a time, to sort out which works and which doesn’t. You’d be lucky if 10 per cent of them were effective, so let’s proceed by assuming a prevalence or prior probability of 10 per cent. Say we observe a ‘just significant’ result, for example, a P = 0.047 in a single test, and declare that this is evidence that we have made a discovery. That claim will be wrong, not in 5 per cent of cases, as is commonly believed, but in 76 per cent of cases. That is disastrously high. Just as in screening tests, the reason for this large number of mistakes is that the number of false positives in the tests where there is no real effect outweighs the number of true positives that arise from the cases in which there is a real effect.
In general, though, we don’t know the real prevalence of true effects. So, although we can calculate the p-value, we can’t calculate the number of false positives. But what we can do is give a minimum value for the false positive rate. To do this, we need only assume that it’s not legitimate to say, before the observations are made, that the odds that an effect is real are any higher than 50:50. To do so would be to assume you’re more likely than not to be right before the experiment even begins.
If we repeat the drug calculations using a prevalence of 50 per cent rather than 10 per cent, we get a false positive rate of 26 per cent, still much bigger than 5 per cent. Any lower prevalence will result in an even higher false positive rate.
The upshot is that, if a scientist observes a ‘just significant’ result in a single test, say P = 0.047, and declares that she’s made a discovery, that claim will be wrong at least 26 per cent of the time, and probably more. No wonder then that there are problems with reproducibility in areas of science that rely on tests of significance.
What is to be done? For a start, it’s high time that we abandoned the well-worn term ‘statistically significant’. The cut-off of P < 0.05 that’s almost universal in biomedical sciences is entirely arbitrary – and, as we’ve seen, it’s quite inadequate as evidence for a real effect. Although it’s common to blame Fisher for the magic value of 0.05, in fact Fisher said, in 1926, that P = 0.05 was a ‘low standard of significance’ and that a scientific fact should be regarded as experimentally established only if repeating the experiment ‘rarely fails to give this level of significance’.
The ‘rarely fails’ bit, emphasised by Fisher 90 years ago, has been forgotten. A single experiment that gives P = 0.045 will get a ‘discovery’ published in the most glamorous journals. So it’s not fair to blame Fisher, but nonetheless there’s an uncomfortable amount of truth in what the physicist Robert Matthews at Aston University in Birmingham had to say in 1998: ‘The plain fact is that 70 years ago Ronald Fisher gave scientists a mathematical machine for turning baloney into breakthroughs, and flukes into funding. It is time to pull the plug.’
The underlying problem is that universities around the world press their staff to write whether or not they have anything to say. This amounts to pressure to cut corners, to value quantity rather than quality, to exaggerate the consequences of their work and, occasionally, to cheat. People are under such pressure to produce papers that they have neither the time nor the motivation to learn about statistics, or to replicate experiments. Until something is done about these perverse incentives, biomedical science will be distrusted by the public, and rightly so. Senior scientists, vice-chancellors and politicians have set a very bad example to young researchers. As the zoologist Peter Lawrence at the University of Cambridge put it in 2007:
hype your work, slice the findings up as much as possible (four papers good, two papers bad), compress the results (most top journals have little space, a typical Nature letter now has the density of a black hole), simplify your conclusions but complexify the material (more difficult for reviewers to fault it!)
But there is good news too. Most of the problems occur only in certain areas of medicine and psychology. And despite the statistical mishaps, there have been enormous advances in biomedicine. The reproducibility crisis is being tackled. All we need to do now is to stop vice-chancellors and grant-giving agencies imposing incentives for researchers to behave badly.