Showing posts with label Firger. Show all posts
Showing posts with label Firger. Show all posts

Sunday, October 18, 2015

Replicability: A Key Standard of Scientific Method Now Found to be Under Existential Threat

Nine key standards, or principles, of scientific research can be gleaned from texts on research method. As I see them, these are:
2. Replicability
3. Control
4. Measurability
5. Validity
6. Reliability (Pending)
7. Objectivity (Pending)
8. Ethics
9. Representativeness.

I will say something about each of the nine standards in later posts. What I want to briefly talk about today is replicability. It turns out that, highly valued as it is, replicability is easier 'said' in conversations about methodology -- defined as the science or theory of method -- than 'done' in actual empirical research. This is what we pick from a study captured in Jessica Firger's account, the link to which I give in the References section below.

First, though, let us recapitulate the accepted meaning and procedures of replicability. Then, second, we will turn to the new study to see how its findings shake the foundations of the scientist's faith (now 'blind faith'?) in the utility and robustness of replicabiliy as a verification process.

First, then, about the centrality of replication to scientific method and to science's "theory of evidence": As a concept of methodology, replicability calls for precision in measurement and accuracy of procedure in all scientific endeavors. That is to say, it calls for the doing of any research project in such a way as to ensure its step-by-step, beginning-to-end, repeatability or 'reproducibility' -- which is what replicability means -- in any subsequent study, by any competent researcher, intended to attest to the veracity of its discoveries or claims. To repeat, it should be repeatable in its exact original form, and should yield the very same results. Thus:
“The research should be repeatable. Any other competent researcher should be able to take your problem and , collecting data under the same circumstances and within the identical parameters as you have observed, achieve results comparable (sic) to those you have been able to secure” (Leedy, 1980: 46; see also Seale, 2004). 
That emphasis on "the very same results", as opposed to 'comparable' results, highlights the ultimate basis for the confirmation or validation of scientific claims. However, the confounding truth, which subverts such a neatly articulated metric, is that several important factors may be suggested as underpinning the likelihood of not achieving such results in the real world which social science, including/particularly sociology, routinely deals with. This is the real world of conscious, querying, skeptical (and often over-researched) human beings unsure of the risks, or real value, of availing themselves and telling everything that researchers wish to be told. And that's even before one factors in variability, for various reasons, in the ways respected scientists implement even widely agreed-upon procedures -- or the procedures that should, in the name of reason and perhaps common sense, be so agreed -- for data capture, processing, analysis and synthesis.

But why all this fuss about replicability or reproducibility to which I believe all seasoned interrogators of methodology have been paying attention for years? Firger, I think, captures the underlying sentiment as well as anyone can or has:
"Most rational humans hold some faith in science. We trust that scientists, through their hard-won expertise, are well equipped to conduct studies that provide proof of why things are the way they are, help solve problems and explain mysteries. Much of that trust is built on an implicit belief that research findings are concrete truths—in other words, that if given the same set of parameters, they would be easy to reproduce. This type of replication is essential to science because it validates key discoveries and helps scientists make progress in their fields of research" (Firger, 2015).
So replicability is a standard worth respecting, implementing and defending. But what if it turns out that in practice we cannot, or have not been doing so? What are we to make of our scientific truths, then? That, at least implicitly, is for me the most fundamental, if implicit, question driving the University of Virginia study, as we will now see.

Second, then: what was the purpose of the new study and what does it tell us? We can discern that purpose and related method from this portion of the study's problem statement, as seen in the abstract:
"Reproducibility is a defining feature of science, but the extent to which it characterizes current research is unknown. We conducted replications of 100 experimental and correlational studies published in three psychology journals using high-powered designs and original materials when available"  (Open Science Collaboration, 2015)
What about the results? The key finding of the study was that the results of less than 50% of the studies subjected to replication were corroborated; that is, independently validated. That is a huge shortfall, as elaborated below:
"Replication effects (Mr = .197, SD = .257) were half the magnitude of original effects (Mr = .403, SD = .188), representing a substantial decline. Ninety-seven percent of original studies had significant results (p < .05). Thirty-six percent of replications had significant results; 47% of original effect sizes were in the 95% confidence interval of the replication effect size; 39% of effects were subjectively rated to have replicated the original result; and, if no bias in original results is assumed, combining original and replication results left 68% with significant effects" (Open Source Collaboration, 2015; see also PSA, September 2015).
Firger (2015) notes problem in this regard: that the threat to replicability may be exacerbated by the existential pressure, perceived by scientists seeking recognition and success (presumably via promotion and tenure, or higher levels of funding, or 'all of the above'), to dump rigor, and all its 'yokes' and promise, in favor of building the capacity to "weave a memorable story out of tenuous science" -- and so to achieve such success.

In the end, perfect replicability, certainly of studies involving conscious human subjects, may not be feasible, even when driven purely by the dictates of scientific rigor, ahead of the capacity for time-travel back to the chronological 'moment' at which the original study of interest was conducted. Incidentally, contrary to Drummond's (2009) claim, this is not to say that 'reproducibility' is a more viable posture for the scientist. The attributes he assigns to reproducibility are in fact, precisely, those of replicability, as widely understood by those who routinely use the latter term (see, for example, the clarification offered by the Replicability Research Group, 2015). What Drummond wrongly sees as the weaknesses of replicability are its ideals and strengths, and what he touts as the strengths of reproducibility are its weaknesses. Reproducibility, defined as he so facilely does, would be a recipe for widespread intellectual fraud.

So what is one to do in the meantime? We will just have to make do with such rigors and caveats as we are able to muster and agree upon. But the rationale for replicating scientific work remains as great and as persuasive as it ever was, and its utility for science and human progress just as enormous. All that we have discovered is its disappointingly low success score, certainly in psychology -- which is a human science. We have also found, alas, that the pressure for recognition and success makes the scientist as much an opportunist as the proverbial self-sacrificing servant of truth. We can work with sharp focus on these personal/human 'frailties' of the scientist -- not necessarily of science -- to raise replicability's score. Replication is (natural) science's ultimate audit tool, which we will abandon only at society's and civilization's own great peril.

Still, we will also have to more consciously and more robustly develop other methods of arriving at the truth -- and the natural scientist, in particular, will have to respect with greater humility and sharper awareness of that potentially debilitating soft underbelly, just discovered, of his/her cherished branch of science. The other methods I have in mind are those which social science and the humanities have been grappling with, self-critically (and against the backdrop of harsh and even dismissive criticism by often sanctimonious natural scientists), for decades, and even more than a century, now. Many of these fall under the qualitative label, as opposed to the quantitative, and include versions of such broad and cross-cutting, timeless and a-disciplinary methods as: observation (naturalistic, participant or device-assisted), comparison, deduction, induction, abduction, focus-group conversations, dialectics, hermeneutics, and, particularly a la Foucault (1973: x-xxiv), genealogy and even archaeology.

REFERENCES:

Drummond, Chris (2009) "Replicability is not Reproducibility: Nor is it Good Science." Proceedings of the Evaluation Methods for Machine Learning Workshop at the 26th ICML, Montreal, Canada

Firger, Jessica (August 28, 2015) "Science's Reproducibility Problem: 100 Psych Studies Were Tested and Only Half Held Up"  Newsweek

Foucault, Michel (1973) The Order of Things: An Archaeology of Human Sciences. New York: Vintage Books.

Leedy, Paul D. (1980) Practical Research: Planning and Design. Second Edition. New York:  Macmillan Publishing Co.,

Open Science Collaboration (2015) "Estimating the Reproducibility of Psychological Science"  Reproducibility Project: Psychology. University of Virginia

PSA - Psychological Science Agenda (September 2015) "Science Paper Shows Low Replicability of Psychology Studies" APA

Replicability Research Group (2015) "Replicability vs Reproducibility." Tel Aviv University, Department of Statistics and Operations Research

Seale, Clive (2004) "Replication/Replicability in Qualitative Research" in The SAGE Encyclopedia of of Social Science Research Methods (Michael S. ewis-Beck, Alan Bryman and Tim Futing Liao, Eds.)

van Rijn, Hedderik and Sabine Scholz (2014) "Replication of Experiment 3 of 'Tracing Attention and the Activation Flow in Spoken Word Planning Using Eye Movements' by A Roelofs (2008,


POSTSCRIPT:

All that I have said above is based on the assumption, not necessarily correct, that:
1. The original study to be subjected to replication provided all the details necessary to do so.
2. Each scientist involved in the replication exercise is/was fully competent to do so, and that the replicating work is itself fully open to replication.

It is also important to observe that the scientist seeking to do replication work may have to come to terms with three potential failure nodes:
1. The failure to replicate a study whose methodological procedures and related tools were not detailed enough to permit faithful or full retracing.
2. The failure to replicate because the researcher did not quite have the requisite competency or capacity -- intellectual, resource, temporal -- to replicate.
3. The failure to replicate because the original work is/was of such a type (partially or totally qualitative, for example) that it is/was not fully, adequately or in any other meaningful way replicable.


END NOTE: Paper updated October 19-20, 2015