Showing posts with label Leedy. Show all posts
Showing posts with label Leedy. Show all posts

Sunday, November 08, 2015

Measurability: A Key Standard of Scientific Research

Nine key standards, or principles, of scientific research can be gleaned from texts on research method. As I see them, these are:


3. Control
4. Measurability
5. Validity
6. Reliability (Pending)
7. Objectivity (Pending)
8. Ethics 
9. Representativeness.

I say something about each of the nine standards in separate posts, as shown in the links above. What I want to briefly talk about here is the measurability standard. 

Measurability 

Measurability means susceptibility to measurement. It means, in other words, that what is to be measured must embody quantities and/or qualities -- or have features, dimensions or characteristics -- that can be subjected to appropriate or 'true' measurement. Concerning the defining characteristics of a measurement, Bell (1999) notes that:
"A measurement tells us about a property of something. It might tell us how heavy an object is, or how hot, or how long it is. A measurement gives a number to that property. Measurements are always made using an instrument of some kind. Rulers, stopwatches, weighing scales, and thermometers are all measuring instruments. The result of a measurement is normally in two parts: a number and a unit of measurement, e.g. ‘How long is it?... 2 metres." 
The measurability rule is anchored on the above conceptions, and so requires that the variables around which the researcher intends to collect data should be measurable, or susceptible to acceptable ‘measurement’ (Leedy, 1980: 46)[1]. This is easier done in the natural sciences than in the social sciences; in quantitative studies than in qualitative. Still, one must, in the social sciences too, endeavor to quantify, measure and evaluate. Indeed, the guiding principle of the measurability rule -- its corollary, in other words -- is this: "What can be measured must be measured." Thus, not measuring what can be measured is not an option allowed anyone.

In quantitative research, the measurability standard features prominently in all methodological procedures that revolve around, or build up to, hypothesis testing. The testability of a hypothesis is indeed a function of the measurability of the variables that constitute it. Such measurability is, in turn, dependent on the indicators chosen to represent them; but one is not entirely free to choose just any indicator(s). They must be such as most peers or reviewers or supervisors or sponsors can be persuaded to accept as valid and reliable. Still, one has some leeway in choosing the indicators one wishes to use, so long as the reviewers, supervisors or funding entities are prepared to see or accept one's findings “in the light" of the chosen indicators.

Viswanathan (2010: 285-210) has interrogated in depth the challenges encountered in "Measure development procedures," particularly in connection with quantitatively-oriented social science research. Among these challenges, which he believes are not sufficiently appreciated, is the underpinning assumption, which he faults, that "a construct can be [readily and invariably] isolated and examined." Here is his line of thinking:
"By measuring or manipulating individual constructs, relationships between constructs are studied and substantive hypotheses about these relationships are tested. The very notion that numbers can be assigned to attributes of people, objects, or events is presumed on being able to study attributes or constructs separate from other constructs. After all, measurement relates to rules for assigning numbers to attributes of people, objects and events. [But] this is not the case for many phenomena. A complex network of constructs may influence a phenomenon and may not be separable into individual constructs for purposes of measurement." 
Measurability procedures implemented in qualitative -- that is, 'non-numerical' -- research attempt, not always successfully, to overcome or circumvent such intractables as Viswanathan has pointed out. These procedures include, for example, “comparative judgement (arranging factors in a hierarchy of importance)” or  “scaling (correct versus incorrect responses to a given set of questions)” (Leedy, 1980: 46). Other ‘non-numerical’ strategies that have been used, according to Worsley (1992: 113)[2] include: “careful classification[3], organizing and combining of field-notes [which] requires patient checking and cross-checking and the application of systematic techniques.” But he acknowledges that “numerical methods are sometimes applied to field-data too” (Worsley, 1992: 113). Indicators are often used as proxies/surrogates for more fuzzy or abstract concepts, such as status (Worsley, 1992: 96). Such indicators are inevitably numerical. 

One may also, in connection with the nominal level of measurement (the simplest level), and even with the ordinal level of measurement, calculate frequencies/percentages, and cross-classify in qualitative research. Scaling, moreover, is not just about correct v. incorrect. Incorporating comparative judgement, one may also use scales of: high v. low, do v. don’t, very high to very low, very good to very bad, strongly agree to strongly disagree.

To be fair to everyone, no one is quite attempting stubbornly to square the circle in this measurability conversation. Still, the challenges of measuring require, and will surely be tempered with, constant and open-minded vigilance -- as well as an abundance of capacity to deal with critiques and contrary views, including (so be it) one's own self-criticism. Bell (1999: 1), like others before her, subsumes this conversation under the theme "uncertainty of measurement". Thus:
"Uncertainty of measurement is the doubt that exists about the result of any measurement. You might think that well-made rulers, clocks and thermometers should be trustworthy, and give the right answers. But for every measurement - even the most careful - there is always a margin of doubt. In everyday speech, this might be expressed as ‘give or take’ ... e.g. a stick might be two metres long ‘give or take a centimetre’."



[1] Sarantakos talks of the need for precision in measurement (pp. 18-26)
[2] Peter Worsley, ed. 1992. The New Introducing Sociology. Revised Third Edition. London: Penguin Books.
[3] Another word for classification, given by Worsley (1992: 97) is categorization; and both, he argues, are an often overlooked “form of measurement” – even if they are “the simplest level” (that is, the nominal level) of measurement. There are in total four levels of measurement, the others being, in an ascending order of complexity: Ordinal (or rank-order), interval and ‘ratio’ scale (Worsley, 1992: 97-98).



REFERENCES

Bell, Stephanie (1999) Measurement Good Practice Guide: A Beginner's manual to Uncertainty of Measurement. (No. 11, Issue 2). Taddington: National Physics Laboratory

Leedy, Paul D. (1980) Practical Research: Planning and Design. Second Edition. New York:  Macmillan Publishing Co.,

Sarantakos, Sartirios (1994) Social Research. London: The Macmillan Press

Viswanathan, Mathu (2010) "Understanding the Intangibles of Measurement in the Social Sciences," pp. 285-312, in Geoffrey Walford, Eric Tucker and Mathu Viswanathan, Eds. (2010) The SAGE Handbook of Measurement. Los Angeles: SAGE 

Worsley, Peter, ed (1992) The New Introducing Sociology. Revised Third Edition. London: Penguin Books


Saturday, November 07, 2015

Universality: A Key Standard of Scientific Research

Nine key standards, or principles, of scientific research can be gleaned from texts on research method. As I see them, these are:

1. Universality
3. Control
4. Measurability
5. Validity
6. Reliability (Pending)
7. Objectivity (Pending)
8. Ethics 
9. Representativeness.

I say something about each of the nine standards in separate posts, as shown in the links above. What I want to briefly talk about here is the universality standard. 

Universality: 

The tern universality has roots in the broader conception of a pervasive universe, of which we are all inescapably a part. More specifically, it is anchored in the related notion of universals. Thus:
"Universals are features (e.g., redness or tallness) shared by many individuals, each of which is said to instantiate or exemplify the universal...The metaphysical issue is whether or not these features exist independently of the particular things that have them:  realists hold that they do; nominalists hold that they do not; conceptualists hold that they do so only mentally"(A Dictionary of Philosophical terms, in Garth Kemerling (2011) The Philosophy Pages. available online).
Thus, moreover, "evolutionary universals", as Talcott Parsons argues, are those similitudes, those patterns of resemblance, which scholarship detects -- recognizes -- in the evolution of human society.

Armstrong (1986, 1989) understands that universals do have particulars, the very constituents or components which make universals possible in the first place. Moreover, the particulars of any one universal resemble one another. As resembling constituents of universals multiply within and across universals, what we end up with is identity, which, to repeat, is borne out of resemblance, and which in the end it morphs out of. And so, one can see, universals are made of common particulars.Thus:
"If we consider ordinary, first-order, particulars, then . . . two things, while remaining two, can resemble exactly. At least exact resemblance is possible (assuming that the Identity of Indiscernibles is not a necessary truth). In the limit, resemblance of particulars does not give identity. But now consider the resemblance of universals. As resemblance of properties [monadic universals] gets closer and closer, we arrive in the limit at identity. Two become one. This suggests that as resemblance gets closer, more and more constituents of the resembling properties are identical, until all the constituents are identical and we have identity rather than resemblance." [This quote is from Armstrung (1989) as found in Pautz (1977). The original article is yet to be accessed]  
So universality is ultimately "standardized" via the metamorphosis of its resembling parts into a more holistic identity recognized by the, or at the very least a, scientific community. The universality standard requires that any research project should be designed and planned in such a way that any competent researcher, not just the one(s) who wrote the proposal,  should be able to successfully undertake it. (Leedy, 1980: 46). In this sense, the research project has a life of its own, independent of any particular researcher, within the confines of the relevant discipline or scientific community.

Universality demands a high level of discipline and transparency in the research habits of all who do research. In a sense, too, universality makes all scientific discoveries part of a common, with standard operating procedures and the common ownership of the discoveries claimed and accumulated across geographies, disciplines and time.

READ: Universal (metaphysics)

ALSO READ: Professor JeeLou Lin "D.M. Armstrong, Universals: An Opinionated Introduction"


To partially sum it all up, let's see in this quotation how the International Council for Science (ICSU) defines the universality of Science:
"The universality of science in its broadest sense is about developing a truly global scientific community on the basis of equity and non-discrimination. It is also about ensuring that science is trusted and valued by societies across the world. As such, it incorporates issues related to the conduct of science; capacity building; science education and literacy; access to data and information and the relationship between science and society..." (see ICSU)
Furthermore, a noteworthy point -- on the importance of doubt and criticism (including self-criticism) in all claims to, or efforts to affirm, scientific discovery -- is made by Michel Paty (2001: 8) who, inspired by Descartes's Discourse on Method, argues that:
"...the idea of of universality, as well as ideas of reason and of demonstrative (and even objective) science, with which it has a constitutive link, carries with it the requirement of its own criticism...[Thus]... the only true knowledge is that knowledge that, for every thinking subject, overcomes the obstacles [posed] by doubt."

REFERENCES:
Armstrong, D.M. (1986) "In Defense of Structural Universals", Australasian Journal of Philosophy 64 (1986) pp. 85-88.

Leedy, Paul D. (1980) Practical Research: Planning and Design. Second Edition. New York:  Macmillan Publishing Co.,


Parsons, Talcott (1964) "Evolutionary Universals in Society" American Sociological Review, Vol. 29, Issue 3 (June, 1964), pp. 339-357


Paty, Michel (2001). "Universality of Science : Historical Validation of a Philosophical Idea." in Habib, S. Irfan and Raina, Dhruv. Situating the history of science : Dialogues with Joseph Needham, Oxford India Paperbacks, p. 303-324, 2001.  [see p. 7 of the pdf version given in this link]

Pautz, Adam (1977) "An Argument Against Armstrong's Analysis of the resemblance of Universals"




~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
CSO 302

Friday, October 30, 2015

Control: A Key Standard of Scientific Method


Nine key standards, or principles, of scientific research can be gleaned from texts on research method. As I see them, these are:

3. Control
4. Measurability
5. Validity
6. Reliability (Pending)
7. Objectivity (Pending)
8. Ethics 
9. Representativeness.

I say something about each of the nine standards in other posts. What I want to briefly talk about here is the control standard. 

Control

By control we mean establishing or maintaining the integrity or identity or boundaries of the sample/specimen/subject/site against all disturbance, unwanted/unintended influence or “contamination” by external/extraneous factors or even by the researcher himself or herself. In effect, control means quality control. 

In experimental design, Yu (n.d.) points out, the kind of control called for is, specifically, variance control  -- which is indeed a form of quality control. According to Leedy, control suggests establishing parameters: “Parameters are important. All research is conducted within an area sealed off by given parametric limitations. By such control, we isolate those factors which are critical to the research. Control is important for replication. An experiment should be repeated under the identical conditions and in the identical way in which it was first carried out. It is also important for consistency with the research design” (Leedy, 1980: 46).

Giorgio (n.d.) notes that, in the laboratory setting, the control function tends to be too narrowly carried out in practice, in a way that even conceptually distorts the meaning of control. Thus:
"...control practically comes to mean manipulate. That is because the laboratory setting is itself the creation of the scientist, and it is deliberately created in order for the researcher to be able to dictate to the situation. Any thing that is left to chance is seen as a demerit. The dictionary con firms these ideas by defining control as exercising “authority or dominating influence over; to hold in restraint or check; to direct or regulate.” Thus a strong sense of control means being able to do things at will, to be able to have desires transformed into realities in as brief a time as it takes to accomplish the task."
But, he argues, "control does not have to be defined so narrowly." Thus, for him, control:
"...could mean to be respectfully present in a situation so that one could respond knowledgeably and appropriately with readiness to the perceived spontaneous happenings of a situation. In this sense, control refers to a certain kind of “intentional presence,” a certain way of looking, an informed being present that at the same time does not actively intervene in the unfolding of the event to be observed... [The] knowledge that the researcher has guides his expectations so that he knows how to interpret or evaluate his ongoing perceptions. The researcher can do this without active intervention, but it is still a form of control because directedness and restraint are used by the researcher in perceiving the situation. Thus control implies observation, that is, it is an educated looking."
Admittedly, in the social sciences, in contrast to the natural sciences, control is more varied and more nuanced, and sometimes more problematic, but not impossible. For example, determining the parameters of, or controlling for variance in, a social research includes such considerations as scope and study limitations. Moreover, one can ensure control over the following, among others: [1] the setting of the interview, or the conditions under which the respondent is interviewed, [2] the accuracy/precision of the questions asked and the responses given. [3] the language of the interview. [4] the quality of training given to interviewers, other research assistants and supervisors. [5] the actual supervision process. [6] the parameters of both the experimental and the control group.

Archaeologists and paleontologists, who do not represent the image of the typical natural scientist, and who routinely work away from the public view, offer outstanding templates for quality control in their persistent quest for new discoveries. They are quite adept at protecting the sites of their ongoing work against contamination and interference, in time and space. In important ways, the work of the crime (or accident) investigator, often partially done in full view of the media and the public, at its best obeys the strictures of the control standard,

Clearly, controls are crucial not just in experimental designs and field observations, but also, and perhaps more dramatically so, in all situations in which the conscious and pre-emptive gathering and establishing of incontrovertible evidence is called for; for example, in the prosecutor's mandate, or in the context of a claim to/of the discovery of a consequential artifact. Both of which are typically subjected to strict and detailed scrutiny, skepticism and even attack.



REFERENCES:

Leedy, Paul D. (1980) Practical Research: Planning and Design. Second Edition. New York:  Macmillan Publishing Co.,

Giorgi, Amedeo (n.d.) "The Role of Observation and Control in Laboratory and Field Research Settings"

Jaikumar, Maheswari (2014) "Experimental Research Design" Slideshare.net

Mann, CJ (2003) "Observational Research Methods. Research Design II: Cohort, Cross Sectional and Case-Control Studies"

Yu, Chong-Ho (n.d.) "Experimental Design as Variance Control"

Sunday, October 18, 2015

Replicability: A Key Standard of Scientific Method Now Found to be Under Existential Threat

Nine key standards, or principles, of scientific research can be gleaned from texts on research method. As I see them, these are:
2. Replicability
3. Control
4. Measurability
5. Validity
6. Reliability (Pending)
7. Objectivity (Pending)
8. Ethics
9. Representativeness.

I will say something about each of the nine standards in later posts. What I want to briefly talk about today is replicability. It turns out that, highly valued as it is, replicability is easier 'said' in conversations about methodology -- defined as the science or theory of method -- than 'done' in actual empirical research. This is what we pick from a study captured in Jessica Firger's account, the link to which I give in the References section below.

First, though, let us recapitulate the accepted meaning and procedures of replicability. Then, second, we will turn to the new study to see how its findings shake the foundations of the scientist's faith (now 'blind faith'?) in the utility and robustness of replicabiliy as a verification process.

First, then, about the centrality of replication to scientific method and to science's "theory of evidence": As a concept of methodology, replicability calls for precision in measurement and accuracy of procedure in all scientific endeavors. That is to say, it calls for the doing of any research project in such a way as to ensure its step-by-step, beginning-to-end, repeatability or 'reproducibility' -- which is what replicability means -- in any subsequent study, by any competent researcher, intended to attest to the veracity of its discoveries or claims. To repeat, it should be repeatable in its exact original form, and should yield the very same results. Thus:
“The research should be repeatable. Any other competent researcher should be able to take your problem and , collecting data under the same circumstances and within the identical parameters as you have observed, achieve results comparable (sic) to those you have been able to secure” (Leedy, 1980: 46; see also Seale, 2004). 
That emphasis on "the very same results", as opposed to 'comparable' results, highlights the ultimate basis for the confirmation or validation of scientific claims. However, the confounding truth, which subverts such a neatly articulated metric, is that several important factors may be suggested as underpinning the likelihood of not achieving such results in the real world which social science, including/particularly sociology, routinely deals with. This is the real world of conscious, querying, skeptical (and often over-researched) human beings unsure of the risks, or real value, of availing themselves and telling everything that researchers wish to be told. And that's even before one factors in variability, for various reasons, in the ways respected scientists implement even widely agreed-upon procedures -- or the procedures that should, in the name of reason and perhaps common sense, be so agreed -- for data capture, processing, analysis and synthesis.

But why all this fuss about replicability or reproducibility to which I believe all seasoned interrogators of methodology have been paying attention for years? Firger, I think, captures the underlying sentiment as well as anyone can or has:
"Most rational humans hold some faith in science. We trust that scientists, through their hard-won expertise, are well equipped to conduct studies that provide proof of why things are the way they are, help solve problems and explain mysteries. Much of that trust is built on an implicit belief that research findings are concrete truths—in other words, that if given the same set of parameters, they would be easy to reproduce. This type of replication is essential to science because it validates key discoveries and helps scientists make progress in their fields of research" (Firger, 2015).
So replicability is a standard worth respecting, implementing and defending. But what if it turns out that in practice we cannot, or have not been doing so? What are we to make of our scientific truths, then? That, at least implicitly, is for me the most fundamental, if implicit, question driving the University of Virginia study, as we will now see.

Second, then: what was the purpose of the new study and what does it tell us? We can discern that purpose and related method from this portion of the study's problem statement, as seen in the abstract:
"Reproducibility is a defining feature of science, but the extent to which it characterizes current research is unknown. We conducted replications of 100 experimental and correlational studies published in three psychology journals using high-powered designs and original materials when available"  (Open Science Collaboration, 2015)
What about the results? The key finding of the study was that the results of less than 50% of the studies subjected to replication were corroborated; that is, independently validated. That is a huge shortfall, as elaborated below:
"Replication effects (Mr = .197, SD = .257) were half the magnitude of original effects (Mr = .403, SD = .188), representing a substantial decline. Ninety-seven percent of original studies had significant results (p < .05). Thirty-six percent of replications had significant results; 47% of original effect sizes were in the 95% confidence interval of the replication effect size; 39% of effects were subjectively rated to have replicated the original result; and, if no bias in original results is assumed, combining original and replication results left 68% with significant effects" (Open Source Collaboration, 2015; see also PSA, September 2015).
Firger (2015) notes problem in this regard: that the threat to replicability may be exacerbated by the existential pressure, perceived by scientists seeking recognition and success (presumably via promotion and tenure, or higher levels of funding, or 'all of the above'), to dump rigor, and all its 'yokes' and promise, in favor of building the capacity to "weave a memorable story out of tenuous science" -- and so to achieve such success.

In the end, perfect replicability, certainly of studies involving conscious human subjects, may not be feasible, even when driven purely by the dictates of scientific rigor, ahead of the capacity for time-travel back to the chronological 'moment' at which the original study of interest was conducted. Incidentally, contrary to Drummond's (2009) claim, this is not to say that 'reproducibility' is a more viable posture for the scientist. The attributes he assigns to reproducibility are in fact, precisely, those of replicability, as widely understood by those who routinely use the latter term (see, for example, the clarification offered by the Replicability Research Group, 2015). What Drummond wrongly sees as the weaknesses of replicability are its ideals and strengths, and what he touts as the strengths of reproducibility are its weaknesses. Reproducibility, defined as he so facilely does, would be a recipe for widespread intellectual fraud.

So what is one to do in the meantime? We will just have to make do with such rigors and caveats as we are able to muster and agree upon. But the rationale for replicating scientific work remains as great and as persuasive as it ever was, and its utility for science and human progress just as enormous. All that we have discovered is its disappointingly low success score, certainly in psychology -- which is a human science. We have also found, alas, that the pressure for recognition and success makes the scientist as much an opportunist as the proverbial self-sacrificing servant of truth. We can work with sharp focus on these personal/human 'frailties' of the scientist -- not necessarily of science -- to raise replicability's score. Replication is (natural) science's ultimate audit tool, which we will abandon only at society's and civilization's own great peril.

Still, we will also have to more consciously and more robustly develop other methods of arriving at the truth -- and the natural scientist, in particular, will have to respect with greater humility and sharper awareness of that potentially debilitating soft underbelly, just discovered, of his/her cherished branch of science. The other methods I have in mind are those which social science and the humanities have been grappling with, self-critically (and against the backdrop of harsh and even dismissive criticism by often sanctimonious natural scientists), for decades, and even more than a century, now. Many of these fall under the qualitative label, as opposed to the quantitative, and include versions of such broad and cross-cutting, timeless and a-disciplinary methods as: observation (naturalistic, participant or device-assisted), comparison, deduction, induction, abduction, focus-group conversations, dialectics, hermeneutics, and, particularly a la Foucault (1973: x-xxiv), genealogy and even archaeology.

REFERENCES:

Drummond, Chris (2009) "Replicability is not Reproducibility: Nor is it Good Science." Proceedings of the Evaluation Methods for Machine Learning Workshop at the 26th ICML, Montreal, Canada

Firger, Jessica (August 28, 2015) "Science's Reproducibility Problem: 100 Psych Studies Were Tested and Only Half Held Up"  Newsweek

Foucault, Michel (1973) The Order of Things: An Archaeology of Human Sciences. New York: Vintage Books.

Leedy, Paul D. (1980) Practical Research: Planning and Design. Second Edition. New York:  Macmillan Publishing Co.,

Open Science Collaboration (2015) "Estimating the Reproducibility of Psychological Science"  Reproducibility Project: Psychology. University of Virginia

PSA - Psychological Science Agenda (September 2015) "Science Paper Shows Low Replicability of Psychology Studies" APA

Replicability Research Group (2015) "Replicability vs Reproducibility." Tel Aviv University, Department of Statistics and Operations Research

Seale, Clive (2004) "Replication/Replicability in Qualitative Research" in The SAGE Encyclopedia of of Social Science Research Methods (Michael S. ewis-Beck, Alan Bryman and Tim Futing Liao, Eds.)

van Rijn, Hedderik and Sabine Scholz (2014) "Replication of Experiment 3 of 'Tracing Attention and the Activation Flow in Spoken Word Planning Using Eye Movements' by A Roelofs (2008,


POSTSCRIPT:

All that I have said above is based on the assumption, not necessarily correct, that:
1. The original study to be subjected to replication provided all the details necessary to do so.
2. Each scientist involved in the replication exercise is/was fully competent to do so, and that the replicating work is itself fully open to replication.

It is also important to observe that the scientist seeking to do replication work may have to come to terms with three potential failure nodes:
1. The failure to replicate a study whose methodological procedures and related tools were not detailed enough to permit faithful or full retracing.
2. The failure to replicate because the researcher did not quite have the requisite competency or capacity -- intellectual, resource, temporal -- to replicate.
3. The failure to replicate because the original work is/was of such a type (partially or totally qualitative, for example) that it is/was not fully, adequately or in any other meaningful way replicable.


END NOTE: Paper updated October 19-20, 2015