Showing posts with label embarrassment. Show all posts
Showing posts with label embarrassment. Show all posts

Tuesday, March 22, 2011

“Charles Darwin” of intelligent design condones plagiarism at Baylor?

Last year, “intelligent design” creationist William A. Dembski, Research Professor of Philosophy at Southwestern Baptist Theological Seminary, said of one of his coauthors at Baylor University, “[Student] — With pro-ID graduate students like this, Darwinian profs don't stand a chance.” Well, [the student] evidently went on to defend a master’s thesis, Studies of Active Information in Search, in the Fall semester of 2010. I downloaded it from Baylor’s document archive, and had only to read to the bottom of the first page to discover plagiarism.*

Almost half of the first chapter is copied, without citation or sign of quotation, from

[Edit 3/24: Actually, at least half is copied, as you easily can see in my markup of the text.] According to the signature page, Marks, who holds the rank of Distinguished Professor of Engineering at Baylor, served on the thesis committee. Yes, I am referring to one of the twenty most influential Christian scholars — the “Charles Darwin” of intelligent design.

Most of the thesis is copied from two published conference papers, and from an article that had been accepted for publication:

  1. [Student], William A. Dembski and Robert J. Marks II, Evolutionary Synthesis of Nand Logic: Dissecting a Digital Organism, 2009
  2. [Student], George Montañez, William A. Dembski, Robert J. Marks II, Efficient Per Query Information Extraction from a Hamming Oracle, 2010
  3. George Montañez, [Student], William A. Dembski, Robert J. Marks II, A Vivisection of the ev Computer Organism: Identifying Sources of Active Information, accepted August 27, 2010; published December 15, 2010
Not only are these papers not cited, but they also are not listed in the bibliography.

Chapter 2 is essentially (2), beginning with the fourth paragraph, and skipping Section II. (The first three paragraphs are in Chapter 1.) It corrects errors that I identified here at Bounded Science. More than half of the paper is mathematical analysis that is beyond most master’s students in computer science. I would guess that [the student’s] contributions were programming, data gathering and visualization, and paper preparation. If I am right, then it is inappropriate for the thesis to give the impression that [the student] did the math.

Chapter 4 is (1) with Section I eliminated, and with several small changes. It haplessly ends with “we have not explored in this paper.” (For those of you who are not scholars, I should explain that theses are not referred to as papers.)

Chapter 3 is not so simply related to its source, presumably because its author is not the lead author of (2). Almost all of its text is present in the article. But the chapter reorders passages of the article. Most notably, it moves empirical results ahead of theoretical analysis, and introduces awkward forward references to the analysis. Also, there are occasional replacements of words with synonyms, as well as deletions and insertions of short phrases. All in all, the chapter looks like the work of someone who was trying, pathetically, to make copy-and-paste pass for original writing. (Seventeen years of teaching experience are talking here.)

Chapter 5 is a double-spaced, one-page conclusion.

Some universities limit how much of a thesis may come from published work, but I can find no indication that Baylor is one of them. Yet the thesis does not state that the chapters are excerpted from published and forthcoming papers. Instead there is an appendix entitled ”Copyrights” that gives, without explanation, copyright release forms bearing the titles, but not the complete lists of authors (required by the publisher), of (1) and (2). Guess which author is missing? Yes, that would be William A. Dembski. I suspect that the forms on file with the publisher bear his name.

To get a hint that Dembski deserves credit for contributions to the thesis, you have to click the Show full item record button of the entry for the thesis in the BEARdocs system, and then figure out the meaning of the identifier.citation fields providing full citations of (1) and (2). In my opinion, burying this information in the metadata of the document retrieval system is unethical.

The copyright release forms indicate that the authors retain the right to use the published material in derivative works, “provided that the source and any IEEE [Institute of Electrical and Electronic Engineers] copyright notice are indicated.” So the thesis perhaps does not violate the copyrights of the publisher. But there is much more to academic integrity than not breaking the law. When you draw on a source, you cite it, even if you are one of the authors. When you copy from a source, you indicate clearly what you are copying. It would have been so easy to end the introduction of the thesis with an indication that Chapter 2 is excerpted, with emendations, from (2), and that Chapter 4 is excerpted from (1). As for Chapter 3, the author should have written it from scratch, and should have cited (3) in all places where the work was not his own.

Now, should you believe that self-citation and self-quotation are optional in scholarly writing, go back to the beginning of this post. It is unethical for a thesis committee member to condone plagiarism, even if it is plagiarism of his own work. The unacknowledged use of (3) in the thesis is egregious. It is clearly wrong to draw on the work of others, and not indicate that you are doing so. (For those of you who are not scholars, I should mention that references to forthcoming publications are common.) There is no wiggle room here. If the document is authentic, then both [the student] and Marks are out-of-bounds ethically. And one really must wonder what was going on with the chairman of the thesis committee, associate professor of computer science Greg Hamerly. Was Marks the de facto chairman? Had Hamerly even bothered to read the handful of papers on active information? If he had, then he is complicit. If he had not, then his performance was shabby, to say the least.

The contact information below should come in handy for any journalist who wants to work on the story. And I encourage readers to let Baylor administrators know what a blight on the reputation of the school the thesis is. Click here now to open your default email application and address the dean of the graduate school, J. Larry Lyon. You will CC the executive vice president and provost, the dean of engineering and computer science, and the chairman of the computer science department.


Thesis author:
[Student]

Thesis committee member:
Robert J. Marks II, Ph.D.
Distinguished Professor of Engineering
(254) 710-7302
Robert_Marks@baylor.edu

William A. Dembski, Ph.D.
Research Professor of Philosophy and Director of the Center for Cultural Engagement
Southwestern Baptist Theological Seminary
817-923-1921 ext.4435
wdembski@designinference.com

J. Larry Lyon, Ph.D.
Dean of the Graduate School
Baylor University
(254) 710-3588
Larry_Lyon@baylor.edu

Elizabeth Davis, Ph.D.
Executive Vice President and Provost
Baylor University
(254) 710-7803
Elizabeth_Davis@baylor.edu

Benjamin S. Kelley, Ph.D., P.E.
Dean of Engineering and Computer Science
Baylor University
(254) 710-3871
Ben_Kelley@baylor.edu

Signed to approve thesis as department chairman:
Donald L. Gaitros, Ph.D.
Professor and Chairman of Computer Science
Baylor University
(254) 710-3876
Don_Gaitros@baylor.edu

Thesis committee chairperson:
Greg Hamerly, Ph.D.
Associate Professor of Computer Science
Baylor University
(254) 710-6846
hamerly@cs.baylor.edu

Thesis committee member:
[completed Ph.D. and joined Baylor in 2009]
Young-Rae Cho, Ph.D.
Assistant Professor of Computer Science
Baylor University
(254) 710-3385
Young-Rae_Cho@baylor.edu


* I emphasize that my remarks are contingent on the authenticity of the document residing in Baylor’s BEARdocs archive on March 19, 2011. When possible, I assess the thesis I retrieved, and not persons. I have verified that Baylor requires entry of theses into BEARdocs. A metadatum indicates that the thesis was added on January 5, 2011, which is consistent with December 2010 graduation. I contacted [the student] by email to ask if the document were the final draft of his thesis, and he responded, “I have not verified the document, but it should be the final draft.”

Thursday, July 29, 2010

Feeling charitable toward Baylor’s IDC cubs

The reason I come off as a nasty bastard on this blog is that I harbor quite a bit of anger toward the creationist bastards who duped me as a teenager. The earliest stage of overcoming my upbringing was the worst time of my life. I wanted to die. Consequently, I am deadly serious in my opposition to “science-done-right proves the Bible true” mythology. William A. Dembski provokes me especially with his prevarication and manipulation. He evidently believes that such behavior is moral if it serves higher ends in the “culture war.” My take is, shall we say, more traditional.

When Robert J. Marks II, Distinguished Professor of Engineering at Baylor University, and Fellow of the Institute of Electrical and Electronics Engineers (IEEE), began collaborating with Dembski, I did not rush to the conclusion that he was like Dembski. But it has become apparent that he is willing to play the system. For instance, Dembski was miraculously elevated to the rank of Senior Member of the IEEE, which only 5% of members ever reach, in the very year that he joined the organization. To be considered for elevation, a member must be nominated by a fellow.

Although Marks was the founding president of the progenitor of the IEEE Computational Intelligence Society, which addresses evolutionary computation (EC), he and his IDCist collaborators go to the IEEE Systems, Man, and Cybernetics Society for publication. He is fully aware that reviewers there are unlikely to know much about EC, and are likely to give the benefit of the doubt to a paper bearing his name. I would love to see him impugn the integrity of his and my colleagues in the Computational Intelligence Society by claiming that they don’t review controversial work fairly. But it ain’t gonna happen.

I’ve come to see Marks as the quintessential late-career jerk, altogether too ready to claim expertise in an area he has never engaged vigorously. He is so cocksure as to publish a work of apologetics with the title Evolutionary Computation: A Perpetual Motion Machine for Design Information? (Chap. 17 of Evidence for God, M. Licona and W. A. Dembski, eds.). He states outright some misapprehensions that are implicit in his technical publications. Here’s the whopper: “A common structure in evolutionary search is an imposed fitness function, wherein the merit of a design for each set of parameters is assigned a number.” Who are you, Bob Marks, to say what is common and what is not in a literature you do not follow? Having scrutinized over a thousand papers in EC, and perused many more, I say that you are flat-out wrong. There’s usually a natural, not imposed, sense in which some solutions are better than others. Put up the references, Distinguished Professor Expert, or shut up.

Marks and coauthors cagily avoid scrutiny of their (few) EC sources by dumping on the reviewers references to entire books, i.e., with no mention of specific pages or chapters. This is because their EC veneer will not withstand a scratch. The chapter I just linked to may seem to contradict that, given its references to early work in EC by Barricelli (1962), Crosby (1967), and Bremmerman [sic] et al. (1966). [That's Hans-Joachim Bremermann.] First, note the superficiality of the references. Marks did not survey the literature to come by them. The papers appear in a collection of reprints edited by David Fogel, Evolutionary Computation: The Fossil Record (IEEE Press, 1998). Marks served as a technical editor of the volume, just as I did, and he should have cited it.

Although Marks is an electrical engineer, he has been working with two of Baylor’s graduate students in computer science, Winston Ewert and George Montañez. I would hazard a guess that there is some arrangement for the students to turn their research with Marks into masters’ theses. I’ve been sitting on some errors in their most recent publication, Efficient Per Query Information Extraction from a Hamming Oracle, thinking that the IDC cubs would get what they deserved if they included the errors in their theses. Well, I’ve got a soft spot for students, and I’m feeling charitable today. But there’s no free lunch for Marks. He has no business directing research in EC, his reputation in computational intelligence notwithstanding, and I hope that the CS faculty at Baylor catch on to the fact.

First reading

On first reading the paper, I was deeply annoyed by the combination of a Chatty-Cathy, self-reference-laden introduction focusing on “oracles,” irrelevant to the majority of the paper, with a non-survey of the relevant literature in the theory of EC. Ewert et al. dump in three references to books, without discussion of their content, at the beginning of their 4-1/2 page section giving Markov-chain analyses of evolutionary algorithms. It turns out that one of the books does not treat EC at all — I contacted the author to make sure.

As I have discussed here and here, two of the algorithms they analyze are abstracted from defective Weasel programs that Dawkins supposedly used in the mid-1980's. It offends me to see these whirlygigs passed off as objects worthy of analysis in the engineering literature.

Yet again, they express the so-called average active information per query as $$I_\oplus = {{I_\Omega} \over Q} = \frac{\log N^L}{Q} = {{L \log N} \over Q},$$ where Q is not the simple random variable it appears to be, but is instead the expected number of trials (“queries”) a procedure requires to maximize the number of characters in a “test” string that match a “target” string. Strings are over an alphabet of size N, and are of length L. Unless you have something to hide, you write $$I_\oplus ={{L \log N} \over {E[T]}},$$ where T is the random number of trials required to obtain a perfect match of the target. This is a strange idea of an average, and it appears that a reviewer said as much. Rather than acknowledge the weirdness overtly, Ewert et al. added a cute “yeah, we know, but we do it consistently” footnote. Anyone without a prior commitment to advancing “intelligence creates active information” ideology would simply flip the fraction over to get the average number of trials per bit of endogenous information IΩ, $$\frac{1}{I_\oplus} = E\left[{T \over {I_\Omega}}\right] = {{E[T]} \over {L \log N}}.$$ This has a clear interpretation as expected performance normalized by a measure of problem hardness. But when it’s “active information or bust,” you’re not free to go in any sensible direction available to you. I have to add that I can’t make a sensible connection between average active information per query and active information. Given a bound K on the number of trials to match the target string, the active information is $$I_+ = \log \Pr\{T \leq K\} + {L \log N}.$$ Do you see a relationship between I+ and I that I’m missing?

By the way, I happened upon prior work regarding the amount of information required to solve a problem. The scholarly lassitude of the IDC “maverick geniuses” glares out yet again.

Second reading

On second reading, I bothered to do sanity checking of the plots. I saw immediately that the surfaces in Fig. 2 were falling off in the wrong directions. For fixed alphabet size N, the plots show the average active information per query increasing as the string length L increases, when it obviously should decrease. The problem is harder, not easier, when the target string is longer. Comparing Fig. 5 to Figs. 3 and 4, it’s easy to see that the subscripts for N and L are reversed somewhere. But what makes Fig. 3 cattywampus is not so simple. Ewert et al. plot $$I_\oplus(L, N) = \frac{L \log N}{E[T_{N,L}]}$$ instead of $$I_\oplus(L, N) = \frac{L \log N}{E[T_{L,N}]}.$$ That is, the matrix of expected numbers of trials to match the target string is transposed, but the matrix of endogenous information values is not.

The embarrassment here is not that the cubs got confused about indexing of square matrices of values, but that a team of four, including Marks and Dembski, shipped out the paper for review, and then submitted the final copy for publication, with nary a sanity check of the plots. From where I sit, it appears that Ewert and Montañez are getting more in the way of indoctrination than advisement from Marks and Dembski. Considering that various folks have pointed out errors in every paper that Marks and Dembski have coauthored, you’d think the two would give their new papers thorough goings-over.

It is sad that Ewert and Montañez probably know more about analysis of algorithms than Marks and Dembski do, and evidently are forgetting it. The fact is that $$E[T_{L,N}] = \Theta(N L \log L)$$ for all three of the evolutionary algorithms they consider, provided that parameters are set appropriately. It follows that $$I_\oplus = \Theta\left(\frac{L \log N}{N L \log L}\right) = \Theta\left(\frac{\log N}{N \log L}\right).$$ In the case of (C), the (1, λ) evolutionary algorithm, setting the mutation rate to 1 / L and the number of offspring λ to N ln L does the trick. (Do a lit review, cubs — Marks and Dembski will not.) From the perspective of a computer scientist, the differences in expected numbers of trials for the algorithms are not worth detailed consideration. This is yet another reason why the study is silly.

Methinks it is like the OneMax problem

The optimization (not search) problem addressed by Ewert et al. (and the Weasel program) is a straightforward generalization of a problem that has been studied heavily by theorists in evolutionary computation, OneMax. In the OneMax problem, the alphabet is {0, 1}, and the fitness function is the number of 1's in the string. In other words, the target string is 11…1. If the cubs poke around in the literature, they’ll find that Dembski and Marks reinvented the wheel with some of their analysis. That’s the charitable conclusion, anyway.

Winston Ewert and George Montañez, don’t say the big, bad evilutionist never gave you anything.