Tuesday, September 26, 2023

Retroaction is not quite a retraction

 In a 2021 post Cauchy priors I made a gigantic blunder in misinterpreting a twitter response to a question posed by Victor Chernozhukov about  the mean E[X|X+Y] when X is standard Gaussian, and Y is independent and standard Cauchy.  I reformulated this as:  suppose Y|T ~ N(T,1) and T ~ Cauchy, what is E(T|Y=y)?  This is a standard Bayesian problem with the idea of Cauchy priors going back to Jeffreys and explored more recently by Berger and others.  My blunder was elementary and involved failing to remember  that the normalizing factor for the conditional density was dependent on y.  When this was fixed, I get the figure below.  To accentuate the flat portion of the posterior mean I've reduced the scale of the Cauchy to be 0.1 rather than 1.  The interpretation of figure is quite intuitive:  when y is near zero and therefore in agreement with the prior, the posterior mean is aggressively shrunken toward zero.  However, when |y| is far from zero, the prior says, "well, that could happen" and the posterior eventually looks indistinguishable  from y.



There is a mildly amusing story associated with how I came to revisit this problem.  I have been reading a recent JPE paper A/B Testing with Fat Tails that employs Student t  priors with low degrees of freedom in an essential way.  Having totally forgotten about the previous blog post, I proceeded to investigate how to compute this posterior mean, and not surprisingly my initial attempts faltered a bit, so I started to google around to see what was "out there" in webland.  Early on I found a nice paper by Guy Nason that dealt with the case of Student on 3 dfs.  It mentioned that there was a 1939 David Kendall paper that treated the Cauchy case.  This must have been written when Kendall was still a grad student.  It involves some quite exotic complex analysis, and among others cites a 1935 paper by Robert Oppenheimer!  If I interpret Nason correctly, the Kendall paper produces a "closed form" expression for the marginal density of a Cauchy mixture of Gaussians.  Kendall comments rather drolly that the expression isn't useful for computations because there was no tabulated  version of the erfc function for complex arguments.  This lack has been rectified in the intervening years, but my attempts in R, and then in Mathematica to check that this Kendall's expression integrated to one failed.  Instead, the integral seemed to diverge slowly.  I would be grateful for any and all suggestions about this, but I rather expect that it is all lost in the mists of time.

Meanwhile, fortunately, it is easy to cook up a numerical version of the posterior mean solution that I will append here:

# Berger problem

s <- 0.1

f <- function(t,x) dnorm(x-t, sd = 1) * dt(t/s,1)/s

k <- function(x) integrate(f, -Inf,Inf, x = x)$value

g <-function(t,x) t * f(t,x)

h <- function(x) integrate(g, -Inf,Inf, x = x)$value

x <- -100:100/10

m <- x

for(i in 1:length(x)) m[i] <- h(x[i])/k(x[i])

png("Cauchy.png")

plot(x, m, type = "l", xlab = "y", ylab = expression(E~theta|Y==y))

abline(c(0,1),col = 2)

abline(h = 0,col = 2)

dev.off()



Wednesday, July 12, 2023

There is no discussable subject (of the first order)


 I joined Twitter in 2009 in the futile hope that it would lead me to a Korean Taco truck on my first ever visit to LA.  It didn't occur to me to tweet until 2021 when I decided that it was time to launch a quixotic attempt to get Bialetti to revive their legendary pasta machine.  This failed miserably too, although the Guardian food columnist Racheal Roddy was very nice about it.

Since then I've tweeted a few times always in response to something someone else had written.  This led me to wonder why I couldn't bring myself to originate a tweet.  The answer to this query appeared to me yesterday in the form of a talk delivered by Frank Ramsey in 1925 that appears as the Epilogue in the collection of Ramsey's papers edited by R.B. Braithwaite titled: "The Foundations of Mathematics"




Wednesday, December 7, 2022

 I've been reading about the Rasch model of item response in educational testing, in preparation for writing a brief section about it for the empirical Bayes book.  Eventually, I recalled that Edgeworth had an amusing paper about this sort of thing, from which I quote the final paragraph.


To examiners at least it will be interesting to test the accuracy of the instrument with which they work. The statistical study may beguile the monotony of their task. The " charm severe of numbers " is celebrated by Wordsworth as


                    "Especially perceived when nature droops 

                     And feeling is suppressed." 


The poet is evidently describing in prophetic words words the condition of examiners, and prescribing their solace. More tropically another inspired bard has indicated the paregoric use of an interest in statistics. In one of the beautiful pictures with which Homer has adorned the shield of Achilles, the ploughman of the good old times, as he finishes each furrow, and turns to begin a new one, is presented with a refreshing cup of honey-sweet wine. So they who plough in the modern metaphorical sense, may, in the pauses of their labours, be refreshed with the cup of statistical science, which I have endeavoured to sweeten. 


F.Y. Edgeworth (1890) The Element of Chance in Competitive Examinations, JRSS, 664-663.

Wednesday, August 10, 2022

Poetry makes nothing happen

 Peter Hull posted on twitter this fragment from a paper by Don Rubin  that perfectly encapsulates the W.H. Auden maxim:  Poetry makes nothing happen.  

http://www.asasrms.org/Proceedings/y1975/Bayesian%20Inference%20for%20Causality%20-%20The%20Importance%20of%20Randomization.pdf




Sunday, May 15, 2022

Almost a Haiku

" Good sense is dead, its child, science killed it to find out how it was made."

From the novel, Innocence by Penelope Fitzgerald, the phrase is attributed to Antonio Gramsci..


Innocence  is a truly brilliant novel, with a sensibility somewhere between Jane Austen and Henry James.  It is strange that someone, preferably Paolo Sorrentino, hasn't made a movie of this novel.

Tuesday, November 16, 2021

Hansen's Gauss-Markov Theorem


 

Edgeworth's 1920 paper "The Element of Chance in Competitive Examinations" mocks excessive reliance on "reasoning with the aid of the gens d'arme's hat -- from which as from a conjuror's, so much can be extracted."  In this spirit Bruce Hansen's recent paper, "A Modern Gauss-Markov Theorem" argues that the econometrics slogan  "OLS is BLUE" can be modified to "OLS is BUE", that is that we need not restrict attention to linear estimators, OLS can be considered minimum variance unbiased in a suitable class of more general regression models.  

Since I'm thanked in the acknowledgments, I thought it might be prudent to make explicit a few reservations I have about Bruce's version of the GMT.  Here then is my unexpurgated original comment on an earlier draft of the paper.

Bruce,

I hope that you won’t mind an unsolicited comment on your recent Gauss-Markov paper.  I was wandering around somewhat aimlessly yesterday looking for recent work on model averaging for a refereeing task, and it attracted my attention.  (Spoiler alert:  I’ve always hated the GM Thm since it seemed to restrict attention to such a small class of estimators that its optimality claim was nearly vacuous.)

There is of course the (Rao?) result that ols is MVUE in the Gaussian linear model, but you want to say something much stronger, that it is MVUE in a much bigger class of models, but then the qualifiers become critical.  I think that I understand where you are coming from, but I wonder whether you might be misleading the youth of econometrics by the way that you develop the argument.  Your “for all F in calF_2”  is quite strong.  Of course median regression can be much more efficient than mean regression in iid error linear models and both are unbiased when the errors are symmetric.  When errors are iid and not symmetric then median regression is biased, but only the intercept is biased, the slope parameters are still potentially much more efficient than the mean regression estimates would be. Here, I don’t mean to suggest that there is anything special about median regression — a plethora of other estimators would serve as well.  There is merit, I concede, in the idea that “if you want to estimate a mean you should use the sample mean, etc” — I’ve heard this from Lars several times, but on the other side of the argument there is the infamous Bahadur-Savage result that the mean is never identified, in the sense that slight perturbations of the tails of the population distribution can make it bounce around arbitrarily. Of course, this depends upon what “slight” might mean.  Your “for all F…” condition and unbiasedness for any linear contrast gets us back very close to requiring linearity, it seems.

The paper is fine, I just think that it might need a surgeon general’s warning of some sort.

Best

Roger

PS.  The photo, taken recently at the Museum of the History of Paris (Carnavalet), depicts a metal, Napoleanic era, hat of the type that Edgeworth presumably had in mind.


PPS (added March 12 2022)  There are now two papers circulating one by Steve Portnoy and the other by Benedikt Potscher showing that the unbiasedness condition of Hansen admits only linear estimators, so my "very close" in the note above could be strengthened a bit.  It also occurred to me after writing the original message that the 1757 proposal of Boscovich defines an estimator that can have superior (to OLS) asymptotic MSE performance and is asymptotically unbiased for iid error linear models.  Details are given here.

Friday, November 5, 2021

UNIX is 50!

 



I don't think that there was anything remotely as influential in my research  experience as the existence of UNIX.  When I arrived at Bell Labs in 1976 UNIX was still in its infancy but already there were rumblings of a new statistics  language called S that would revolutionize my world.  In my last few years at Murray Hill my office was across the hall from Rick Becker's, so I was able learn S from one of its original authors.  When I returned to UIUC in 1983 it was a struggle to maintain my access to S and UNIX.  I recall the director of campus computing services telling me at that time that "UNIX wasn't appropriate  for educational institutions because it was too flexible."  Eventually accounts on various Vaxen were created and life went on with a commercial manifestation of S called Splus.  In 1989 on a yearlong  sabbatical adventure I was even able to maintain my dependence on UNIX with a dubious version by SCO  on a Zenith portable.  Sometime in 1999 I made the transition to R, and have never looked back.  Without UNIX all of this would have been almost unthinkable.