Ask HN: Would you read a statistics textbook?
Posted by usernametaken29 2 days ago
I had a classic Bayesian statistical education throughout my undergraduate and grad years and I’ve come to conclude that the plethora of books offered are pretty bad.
For me personally statistics is intuitive if illustrated properly. A great example is https://seeing-theory.brown.edu/
I was wondering whether it would make sense to turn the intuitions behind statistics into a book. Would anyone read it? Do people still read stats books or would it be mostly a waste of effort? If you wanted to teach/help people understand statistics, what resources would you consider?
Comments
Comment by gus_massa 1 day ago
1) The main book, that has a complete explanation and is well ordered. It's for learning.
2) Tha Landau book, that is super short and hard. It's only to check you didn't miss any important formula or topic.
3) There Feynman book, that is anassorted colection of fairytales for physicist. It's a pleasure to read it but you must already read 1 to understand it.
4) The Shaum book, that is almost a colection of exercices. Some people hate it. Some people love it. I like it as a companion to theother books.
I guess you are complaining that 1 is boring and want to write 3. It's a good idea, but it's harder than expected.
Comment by msla 1 day ago
1. Basic introduction.
2. Reference tome, which has absolutely everything.
3. Cookbook with style advice for the more advanced student, which assumes you've read 1 and can look up various details in 2.
These days, 2 would be a wiki and 1 would likely be a bunch of pages on that wiki, but it's still good if you have someone sit down and write 3.
Comment by WCSTombs 1 day ago
[1]: https://diataxis.fr/
Comment by throwaway81523 1 day ago
Comment by lukasbm 1 day ago
Comment by ssivark 1 day ago
Comment by eru 1 day ago
Modelling distributions explicitly sounds nice, yes.
Comment by ssivark 1 day ago
If you are not doing something crazy, most reasonable people would agree with your judgement. Conversely, if you are making non-obvious inferences where reasonable people disagree, you are in murky water and no sophisticated statistical method will save you. Math is not magic; theorems merely recycle (launder) modeling assumptions into results.
Comment by jbs789 23 hours ago
Comment by stackghost 23 hours ago
Only people with prior education/training in statistics are capable of doing this. The people who don't need a textbook.
Something like 60% of US adults read at or below the 6th grade level, and 25% of US adults struggle to comprehend graphs or charts entirely. Someone who has no idea what a standard deviation is can't intuit about distributions. I think you're dramatically overestimating the average person.
Comment by ssivark 23 hours ago
I disagree vehemently with this claim. I could cite my experience in teaching this topic to liberal arts / humanities college students in the US, but it is really more obvious than that. Anyone can understand a histogram easily and far more intuitively than they can understand the formula for a standard deviation and whether it must divide by N or N-1. Statistics courses and textbooks get stuck on that kind of pedantry, and most students end up missing the forest for the trees.
Comment by discardable_dan 23 hours ago
Comment by actualeff0rt 1 day ago
Maybe probability and statistics are a skill issue on my behalf, but what I absolutely loathe is the absolute lack of standardisation when it comes to notation in measure theory. Every textbook does it differently. All of them assume that their notation is the one everyone uses. Nobody bothers to explain _what_ the notation means. If you ask me, every bit of new notation should be introduced with a sentence or two on "how to read this symbol in your head" - especially when there are indices, subscripts and superscripts involved. It's especially terrible for measure theory because there's so much "implicit" information you're supposed to gather from the context - but in a way I understand it, because if every bit of notation of absolute and complete, I imagine it would be quite hard to type up.
Anyways, my rant on measure theory notation aside - I would absolutely read yet another Prob/Stats textbook. But unfortunately I will also drop it really quickly if the author doesn't show me any "notation-sympathy" :)
Comment by montalbano 1 day ago
This is widely regarded as the most accessible intro textbook to Bayesian statistics.
Comment by stdbrouw 1 day ago
Comment by geokon 1 day ago
- The language is often very vague and imprecise, so it's more difficult than it feels it should be
- Concepts are sometimes introduced "at random" in ways that only really make sense in hindsight. So as you're reading you're left scratching your head as to why something was brought up.
- There are constant philosophical and historical digressions that seem to hold deeper meaning, but maybe once you know the topic already.
- Similarly, constantly talking about people that take issue with the method. People not liking the method is a constant theme (they seem really butthurt about this?). But the craziest part is this all done before you even really understand what the method is!!
- The editor must have placed some strict requirement of saying "Bayesian" at least five times per page.
- No index. Useless table of context. But lots of end-notes you feel compelled to flip to constantly
Overall it feels like a textbook written to impress other statistics professors - and as an outlet for the author to air some frustrations with how people do statistics (which may be completely valid!)
The overall structure and objectives seem solid for the most part. It's just a lot of the details aren't great. The problems (so far) have been good. The examples in the text are fun and compelling, but you have to do your own legwork to actually pick through all the prose and tie the pieces together - to figure how it fits together mathematically. Fortunately AI helps as a tutor
Comment by stdbrouw 13 hours ago
Comment by geokon 5 hours ago
Comment by howunfortunate 1 day ago
I read a lot of informational things, but math / stats / software has always felt like an area where a book is just the wrong format.
If I were you I'd make an interactive website like SQLZoo or a video series like StatQuest.
Those are educational formats that really clicked with me for whatever reason.
Comment by jan_m_savage 17 hours ago
- put a lot of 'teaching' into the book. Don't write so the learner memorizes; instead, motivate each discussion so that the learner understands, and gains what Wirth calls _coverage_.
- ensure there are no errors; this will turn off learners.
- keep a light tone, be yourself not a stuffed shirt. A good model for such writing is Nassim Taleb's non-technical books.
- never seem to be in a hurry. Many good books were ruined because midway the author lost interest and just hurried up.
- don't be too verbose, and neither too concise. Every sentence should do required work; learn how to write (if needed) and practice daily.
- ensure the book uses good, comfortable fonts, and good typography. Must learn this. Many good books ruined due to lack of fluency effect.
If you conscientiously do these things, there will always be readers for your book.
Good luck!
Comment by mdspan 2 days ago
I think more resources like seeing-theory would be great since stats books are almost universally dry (Blitzstein being a notable exception), but I'm not sure how easily more advanced concepts lend themselves to visual explanation in a way that's digestible for a non-stats person.
Comment by ivansavz 21 hours ago
It's still complicated but the visualizations help a lot!
Comment by usernametaken29 1 day ago
Comment by diogenes_atx 15 hours ago
Ian Ayres (2007) Super Crunchers: Why Thinking by Numbers is the New Way to Be Smart.
Joseph Healey (2020) Statistics: A Tool for Social Research [This is the textbook used in the class, which I highly recommend].
John Allen Paulos (1988) Innumeracy: Mathematical Illiteracy and Its Consequences.
Nate Silver (2012) The Signal and the Noise: Why So Many Predictions Fail, but Some Don't.
Nassim Nicholas Taleb (2005) Fooled by Randomness, second edition.
Nassim Nicholas Taleb (2010) The Black Swan: The Impact of the Highly Improbable, second edition.
Comment by jldugger 1 day ago
At some point in my professional career I started reading non-fiction books and even bought a used statistics textbook for 10 bucks on abebooks. I didn't end up actually reading it until 12 years later during the COVID lockdown. I ended up shooting for 10 pages a day, 7 days a week. If those 10 pages included review exercises, it would be a long night.
Could just be the right book at the right time, but this one really helped me understand stuff beyond the normal HS math stuff, like RMS-error, calculating correlation, the difference between standard error and standard deviation, the relationship between sample size and standard error, t-tests, and chi-squared. Working as an SRE/release engineer, this stuff really helped me overcome a lot of _bad_ canary data analysis my predecessors had constructed.
That book was the 3rd edition of Statistics by Freedman et al.[1] One thing I want to complement was getting the pedagogy right. Most chapters have strong narrative hooks, several "check your knowledge" problems, review exercises, and post chapter bullet points to assist with spaced repetition. There's even a series of "special" review exercises covering entire sections of the book, ie exams.
For the HN crowd I should also probably note that the book is almost entirely non-bayesian and not intended to prepare readers for further coursework. You will not learn normal phraseology like "IID," "random variable" or "kernel".
[1]: https://www.amazon.com/dp/B00SLB5Q72?lv=shuf&channelId=520&p...
Comment by a_bonobo 23 hours ago
Where would your book fit into this?
Comment by NishanStepak 12 hours ago
Comment by dmwood 13 hours ago
Comment by thastings 1 day ago
Comment by RobGR 1 day ago
Comment by pks016 14 hours ago
But, It would be helpful for beginners. I work with students and I know a lot (like a lot) of students (not math or stats major, other disciplines) are scared of statistics. They are always looking for resources to learn.
Comment by Eridanus2 2 days ago
Comment by jadermcs 1 day ago
Another exemple of a successful visual pedagogical content is: https://www.byhand.ai/
Comment by ludicrousdispla 1 day ago
If you could do something similar for bayesian statistics I think that would be useful, but not necessarily popular.
Comment by philosopherNoob 1 day ago
Comment by sunir 16 hours ago
Comment by junon 1 day ago
However going the 'visual' route might be enough for me to pick it up.
Comment by throwaway81523 1 day ago
Comment by jsw97 23 hours ago
Comment by 1970-01-01 19 hours ago
Comment by clutter55561 1 day ago
But beware of opinions.
Don’t let people put you down, especially here in HN, where people are perceived to smart. Smart doesn’t equal sensible or unbiased.
Many books are written to scratch the itch of the author. Just like an open source project. It is a work of love.
Comment by lormayna 23 hours ago
Comment by aghuang 2 days ago
More generally, I would buy a statistics if it is linked to today's interesting technological breakthroughs and also if it comes as a distilled version for beginners.
Comment by pessimizer 1 day ago
Can you write one that's more worth reading than the standard ones? Don't answer that question, just prove it.
Comment by stared 1 day ago
You will also see hos long it takes - and what is thd difference between an idea and making it real.
Comment by itake 1 day ago
Comment by xtiansimon 20 hours ago
Comment by bigdict 1 day ago
Comment by gignico 1 day ago
Comment by BoredomIsFun 1 day ago
Comment by rootsudo 1 day ago
Comment by gdulli 1 day ago
Comment by hollowturtle 1 day ago
Comment by whattheheckheck 1 day ago
Comment by rramadass 1 day ago
However the link you have provided is not the way; it is low on content and high on pretty distractions. Use all sorts of diagrams and graphs primarily, with animations only where required. The key is to always relate to something in the real world so one can see its actual relevance. Also tie it back to other fields of mathematics so one can see how they all come together.
A good example to study is How to Measure Anything: Finding the Value of Intangibles in Business by Douglas Hubbard. Detailed review at - https://www.lesswrong.com/posts/ybYBCK9D7MZCcdArB/how-to-mea...
And of course Nassim Taleb's works are a good source of inspiration. Here is a great video summarizing Taleb's ideas nicely Pareto, Power Laws, and Fat Tails - https://www.youtube.com/watch?v=Wcqt49dXtm8
Comment by sdcfgy 1 day ago
Crap jokes aside, I mostly use Statistics In A Nutshell. It’s pretty ok. I say that as a member of the RSS for 30 odd years.
Comment by Saline9515 1 day ago
Ready the theory is nice but you only learn by solving problems that uses it.
Comment by exe34 1 day ago
I like Statistical Rethinking, but unfortunately the solver was very slow and I didn't end up using it much - not the author's fault, it's probably my computer that's way too old.
Comment by aspectmin 1 day ago
Good luck if you do this.
Comment by tryauuum 15 hours ago
Comment by eimrine 1 day ago
Comment by chermi 13 hours ago
Statistics is big. It's hard to answer without knowing what you plan on covering. I agree that most books (I've seen) suck. Wasserman all of statistics is the best I know, for my purposes. The problem I've found with most statistics books is trying to be too cute and clever by catering to a certain crowd thereby justifying holes that make it harder to truly understand and making the subject seem an incoherent patchwork.
The other problem that's even harder to solve is that the majority of people picking up a statistics book don't really think they need to learn statistics, just certain pieces. Which makes statistics seem less coherent and thus furthering the perception that statistics is in fact incoherent, leading to special background books that say "here's all you really need to know about statistics."
The final problem is statistics IS kind of incoherent as most often presented, especially when it tries to be what I'd roughly call "backward compatible". Why is so much time spent on p-values, for example? Is that really what a consistent modern perspective of statistics entails[1]? But backward compatibility forces it's inclusion, because that's what people still use and have used because they never got taught of "philosophy" of statistics, but rather rules. If you don't teach those rules everyone else uses, you're making them spend even more time on statistics than the small amount they're already unhappy spending. So the lowest common denominator is taught, which is incoherent. It's a vicious cycle that is not the fault of statisticians.
To an outsider there's not unifying dominant through lines because the field is justified largely through its application. There's no obvious "philosophy" or vibe of how to think about statistics. What I mean is that one approaching the field doesn't really grow more comfortable with it and doesn't really feel like they're advancing to a more coherent view, so they're less motivated to try to grok it, because it simply doesn't look like there is something to grok. With all of previous factors contributing to this.
So I guess the point of my ramble is that I think what's missing the promise of a reward for learning statistics "properly", really understanding it. In my biased view this is largely because it's not often presented to the non-professional statistics student as if there is a clear thing to understand, more just a collection of tools. And I think this is partly because the author already knows the reader doesn't want to spend time on it. So my advice would be, come up with a clear story of what statistics is. Promise and fulfill the promise that it is worth spending the time in a more abstract world for a little bit because the end is worth it. That it will save time and frustration in the long run because it won't just be a collection of basically faith-based tools with mystical rituals they will feel uncomfortable with for the rest of their careers. Tell that coherent story and then you can demand more commitment from your audience[2]. Maybe that cuts down your audience size but I think the net effect will be more people understanding statistics and what it actually is.
What would be the main threads and themes in your book? How would you justify spending time on it? What would require as necessary background? I don't understand the subject enough to offer any opinion. The most coherent things to me are convergence types and bounds. And then do you include computational approaches? For example in practice I'd say 90%+ of people would be better off using bootstrapping analysis for errors in most real-world cases compared to the standard "approved" approaches, but that's not very satisfying and the theory of it is certainly hard to integrate cleanly. TL;DR I don't envy anyone writing a statistics textbook.
[1] This is just a single aspect, but it gets to a broader problem. This is a really hard problem to overcome. So much of academia is taught p-values and to so much of academia that is basically the hardest math they know, and they work very hard to follow the rules they were taught, which to them is statistics. I don't know how you tell them to re-learn something especially when the answer is more math. That makes them even more hesitant and less likely to fully embrace a deeper understanding of the field.
[2] maybe that's the larger point. The analogy is maybe teaching calc vs. algebra based physics. Algebra based physics is harder to teach and learn, and it is less coherent and complete. There is not much meat in it, mostly rote tools. IMO conservation laws like energy and momentum and solving shit from there would leave a better taste for what physics is, but instead they are often used in one lecture to derive kinematic equations which are then elevated because students can memorize them. I guess it's better than nothing but it leaves people with the wrong impression of what physics is, just like I think most statistics leave a bad impression of what statistics is. Make the statistical equivalents of something like conservation laws the central objects. Something that makes it a coherent story that builds and rewards.
Comment by PolGraciaSerrah 1 day ago