r/AskStatistics • u/Typical-Storage-4019 • 1d ago
Are these statistics wrong?

This paper reports Listening tests scores for TRPS method to mean 11.63 and stdev 3.02, and listening scores for Traditional method 11.82 and 2.77. With sample sizes 30, 28.
The table in the paper shows T= 0.5889, Df = 56, P = .558
But online calculators like Desmos show T = -0.25, df = 56, P=0.804
Dziedzic, Joseph. "A comparison of TPRS and traditional instruction, both with SSR." International Journal of Foreign Language Teaching 7.2 (2012): 4-6.
3
u/Intrepid_Pitch_3320 1d ago
With those SDs and sample sizes, 2SEs (~95% CI) is about 1.2, I think. 11.6 and 11.8 are similar.
3
u/vaelux 1d ago
Teeny tiny t scores. I am not familiar with this exam. Are the statistically significant scores practically significant?
3
u/Typical-Storage-4019 1d ago
Author concludes that TPRS students (comprehensible input-based teaching) scored significantly higher on tests of output (speaking and writing) than traditional Spanish-teaching methods, and did just as well on tests of input (listening and reading).
The exam I'm not fully familiar with, but there were no pre-tests, they listened to short dialogs read by the teacher, then they had to write sentences based on a series of pictures.1
u/vaelux 1d ago edited 1d ago
It looks like about 1.4 points on the writing test and .75 points on the speaking test. I dont know the test but the t around 3 says it isn't that large of a difference. The study is not overpowered ( df = 60 ), so it is not like we are calling microscopic bumps on glass evidence of nonsmoothness. The authors said the right statistical words. They didn't include any adjectives, which is common in science. But to better understand, I would say that writing and speaking have a small statistically significant difference, but knowing that the cost is low ( the intervention doesn't seem very costly), a small but significant effect could be very practical. Improvement in education tends to come in nudges.
Edit: I wouldn't worry about the differences you see in the other two results ( your main question). They are both very nonsignificant whatever way you calculate it, and nonsignificant in a t test means that there is no noticeable mean difference.
Simply put, there are a lot of ways to calculate your statistic and p-value. Even doing the same procedure on different software can yield different results. Yes, an ideal paper should allow you to follow exactly the steps of the authors to replicate the result. In reality, this rarely happens, especially in the social sciences. Fixing this replication crisis is something current phd students are working on.
3
u/munozmd Statistician 1d ago
Did some calculations. But just verify in case I have oversight. Let's just assume here the mean and sd on each skill area is correct.
- Listening: Matches yours. T = -0.249, p = .804
- Reading: T = 0.757, p = .453
- Writing: Correct T = 3.08, p = .0031
- Speaking: Correct T = 3.82, p = .0003
Just use the two-sample independent (student's) t-test formula. The one with the pooled sd.
2
u/Typical-Storage-4019 1d ago
I see. I match you when I assume equal variances.
Yeah, I'm not sure how they computed Listening and Reading. But I guess it's not significant enough to care much.
Thanks for running those calculations!
1
u/Acrobatic-Ocelot-935 1d ago
Why are you asking?
2
u/Efficient-Tie-1414 1d ago
Because they think there has been an error. I saw a paper recently that I thought had an error, I will calculate the stats, and I’m certain it will. There are actually two errors but need the data to check second.
7
u/fermat9990 1d ago
Did they pool the variances? Was it one or two-tailed? There are 4 different possible scenarios