← Back to Manuscript Reviews

Revisiting representativeness classic paradigms: Replication and extensions of nine experiments in Kahneman and Tversky (1972)

Manuscript Reviews

Recommended citation (APA 7)

Röseler, L. (2023, March 28). Review: Revisiting representativeness classic paradigms: Replication and extensions of nine experiments in Kahneman and Tversky (1972) [Peer review]. Open Review Tracker. https://lroesele.zivgitlabpages.uni-muenster.de/reviews/manuscript-reviews/qjep-representativeness.html (accessed September 16, 2026).

Manuscript title: Revisiting representativeness classic paradigms: Replication and extensions of nine experiments in Kahneman and Tversky (1972)
Invitation to Review date: 07.11.2023
Review submission date: 16.11.2023
Review type: original submission / revision / other

I have reviewed the manuscript “Revisiting representativeness classic paradigms: Replication and extensions of nine experiments in Kahneman and Tversky (1972)”. The authors provide a very detailed and transparent report of their replication study, which they extended in numerous valuable ways (e.g., by computing the reliability of people’s reliance on the heuristic). With respect to the conceptualization and execution of a replication study, I believe that this type of replication study is state-of-the art replication research. I apologize for my forwardness but I hope that the Quarterly Journal of Experimental Psychology takes this as a reason to actively encourage the submission of replication studies, data transparency, study preregistration etc. (currently, the TOP factor is far below what I would consider field standards; https://topfactor.org/journals/quarterly-journal-of-experimental-psychology).

I strongly recommend this report to be published. However, I also recommend that the authors revise the manuscript with regards to clarity and several small but sometimes confusing typos. I have all remarks listed below. Despite there being many, I think that they can all be easily resolved. I hope that my comments can help the authors improve the manuscript.

Finally, please note my disclosures at the bottom. I apologize for an inconsistent use of page numbers. There seems to be something wrong with the created pdf proof (see specific remarks). Please contact me if any of my comments and suggestions are unclear.

General Issues

  1. I think one red thread in my comments below is the order of presentation. Several questions arose during my read through that were answered later. I am not sure how to resolve this but I think slightly increasing redundancy by mentioning them (e.g., small telescopes approach) multiple times should be fine.
  2. I found some of the effects a bit difficult to understand. I think that it is not entirely your responsibility to explain everything as this is primarily a replication but I think that elaborating, or at least providing references to exhaustive explanations on the heuristics and why certain answers are wrong could help readers (it would have helped me at least).

Remarks in no specific order

  1. Keywords: There should be a semi-colon in the keywords after “probability” but it appears there is a comma in superscript
  2. P. 5, l. 32: replace “) (“ with “;”
  3. P. 5, l. 44 “extroverted“: I think this should say extraverted
  4. P. 6, l. 5-8: There is a page number missing for the direct quote (“representativeness is a relation…”)
  5. P. 7, l. 35: Just FYI: I think that this is an excellent idea to extend the original study. Reliability of JDM tasks is a big problem. There is research on internal consistency of anchoring by Teovanović (although this is strictly speaking not anchoring but advice taking), Schultze, and by myself.
    1. Advice Taking
      1. Teovanović, P. (2019). Individual differences in anchoring effect: Evidence for the role of insufficient adjustment. Europe's Journal of Psychology, 15(1), 8.
      2. Schultze, T., Gerlach, T. M., & Rittich, J. C. (2018). Some people heed advice less than others: Agency (but not communion) predicts advice taking. Journal of Behavioral Decision Making, 31(3), 430-445.
    2. Anchoring
      1. Weber, L., & Röseler, L. (2022). Testing the Reliability of Anchoring Susceptibility Scores.
      2. Röseler, L., Weber, L., Helgerth, K., Stich, E., Günther, M., Wagner, F. S., & Schütz, A. (2022). Measurements of Susceptibility to Anchoring are Unreliable: Meta-Analytic Evidence From More Than 50,000 Anchored Estimates.
      3. Röseler, L., Schütz, A., Dolling, I. K., Friedinger, K., Hösch, Y., Hügel, J. C., ... & Röseler, J. J. (2020). The Stepwise Anchoring Paradigm: Measuring Reliable Components of Anchoring and Adjustment as the Next Step in Moderator Research.
    3. Internal consistency of JDM tasks is generally a big question. Nobody repots reliability but lots of people “find” personality moderators.
      1. Parsons, S., Kruijt, A. W., & Fox, E. (2019). Psychological science needs a standard practice of reporting the reliability of cognitive-behavioral measurements. Advances in Methods and Practices in Psychological Science, 2(4), 378-395.
  6. P. 9, l. 25 “receive five out nine”:
    1. I think it should say five out *of* nine. Also, I do not quite understand the procedure. Wouldn’t 668*5=371… be higher than 334? What about drop-out and exclusion criteria? Have you planned to apply these to the ~371 people so that at least 334 will remain for the main analyses?
    2. Why did you not assign all participants to all nine problems? Wouldn’t this have saved time (less people in total would have had to fill out the consent form, demographic questions, and the decision styles scale). [You explain that on pdf page 30/66 “It would be cognitively demanding”, maybe mention that a second time here.]
  7. P. 9 Power analysis: I suggest that you include more detail in this section instead of referring to the supplementary materials. If you decide against this decision, could you please be more specific and provide a link, page number, or something that tells the reader where to look? Ideally, there should also be a direct link to a reproducible power analysis here. [This one: https://osf.io/ge9n4]
  8. P. 9, l. 42: “participants who correctly guessed the hypothesis of this study in the funneling section”: Have you specified this anywhere? Were there keywords that had to have been mentioned? In my experience, this is no black-and-white criterion but should be detailed beforehand and coded by multiple people afterwards.
  9. P. 9 bottom “In the analysis with exclusions, all effects have confidence intervals that included the SESOI, whereas in the analysis with exclusions, all but two effects had…”: One of them has to say “without exclusions”.
  10. P. 10 top “using the small telescopes approach”: You did not mention this before (or only in the supplements which I am yet to find). I recommend again that you explain the power analysis in more detail in the main manuscript and quote the small telescopes approach paper for readers unfamiliar with it.
  11. P. 10 top “We initially planned to focus our reporting on the full sample, yet decided to include the exclusions reporting in the main, and move the full sample reporting to the supplementary, as we felt that the exclusion criteria likely reflects more accurate results given the cognitively demanding nature of the experiment.“: I found this sentence a bit confusing. Does initially mean in the preregistration here? Why did you change your mind? Criteria reflect (without s). I think you mean that people should be concentrated during the experiment for it to work, which is why you report results for the sample with exclusion criteria applied?
  12. P. 10, l. 15: With N = 623, can you please report power for the smallest non-null effect? I am asking for this as there is a chance of power deflation with multiple (possibly independent) tests. So 8 effects with 95% power each will – given that they are not internally consistent – lead to only .95^8=.66..% power.
  13. P. 11, Table 1:
    1. Can you provide intervals for the sample size range for the different problems? 1500 in total does not seem very informative to me.
    2. The same applies to the replication sample size. Wouldn’t the range of respondents for each of the scenarios be more informative? You also do not include information about the “5 out of 9 scenarios” here which makes it appear that there were 623 people for each scenario, but this is not the case, is it?
    3. Geographic origin: Is there a difference between US and US American? If so, could you explain it, maybe in the table’s note?
  14. Table 2:
    1. It says “Samp-population similarity”. Do you mean “sample-population”?
    2. Did I understand correctly that you computed Cohen’s h from the reported statistics? If so, can you please explain somewhere how you computed the effect size (e.g., R-code, citation of package and its version, page numbers for the original statistics). For some effect sizes, there are different ways to compute them and they do not always converge and this would facilitate reproducibility.
  15. Table 3:
    1. I find the predictions and problem wording for #1 difficult to understand, could you please clarify: All families were surveyed and result of 72 families are reported. How many families are there? If 72 were already reported and the city is small, all other orders might be less likely. I think the description of the problem leaves a few things unsaid, as is the case in many of these tasks (e.g., due to this Wason Selection Tasks lead to non-interpretable data in almost all cases, in my opinion).
    2. You have three lines vs. one example (gbgbbg) in the problem column but only three different examples in the predictions. From the problem column I would expect there to be 3 items (gbgbbg vs bgbbbb, gbgbbg vs bbbggg, and gbgbbg vs. gbbgbg. But in the predictions column you only have two predictions (1a and 1b): gbgbbg vs bgbbbb and gbbgbg vs. bbbggg. I know that this would clarify upon looking at the questionnaire and the text below but I found it confusing in the table.
    3. #2: Thank you for pointing this example out. I had initial difficulties and simulated what I understood the issue to be. According to my calculations, this is a matter of 1.0% vs. 0.9% likelihood. Did I get something wrong here? If not, I would dismiss this problem as practically irrelevant and beyond the scope of what should be considered a heuristic.
      R-Code:
      set.seed(1)
      a <- replicate(1000000, rbinom(n = 1, size = 100, prob = .45))
      b <- replicate(1000000, rbinom(n = 1, size = 100, prob = .65))
      mean(a == 55)
      mean(b == 55)
      var(a)
      var(b)
    4. #4: It says in “predictions” “Type II distribution is perceived as more probable than Type II distribution.”: I believe that some of these problems are very difficult to understand and I think that you agree as you wrote about their cognitive nature. These types of errors make the manuscript unnecessarily complicated and very frustrating to read. I strongly recommend each of you to go through the manuscript looking for these kinds of errors.
  16. P. 17 overlapping categories: Why did you not include an extension where categories do not overlap?
  17. P. 17 / 26 [referring to pdf page numbers from here on] l. 20-25: Alphas and betas are not properly formatted in my pdf-viewer. There are only squares.
  18. There is something wrong with the page numbers. After page 21 of 65 // 21 // 22/66 it changes to page 22 of 65 // 14 // 23/66. I hope that you will still find which parts I have been referring to.
  19. P. 28 l. 33: Proportions are formatted inconsistently (1/6 and ⅙).
  20. P. 29 “the problem read as follows”: An “s” is missing here.
  21. P. 29 bottom: I think there are papers that discuss the language issues in all of these tasks. After all, most of these phenomena raise from two people having different concepts of a words (e.g., “random”) and one explaining to the other that their concept is wrong. I searched for it but I could not find it, sorry!
  22. P. 30: “referred to in the supplementary file”: Again, and also for all further cases, please indicate which file, which page, etc.
  23. P. 42 l. 15: Please also indicate what other software and versions you used (e.g., esc and powerAnalysis packages and R) and cite it.
  24. P. 47 l. 31:
    1. I recommend that you first report whether the items are intercorrelated (ideally some reliability analysis would be informative, but I am not sure whether your incomplete design allows for that; section “Associations and Comparisons Between Problems). Because if the problems are not correlated, they cannot be expected to be correlated with anything else (reliability is necessary for validity).
    2. What are the results if you stick to the preregistration? Have you calculated them and can refer to them somewhere (supplements, footnote)?
  25. Can you please list at some point in the manuscript all deviations from the preregistration and discuss why you deviated and how they affected the results?
  26. I think that you could cluster the problems using actual clustering analysis instead of visual inspection of pairwise comparisons. I also think that this might be beyond the scope of this manuscript. Nevertheless, I like your suggestion to cluster these problems very much.
  27. I am unsure about this but the title looks a bit odd to me (English is not my mother tongue). My unbinding suggestion is: “Revisiting representativeness heuristics: Replication and extensions of nine classic experiments…”.
  28. Data and code:
    1. You saved the raw data in .sav and processed data in .csv format. Is there a reason why you went with the SPSS-format in the former case? If not, can you please also provide raw data in a non-proprietary format (e.g., also .csv)?
    2. I could not knit the .Rmd file (KT1972replicatoin_v7-L.Rmd). I downloaded the file and put it in the same directory as the raw dataset (“Kahneman & Tversky-1972…sav”). Here is the render output. Can you please fix this or explain in the README file how to avoid this error? Due to this error, I did not check any of the results and calculations.

      processing file: KT1972replication_v6-L.Rmd

|.. | 5% [packages]

Quitting from lines 32-50 [packages] (KT1972replication_v6-L.Rmd)

Error in `contrib.url()`:

! versuche CRAN ohne einen Spiegelserver zu nutzen [engl: trying to reach CRAN without a mirror server]

Backtrace:

1. utils::install.packages(new.pkgs, dependencies = TRUE)

2. utils::contrib.url(repos, "source")

Ausführung angehalten [engl: execution stopped]

  1. I saw that you have published the pre-print already. I suggest that you also publish the accepted version of the manuscript as a pre-print as the journal is not a diamond open access journal (see https://v2.sherpa.ac.uk/romeo/search.html).

General recommendations for replication studies

  1. In the case of discrepancies, I recommend that the authors of the replication study contact the authors of the original study. Due to the replication success I do not think that any insight will be gained from this procedure here.
  2. I recommend that the authors make a comment on the original study via pubpeer.com or alternative systems and include information about the replication study, materials, and outcomes. Via the FORRT Replications and Reversals and the Replication Database (which will include this study eventually), we plan to create pubpeer comments with an automated bot (just like statcheck.io did).
  3. Ideally, journals that published original studies should also publish corresponding replications (pottery barn rule, https://thehardestscience.com/2012/09/27/a-pottery-barn-rule-for-scientific-journals/). If the authors have already submitted the manuscript to the original journal and it was rejected, I (as a reader) would like to know about the reasons. If they did not submit the study to the original study’s journal, I would like to know why they decided against this.

Disclosures

  1. I have been working together with Gilad Feldman on a meta-analytical project about replication studies. Personally, I am a big fan of his detailed and numerous replication efforts. I tried not to let this bias my judgment.
  2. I did not run a reproducibility check due to the close review deadline. I strongly recommend the journal to do a reproducibility check on all of their published manuscripts. I may be able to do one provided that I am granted more time.
  3. I ran the manuscript through statcheck.io but no statistics could be recognized. I believe that this is due to the journal formatting and I encourage you to run the manuscript through statcheck themselves if they have not done so yet.
  4. I did not compare the preregistration with the actual methods and analyses. I recommend the authors to systematically document all deviations.
Download originalWord document