The Effect of Third Party Implementations on Reproducibility
Abstract
Domain fit: AI-adjacent · Paper appears method- or tooling-adjacent to AI workflows with partial ecosystem coverage.
Reproducibility of recommender systems research has come under scrutiny during recent years. Along with works focusing on repeating experiments with certain algorithms, the research community has also started discussing various aspects of evaluation and how these affect reproducibility. We add a novel angle to this discussion by examining how unofficial third-party implementations could benefit or hinder reproducibility. Besides giving a general overview, we thoroughly examine six third-party implementations of a popular recommender algorithm and compare them to the official version on five public datasets. In the light of our alarming findings we aim to draw the attention of the research community to this neglected aspect of reproducibility.
Results and benchmarks
Reproducibility of recommender systems research has come under scrutiny during recent years.
Benchmark evidence is limited
Evidence graph: 3 refs, 3 links.
Utility signals: depth 70/100, grounding 75/100, status medium.
Implementation
Historical official implementation (not recommended for new builds)
Only a historical official implementation is available
Use with caution for new projects; verify against current tooling and maintained community alternatives.
hidasib/GRU4Rec · 810 stars · Last push Aug 24, 2023
pcerdam/KerasGRU4Rec is the closest maintained adjacent implementation (Community adoption signal (106 stars)). It is not paper-verified; validate algorithm and evaluation setup against the paper before trusting reported metrics. Community adoption signal: 106 GitHub stars.
Open hidasib/GRU4Rec- Adjacent implementations are not paper-verified
- Recommended repository is adjacent and not paper-verified.
- Adjacent implementation match confidence is low.
- No direct maintained implementation is currently verified.
- Only historical official repository was found: hidasib/GRU4Rec.
- No maintained paper-verified implementation met reliability thresholds.
Compare implementation paths
Compare maintenance quality, reproducibility coverage, and evidence confidence before choosing a reproduction baseline.
- Maintenance
- Stale
- Confidence
- High
- Reproducibility
- Limited
- Stars
- 810
- Last push
- Aug 24, 2023 (1098d)
Official implementation from Papers with Code · Repository link is mentioned in the paper metadata
- No push in 12+ months
- No CI pipeline detected
- No tagged releases
- Maintenance
- Stale
- Confidence
- High
- Reproducibility
- Limited
- Stars
- 106
- Last push
- Jul 30, 2024 (757d)
Official implementation from Papers with Code · Repository link is mentioned in the paper metadata
- No push in 12+ months
- No CI pipeline detected
- No tagged releases
- Maintenance
- Stale
- Confidence
- High
- Reproducibility
- Limited
- Stars
- 83
- Last push
- Aug 24, 2023 (1098d)
Official implementation from Papers with Code · Repository link is mentioned in the paper metadata
- No push in 12+ months
- No CI pipeline detected
- No tagged releases
Reproduction readiness
Major work
No dependency manifest, manual reconstruction required
- hidasib/GRU4Rec has no requirements.txt, environment.yml, pyproject.toml, or Dockerfile.
- You will need to reverse-engineer dependencies from import statements in the source code.
- Last push was 1098 days ago.
Hardware requirements
- Expect multi-day setup/compute for meaningful reproduction based on current guidance.
Validation caveat
Framework baselines
- Hugging Face Transformers training guide
Modern transformer training baseline.
- PyTorch nn.Transformer docs
Reference transformer building block implementation.
Repositories and ecosystem
Closest related implementations
These are not paper-verified. Use them as reference points when no direct implementation is available.
- pcerdam/KerasGRU4Rec Adjacent · Confidence: Low · 106 stars
Community adoption signal (106 stars)
Official
- paxcema/KerasGRU4RecConfidence: High
Keras implementation of GRU4Rec session-based recommender system
106 stars · 40 forks · Last push Jul 30, 2024
- hidasib/gru4rec_pytorch_officialConfidence: High
Official implementation of the GRU4Rec algorithm in PyTorch
83 stars · 12 forks · Last push Aug 24, 2023
- hidasib/gru4rec_tensorflow_officialConfidence: High
Official implementation of the GRU4Rec algorithm in Tensorflow
7 stars · 1 forks · Last push Nov 28, 2023
Community
No additional community repositories detected yet.
Hugging Face artifacts
No trustworthy direct or curated related Hugging Face artifacts were found yet. Use targeted searches to quickly locate candidate models, datasets, and demos.
Datasets
Tip: start with models, then check datasets and spaces if you need evaluation data or demos.
Research context
20
Citations
43
References
Tasks
Implementation, Reproducibility, Scrutiny, Computer science, Recommender system, Data science, Information Systems
Methods
Information retrieval
Domains
None detected
Related papers
- The reproducibility of reported height and body weight in repeated questionnaire surveys.Search on Paper2Code
1995 · Semantic similarity
- Short- and long-term reproducibility of radioisotopic examination of gastric emptyingSearch on Paper2Code
1990 · Semantic similarity
- An Overview of Scrutiny: A Triumph of Context over StructureSearch on Paper2Code
2004 · Semantic similarity
- NHS scrutiny. Powers of observation.Search on Paper2Code
2002 · Semantic similarity
- The reproducibility of scores from three video formats.Search on Paper2Code
1988 · Semantic similarity
- Comparisons of probing depth measurements with Florida probe and conventional periodontal probeSearch on Paper2Code
2005 · Semantic similarity
Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.
Open in HFEPXJump to Paper2Code search queries derived from this paper's research context.
Data includes links from Papers with Code ( CC-BY-SA-4.0 ).