REM-CTX: Automated Peer Review via Reinforcement Learning with Auxiliary Context

Pawin Taechoyotin, Daniel E. Acuna

Preprint / working paper · arXiv preprint arXiv:2604.00248 ·

DOI: 10.48550/arXiv.2604.00248

Can an automated review make better use of figures and scholarly context?

REM-CTX extends review generation beyond manuscript text. It trains a language model to use auxiliary context and tests whether explicit correspondence rewards improve the grounding of generated reviews.

What the study found

  • The full model outperformed the reported baselines on the study's overall review-quality evaluation.
  • Ablations found complementary contributions from the two correspondence rewards.

How the study works

An eight-billion-parameter model is trained with Group Relative Policy Optimization and quality/correspondence rewards, then evaluated on manuscripts across computer, biological, and physical sciences.

Scope and limitations

  • The evaluation concerns the specified metrics and manuscript collections, rather than every editorial setting.
  • This preprint evaluates review-generation behavior; it does not establish that automated reviews can replace human judgment.

Abstract

Most automated peer review systems rely on textual manuscript content alone, leaving visual elements such as figures and external scholarly signals underutilized. We introduce REM-CTX, a reinforcement-learning system that incorporates auxiliary context into the review generation process via correspondence-aware reward functions. REM-CTX trains an 8B-parameter language model with Group Relative Policy Optimization (GRPO) and combines a multi-aspect quality reward with two correspondence rewards that explicitly encourage alignment with auxiliary context. Experiments on manuscripts across Computer, Biological, and Physical Sciences show that REM-CTX achieves the highest overall review quality among six baselines, outperforming other systems with substantially larger commercial models, and surpassing the next-best RL baseline across both quality and contextual grounding metrics. Ablation studies confirm that the two correspondence rewards are complementary: each selectively improves its targeted correspondence reward while preserving all quality dimensions, and the full model outperforms all partial variants. Analysis of training dynamics reveals that the criticism aspect is negatively correlated with other metrics during training, suggesting that future studies should group multi-dimension rewards for review generation.

Abstract from the original work, reproduced under its Creative Commons license. The overview above summarizes the study.

Cite this work

Pawin Taechoyotin, Daniel E. Acuna (2026). REM-CTX: Automated Peer Review via Reinforcement Learning with Auxiliary Context. arXiv preprint arXiv:2604.00248. https://doi.org/10.48550/arXiv.2604.00248

Download BibTeX

View BibTeX
@article{taechoyotin2026remctx,
  title = {REM-CTX: Automated Peer Review via Reinforcement Learning with Auxiliary Context},
  author = {Taechoyotin, Pawin and Acuna, Daniel E.},
  year = {2026},
  publication_date = {2026-03-31},
  journal = {arXiv preprint arXiv:2604.00248},
  doi = {10.48550/arXiv.2604.00248},
  url = {https://arxiv.org/abs/2604.00248}
}

Overview checked September 7, 2026 against the publication record. Publication and preprint dates refer to the linked versions.