SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

Almene De Meran Meguimtsop, Maria Leonor Pacheco, Daniel E. Acuna

Preprint / working paper · arXiv preprint arXiv:2605.29468 ·

DOI: 10.48550/arXiv.2605.29468

Do language models uphold research-integrity norms when misconduct is disguised?

SciIntBench evaluates how language models respond to scientific requests framed as explicit misconduct, covert misconduct, or legitimate work. It measures both refusal of problematic requests and helpfulness on benign ones.

What the study found

  • Models refused overt misconduct more reliably than covert violations.
  • The study reports weaker boundaries in areas including transparency, plagiarism, and fabrication.

How the study works

The benchmark contains 810 prompts across ten responsible-conduct categories and three scientific domains. The paper evaluates 16 models and 12,960 responses.

Scope and limitations

  • Results reflect the tested model versions and prompts, which may differ from later deployments.
  • A benchmark response is evidence about model behavior in that setting, not a measurement of real-world misconduct.

Abstract

Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (RCR) norms or help undermine them. We introduce SciIntBench, an adversarial benchmark of 810 prompts across ten RCR categories and three scientific domains. Each scenario appears as an Overt Adversarial, Covert Adversarial, and Benign version, allowing us to jointly measure framing-sensitive refusal of misconduct and helpfulness on legitimate requests. We evaluate 16 commercial and open-weight LLMs from six providers (2024--2026), producing 12,960 responses. We find that scientific integrity alignment is strongly framing-sensitive: models refuse explicit misconduct far more reliably than covert violations, especially failing when misconduct is presented as a pressure-driven shortcut. Refusals vary by RCR category, with weaker boundaries around transparency, plagiarism, and fabrication.

Abstract from the original work, reproduced under its Creative Commons license. The overview above summarizes the study.

Cite this work

Almene De Meran Meguimtsop, Maria Leonor Pacheco, Daniel E. Acuna (2026). SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing. arXiv preprint arXiv:2605.29468. https://doi.org/10.48550/arXiv.2605.29468

Download BibTeX

View BibTeX
@article{meguimtsop2026sciintbench,
  title = {SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing},
  author = {Meguimtsop, Almene De Meran and Pacheco, Maria Leonor and Acuna, Daniel E.},
  year = {2026},
  publication_date = {2026-05-28},
  journal = {arXiv preprint arXiv:2605.29468},
  doi = {10.48550/arXiv.2605.29468},
  url = {https://arxiv.org/abs/2605.29468}
}

Overview checked September 7, 2026 against the publication record. Publication and preprint dates refer to the linked versions.

The locally hosted PDF is an unchanged copy from the original source, shared under its Creative Commons license. Copyright remains with the credited authors or rights holders.