2603.16417v1.pdf

Status: Completed
9
Pages
19
Total Citations
0
Verified
0
Hallucinations / Not Found
0
Duplicates
Notes
It is this paper https://arxiv.org/pdf/2603.16417, Via Negativa for AI Alignment: Why Negative Constraints Are Structurally Superior to Positive Preferences, by Quan Cheng
Citation Details
Filter:
Status Citation (Found) Matched Data / Notes Actions
Hallucination
Edited
The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
Liu, Y.; Zeng, Z.; et al. (2025)
Advances in Neural Information Processing Systems (NeurIPS)
Raw: [1] Liu, Y., Zeng, Z., et al. (2025). “The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning.” Advances in Neural Information Processing Systems (NeurIPS). arXiv:2506.01347.
Match: The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
Authors: Xinyu Zhu; Mengzhou Xia; Zhepei Wei; Wei-Lin Chen; Danqi Chen; Yu Meng
Venue: arXiv
DOI: 10.48550/arXiv.2506.01347

ISBN:

URL: https://arxiv.org/abs/2506.01347
CrossRef arxiv_static matches title/DOI (score: 1.00), but cited author identities disagree with the official list. Unmatched cited author(s): Liu, Y., Zeng, Z.. Official authors: Xinyu Zhu, Mengzhou Xia, Zhepei Wei, Wei-Lin Chen, Danqi Chen, Yu Meng.
Hallucination
Edited
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
Duan, H.; Yi, Y.; Zhang, Z.; Liu, F.; et al. (2024)
Findings of EMNLP
Raw: [2] Duan, H., Yi, Y., Zhang, Z., Liu, F., et al. (2024). “Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization.” Findings of EMNLP. arXiv:2403.03419.
Match: Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
Authors: Shitong Duan; Xiaoyuan Yi; Peng Zhang; Yan Liu; Zheng Liu; Tun Lu; Xing Xie; Ning Gu
Venue: arXiv
DOI: 10.48550/arXiv.2403.03419

ISBN:

URL: https://arxiv.org/abs/2403.03419
CrossRef arxiv_static matches title/DOI (score: 1.00), but cited author identities disagree with the official list. Unmatched cited author(s): Duan, H., Yi, Y., Zhang, Z., Liu, F.. Official authors: Shitong Duan, Xiaoyuan Yi, Peng Zhang, Yan Liu, Zheng Liu, Tun Lu, Xing Xie, Ning Gu.
Hallucination
Edited
How RLHF Amplifies Sycophancy
Shapira, N.; Levy, M.; Alavi, S. H.; et al. (2026)
arXiv
Raw: [6] Shapira, N., Levy, M., Alavi, S. H., et al. (2026). “How RLHF Amplifies Sycophancy.” arXiv preprint arXiv:2602.01002.
Match: How RLHF Amplifies Sycophancy
Authors: Itai Shapira; Gerdus Benade; Ariel D. Procaccia
Venue: arXiv
DOI: 10.48550/arXiv.2602.01002

ISBN:

URL: https://arxiv.org/abs/2602.01002
CrossRef arxiv_static matches title/DOI (score: 1.00), but cited author identities disagree with the official list. Unmatched cited author(s): Shapira, N., Levy, M., Alavi, S. H.. Official authors: Itai Shapira, Gerdus Benade, Ariel D. Procaccia.
Minor Error
Edited
Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
Zhang, J.; et al. (2024)
arXiv
Raw: [3] Zhang, J., et al. (2024). “Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning.” arXiv preprint arXiv:2404.05868.
Match: Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
Authors: Ruiqi Zhang; Licong Lin; Yu Bai; Song Mei
Venue: arXiv
DOI: 10.48550/arXiv.2404.05868

ISBN:

URL: https://arxiv.org/abs/2404.05868
CrossRef arxiv_static matches title/DOI (score: 1.00), but cited author identities disagree with the official list. Unmatched cited author(s): Zhang, J.. Official authors: Ruiqi Zhang, Licong Lin, Yu Bai, Song Mei.
Verified
Edited
KTO: Model Alignment as Prospect Theoretic Optimization
Ethayarajh, K.; Xu, W.; Muennighoff, N.; Jurafsky, D.; and Kiela, D. (2024)
Proceedings of ICML
Raw: [4] Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D. (2024). “KTO: Model Alignment as Prospect Theoretic Optimization.” Proceedings of ICML. arXiv:2402.01306.
Match: KTO: Model Alignment as Prospect Theoretic Optimization
Authors: Kawin Ethayarajh; Winnie Xu; Niklas Muennighoff; Dan Jurafsky; Douwe Kiela
Venue: arXiv
DOI: 10.48550/arXiv.2402.01306

ISBN:

URL: https://arxiv.org/abs/2402.01306
Verified via arXiv id 2402.01306.
Verified
Edited
Towards Understanding Sycophancy in Language Models
Sharma, M.; Tong, M.; Korbak, T.; et al. (2024)
Proceedings of ICLR
Raw: [5] Sharma, M., Tong, M., Korbak, T., et al. (2024). “Towards Understanding Sycophancy in Language Models.” Proceedings of ICLR. arXiv:2310.13548.
Match: Towards Understanding Sycophancy in Language Models
Authors: Mrinank Sharma; Meg Tong; Tomasz Korbak; David Duvenaud; Amanda Askell; Samuel R. Bowman; Newton Cheng; Esin Durmus; Zac Hatfield-Dodds; Scott R. Johnston; Shauna Kravec; Timothy Maxwell; Sam McCandlish; Kamal Ndousse; Oliver Rausch; Nicholas Schiefer; Da Yan; Miranda Zhang; Ethan Perez
Venue: arXiv
DOI: 10.48550/arXiv.2310.13548

ISBN:

URL: https://arxiv.org/abs/2310.13548
Verified via arXiv id 2310.13548.
Verified
Edited
Why the Valuable Capabilities of LLMs Are Precisely the Unexplainable Ones
Cheng, Q. (2026)
arXiv
Raw: [7] Cheng, Q. (2026). “Why the Valuable Capabilities of LLMs Are Precisely the Unexplainable Ones.” arXiv preprint.
Match: Why the Valuable Capabilities of LLMs Are Precisely the Unexplainable Ones
Authors: Q. Cheng
Venue: arXiv:2603.15238
DOI:

ISBN:

URL: https://arxiv.org/abs/2603.15238
The title, author, and year are consistent with the official arXiv record.
Verified
Edited
On the Proper Treatment of Connectionism
Smolensky, P. (1988)
Behavioral and Brain Sciences
Raw: [8] Smolensky, P. (1988). “On the Proper Treatment of Connectionism.” Behavioral and Brain Sciences, 11(1), 1-23.
Match: On the proper treatment of connectionism
Authors: Paul Smolensky
Venue: Behavioral and Brain Sciences
DOI: 10.1017/s0140525x00052432

ISBN:

URL: https://doi.org/10.1017/s0140525x00052432
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
The Logic of Scientific Discovery
Popper, K. R. (1959)
Routledge
ISBN: 978-0-415-27844-7
Raw: [9] Popper, K. R. (1959). The Logic of Scientific Discovery. Routledge. (Original: Logik der Forschung, 1934.)
Match: The Logic of Scientific Discovery
Authors: Richard C. Jeffrey; Karl R. Popper
Venue: Econometrica
DOI: 10.2307/1907578

ISBN: 978-0-415-27844-7

URL:
Verified via static CrossRef title search (score: 0.99)
Verified
Edited
Antifragile: Things That Gain from Disorder
Taleb, N. N. (2012)
Random House
Raw: [10] Taleb, N. N. (2012). Antifragile: Things That Gain from Disorder. Random House. Chapter 22: Via Negativa. 8
Match: Antifragile: Things That Gain from Disorder
Authors: Nassim Nicholas Taleb
Venue: Random House
DOI:

ISBN:

URL:
The citation references a specific chapter ('Via Negativa') within the identified book.
Verified
Edited
Negative Knowledge: Understanding Professional Learning and Expertise
Gartmeier, M.; Bauer, J.; Gruber, H.; and Heid, H. (2008)
Vocations and Learning
Raw: [11] Gartmeier, M., Bauer, J., Gruber, H., and Heid, H. (2008). “Negative Knowledge: Understanding Professional Learning and Expertise.” Vocations and Learning, 1(2), 87-103.
Match: Negative Knowledge: Understanding Professional Learning and Expertise
Authors: Martin Gartmeier; Johannes Bauer; Hans Gruber; Helmut Heid
Venue: Vocations and Learning
DOI: 10.1007/s12186-008-9006-1

ISBN:

URL: https://doi.org/10.1007/s12186-008-9006-1
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
Constitutional AI: Harmlessness from AI Feedback
Bai, Y.; Kadavath, S.; et al. (2022)
arXiv
Raw: [12] Bai, Y., Kadavath, S., et al. (2022). “Constitutional AI: Harmlessness from AI Feedback.” arXiv preprint arXiv:2212.08073.
Match: Constitutional AI: Harmlessness from AI Feedback
Authors: Yuntao Bai; Saurav Kadavath; Sandipan Kundu; Amanda Askell; Jackson Kernion; Andy Jones; Anna Chen; Anna Goldie; Azalia Mirhoseini; Cameron McKinnon; Carol Chen; Catherine Olsson; Christopher Olah; Danny Hernandez; Dawn Drain; Deep Ganguli; Dustin Li; Eli Tran-Johnson; Ethan Perez; Jamie Kerr; Jared Mueller; Jeffrey Ladish; Joshua Landau; Kamal Ndousse; Kamile Lukosuite; Liane Lovitt; Michael Sellitto; Nelson Elhage; Nicholas Schiefer; Noemi Mercado; Nova DasSarma; Robert Lasenby; Robin Larson; Sam Ringer; Scott Johnston; Shauna Kravec; Sheer El Showk; Stanislav Fort; Tamera Lanham; Timothy Telleen-Lawton; Tom Conerly; Tom Henighan; Tristan Hume; Samuel R. Bowman; Zac Hatfield-Dodds; Ben Mann; Dario Amodei; Nicholas Joseph; Sam McCandlish; Tom Brown; Jared Kaplan
Venue: arXiv
DOI: 10.48550/arXiv.2212.08073

ISBN:

URL: https://arxiv.org/abs/2212.08073
Verified via arXiv id 2212.08073.
Verified
Edited
Large Language Model Unlearning
Yao, Y.; et al. (2024)
Advances in Neural Information Processing Systems (NeurIPS)
Raw: [13] Yao, Y., et al. (2024). “Large Language Model Unlearning.” Advances in Neural Information Processing Systems (NeurIPS). arXiv:2310.10683.
Match: Large Language Model Unlearning
Authors: Yuanshun Yao; Xiaojun Xu; Yang Liu
Venue: Advances in Neural Information Processing Systems 37
DOI: 10.52202/079017-3346

ISBN:

URL: https://doi.org/10.52202/079017-3346
Verified via static CrossRef title search (score: 0.99)
Verified
Edited
Mind over Machine: The Power of Human Intuition and Expertise in the Era of the Computer
Dreyfus, H. L.; and Dreyfus, S. E. (1986)
Free Press
Raw: [14] Dreyfus, H. L. and Dreyfus, S. E. (1986). Mind over Machine: The Power of Human Intuition and Expertise in the Era of the Computer. Free Press.
Match: Mind over Machine: The Power of Human Intuition and Expertise in the Era of the Computer
Authors: Hubert L. Dreyfus; Stuart E. Drey-fus; Lotfi A. Zadeh
Venue: IEEE Expert
DOI: 10.1109/mex.1987.4307079

ISBN:

URL: https://doi.org/10.1109/mex.1987.4307079
Verified via static CrossRef title search (score: 0.99)
Verified
Edited
Discovering Language Model Behaviors with Model-Written Evaluations
Perez, E.; Ringer, S.; et al. (2023)
Findings of ACL
Raw: [15] Perez, E., Ringer, S., et al. (2023). “Discovering Language Model Behaviors with Model-Written Evaluations.” Findings of ACL. arXiv:2212.09251.
Match: Discovering Language Model Behaviors with Model-Written Evaluations
Authors: Ethan Perez; Sam Ringer; Kamile Lukosiute; Karina Nguyen; Edwin Chen; Scott Heiner; Craig Pettit; Catherine Olsson; Sandipan Kundu; Saurav Kadavath; Andy Jones; Anna Chen; Benjamin Mann; Brian Israel; Bryan Seethor; Cameron McKinnon; Christopher Olah; Da Yan; Daniela Amodei; Dario Amodei; Dawn Drain; Dustin Li; Eli Tran-Johnson; Guro Khundadze; Jackson Kernion; James Landis; Jamie Kerr; Jared Mueller; Jeeyoon Hyun; Joshua Landau; Kamal Ndousse; Landon Goldberg; Liane Lovitt; Martin Lucas; Michael Sellitto; Miranda Zhang; Neerav Kingsland; Nelson Elhage; Nicholas Joseph; Noemi Mercado; Nova DasSarma; Oliver Rausch; Robin Larson; Sam McCandlish; Scott Johnston; Shauna Kravec; Sheer El Showk; Tamera Lanham; Timothy Telleen-Lawton; Tom Brown; Tom Henighan; Tristan Hume; Yuntao Bai; Zac Hatfield-Dodds; Jack Clark; Samuel R. Bowman; Amanda Askell; Roger Grosse; Danny Hernandez; Deep Ganguli; Evan Hubinger; Nicholas Schiefer; Jared Kaplan
Venue: Findings of the Association for Computational Linguistics: ACL 2023
DOI: 10.18653/v1/2023.findings-acl.847

ISBN:

URL: https://doi.org/10.18653/v1/2023.findings-acl.847
Verified via static CrossRef title search (score: 0.99)
Verified
Edited
Simple Synthetic Data Reduces Sycophancy in Large Language Models
Wei, J.; et al. (2023)
arXiv
Raw: [16] Wei, J., et al. (2023). “Simple Synthetic Data Reduces Sycophancy in Large Language Models.” arXiv preprint arXiv:2308.03958.
Match: Simple synthetic data reduces sycophancy in large language models
Authors: Jerry Wei; Da Huang; Yifeng Lu; Denny Zhou; Quoc V. Le
Venue: arXiv
DOI: 10.48550/arXiv.2308.03958

ISBN:

URL: https://arxiv.org/abs/2308.03958
Verified via arXiv id 2308.03958.
Verified
Edited
BNF: As Simple as Fine-tuning: LLM Alignment via Bidirectional Negative Feedback Loss
Han, Y.; et al. (2024)
OpenReview
Raw: [17] Han, Y., et al. (2024). “BNF: As Simple as Fine-tuning: LLM Alignment via Bidirectional Negative Feedback Loss.” OpenReview.
Match: BNF: As Simple as Fine-tuning: LLM Alignment via Bidirectional Negative Feedback Loss
Authors:
Venue:
DOI:

ISBN:

URL:
The work exists and is typically associated with preprint or conference submission platforms such as OpenReview or ArXiv.
Verified
Edited
Negative Knowledge, Expertise and Organisations
Parviainen, J.; and Eriksson, M. (2006)
International Journal of Management Concepts and Philosophy
Raw: [18] Parviainen, J. and Eriksson, M. (2006). “Negative Knowledge, Expertise and Organisations.” International Journal of Management Concepts and Philosophy, 2(2), 140-153.
Match: Negative knowledge, expertise and organisations
Authors: Jaana Parviainen; Marja Eriksson
Venue: International Journal of Management Concepts and Philosophy
DOI: 10.1504/ijmcp.2006.010265

ISBN:

URL: https://doi.org/10.1504/ijmcp.2006.010265
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
Polanyi’s Revenge and AI’s New Romance with Tacit Knowledge
Kambhampati, S. (2021)
Communications of the ACM
Raw: [19] Kambhampati, S. (2021). “Polanyi’s Revenge and AI’s New Romance with Tacit Knowledge.” Communications of the ACM, 64(10), 31-33.
Match: Polanyi's revenge and AI's new romance with tacit knowledge
Authors: Subbarao Kambhampati
Venue: Communications of the ACM
DOI: 10.1145/3446369

ISBN:

URL: https://doi.org/10.1145/3446369
Verified via static CrossRef title search (score: 1.00)