| Status | Citation (Found) | Matched Data / Notes | Actions |
|---|---|---|---|
|
Hallucination
Edited
|
Prefill-based jailbreak: A novel approach of bypassing llm safety boundary. Yakai Li; Jiekang Hu; Weiduan Sang; Luping Ma; Jing Xie; Weijuan Zhang; Aimin Yu; Shijie Zhao; Qingjia Huang; Qihang Zhou (2025) arXiv |
Raw: Yakai Li, Jiekang Hu, Weiduan Sang, Luping Ma, Jing Xie, Weijuan Zhang, Aimin Yu, Shijie Zhao, Qingjia Huang, and Qihang Zhou. Prefill-based jailbreak: A novel approach of bypassing llm safety boundary. arXiv preprint arXiv:2504.21038, 2025.
Match: Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
Venue: arXiv DOI: 10.48550/arXiv.2504.21038 ISBN: URL: https://arxiv.org/abs/2504.21038
Identifier resolves to a real record, but the cited title is substantially different from the official title (similarity 0.49) and cited authors do not match — metadata chimera (real id + wrong title/authors).
|
|
|
Minor Error
Edited
|
Let’s verify step by step. Hunter Lightman; Vineet Kosaraju; Yuri Burda; Harrison Edwards; Bowen Baker; Teddy Lee; Jan Leike; John Schulman; Ilya Sutskever; Karl Cobbe (2023) International Conference on Learning Representations (ICLR) |
Raw: Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. In International Conference on Learning Representations (ICLR), 2023.
The paper is a well-known work (arXiv:2305.20050). The citation correctly identifies the work but contains a factual error regarding the conference year (2023 vs 2024) and minor author spelling variations.
|
|
|
Verified
Edited
|
A StrongREJECT for empty jailbreaks. Pieter Abbeel; Dillon Bowen; Scott Emmons; Elvis Hsieh; Qingyuan Lu; Sana Pandey; Alexandra Souly; Justin Svegliato; Sam Toyer; Tu Trinh; Olivia Watkins (2024) Advances in Neural Information Processing Systems (NeurIPS) |
Raw: Pieter Abbeel, Dillon Bowen, Scott Emmons, Elvis Hsieh, Qingyuan Lu, Sana Pandey, Alexandra Souly, Justin Svegliato, Sam Toyer, Tu Trinh, and Olivia Watkins. A StrongREJECT for empty jailbreaks. In Advances in Neural Information Processing Systems (NeurIPS), NeurIPS 2024, page 125416-125440, 2024.
Match: A StrongREJECT for Empty Jailbreaks
Venue: Advances in Neural Information Processing Systems 37 DOI: 10.52202/079017-3984 ISBN: URL: https://doi.org/10.52202/079017-3984
Verified via static CrossRef title search (score: 1.00)
|
|
|
Verified
Edited
|
gpt-oss-120b & gpt-oss-20b model card. Sandhini Agarwal; Lama Ahmad; Jason Ai; Sam Altman; Andy Applebaum; Edwin Arbus; Rahul K Arora; Yu Bai; Bowen Baker; Haiming Bao (2025) arXiv |
Raw: Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K Arora, Yu Bai, Bowen Baker, Haiming Bao, et al. gpt-oss-120b & gpt-oss-20b model card. arXiv preprint arXiv:2508.10925, 2025.
Match: gpt-oss-120b & gpt-oss-20b Model Card
Venue: arXiv DOI: 10.48550/arXiv.2508.10925 ISBN: URL: https://arxiv.org/abs/2508.10925
Verified via arXiv id 2508.10925.
|
|
|
Verified
Edited
|
Does refusal training in llms generalize to the past tense? Maksym Andriushchenko; Nicolas Flammarion (2024) arXiv |
Raw: Maksym Andriushchenko and Nicolas Flammarion. Does refusal training in llms generalize to the past tense? arXiv preprint arXiv:2407.11969, 2024.
Match: Does Refusal Training in LLMs Generalize to the Past Tense?
Venue: arXiv DOI: 10.48550/arXiv.2407.11969 ISBN: URL: https://arxiv.org/abs/2407.11969
Verified via arXiv id 2407.11969.
|
|
|
Verified
Edited
|
Jailbreaking leading safety-aligned LLMs with simple adaptive attacks. Maksym Andriushchenko; Francesco Croce; Nicolas Flammarion (2024) arXiv |
Raw: Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. Jailbreaking leading safety-aligned LLMs with simple adaptive attacks. arXiv preprint arXiv:2404.02151, 2024.
Match: Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Venue: arXiv DOI: 10.48550/arXiv.2404.02151 ISBN: URL: https://arxiv.org/abs/2404.02151
Verified via arXiv id 2404.02151.
|
|
|
Verified
Edited
|
Refusal in llms is mediated by a single direction. Andy Arditi; Oscar Obeso; Aaquib Syed; Daniel Paleka; Nina Panickssery; Wes Gurnee; Neel Nanda (2024) Advances in Neural Information Processing Systems (NeurIPS) |
Raw: Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. Refusal in llms is mediated by a single direction. In Advances in Neural Information Processing Systems (NeurIPS), pages 136037-136083, 2024.
Match: Refusal in Language Models Is Mediated by a Single Direction
Venue: Advances in Neural Information Processing Systems 37 DOI: 10.52202/079017-4322 ISBN: URL: https://doi.org/10.52202/079017-4322
Verified via static CrossRef title search (score: 0.93)
|
|
|
Verified
Edited
|
Open technical problems in open-weight ai model risk management. Stephen Casper; Kyle O’Brien; Shayne Longpre; Elizabeth Seger; Kevin Klyman; Rishi Bommasani; Aniruddha Nrusimha; Ilia Shumailov; Sören Mindermann; Steven Basart (2025) Social Science Research Network |
Raw: Stephen Casper, Kyle O’Brien, Shayne Longpre, Elizabeth Seger, Kevin Klyman, Rishi Bommasani, Aniruddha Nrusimha, Ilia Shumailov, Sören Mindermann, Steven Basart, et al. Open technical problems in open-weight ai model risk management. Social Science Research Network, 2025.
Match: Open Technical Problems in Open-Weight AI Model Risk Management
Venue: DOI: 10.2139/ssrn.5705186 ISBN: URL: https://doi.org/10.2139/ssrn.5705186
Verified via static CrossRef title search (score: 0.99)
|
|
|
Verified
Edited
|
Jailbreaking black box large language models in twenty queries. Patrick Chao; Alexander Robey; Edgar Dobriban; Hamed Hassani; George J Pappas; Eric Wong (2025) IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) |
Raw: Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries. In IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), pages 23-42, 2025.
Match: Jailbreaking Black Box Large Language Models in Twenty Queries
Venue: 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) DOI: 10.1109/satml64287.2025.00010 ISBN: URL: https://doi.org/10.1109/satml64287.2025.00010
Verified via static CrossRef title search (score: 1.00)
|
|
|
Verified
Edited
|
Sockpuppetting: Jailbreaking llms without optimization through output prefix injection. Asen Dotsinski; Panagiotis Eustratiadis (2026) arXiv |
Raw: Asen Dotsinski and Panagiotis Eustratiadis. Sockpuppetting: Jailbreaking llms without optimization through output prefix injection. arXiv preprint arXiv:2601.13359, 2026.
Verified via arXiv identifier 2601.13359.
|
|
|
Verified
Edited
|
The llama 3 herd of models. Abhimanyu Dubey; Abhinav Jauhri; Abhinav Pandey; Abhishek Kadian; Ahmad Al-Dahle; Aiesha Letman; Akhil Mathur; Alan Schelten; Amy Yang; Angela Fan (2024) arXiv |
Raw: Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
Match: The Llama 3 Herd of Models
Venue: arXiv DOI: 10.48550/arXiv.2407.21783 ISBN: URL: https://arxiv.org/abs/2407.21783
Verified via arXiv id 2407.21783.
|
|
|
Verified
Edited
|
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. Daya Guo; Dejian Yang; Haowei Zhang; Junxiao Song; Ruoyu Zhang; Runxin Xu; Qihao Zhu; Shirong Ma; Peiyi Wang; Xiao Bi (2025) arXiv |
Raw: Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025.
Match: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Venue: arXiv DOI: 10.48550/arXiv.2501.12948 ISBN: URL: https://arxiv.org/abs/2501.12948
Verified via arXiv id 2501.12948.
|
|
|
Verified
Edited
|
ClearHarm: A more challenging jailbreak dataset Oskar Hollinsworth; Ian McKenzie; Tom Tseng; Adam Gleave (2025) |
Raw: Oskar Hollinsworth, Ian McKenzie, Tom Tseng, and Adam Gleave. ClearHarm: A more challenging jailbreak dataset, 2025. URL https://huggingface.co/datasets/AlignmentResearch/ClearHarm
Match: ClearHarm: A more challenging jailbreak dataset
Venue: Hugging Face / Alignment Research Center (FAR AI) DOI: ISBN: URL: https://huggingface.co/datasets/AlignmentResearch/ClearHarm
The work is a verified dataset release on Hugging Face.
|
|
|
Verified
Edited
|
Best-of-n jailbreaking. John Hughes; Sara Price; Aengus Lynch; Rylan Schaeffer; Fazl Barez; Sanmi Koyejo; Henry Sleight; Erik Jones; Ethan Perez; Mrinank Sharma (2024) arXiv |
Raw: John Hughes, Sara Price, Aengus Lynch, Rylan Schaeffer, Fazl Barez, Sanmi Koyejo, Henry Sleight, Erik Jones, Ethan Perez, and Mrinank Sharma. Best-of-n jailbreaking. arXiv preprint arXiv:2412.03556, 2024.
Match: Best-of-N Jailbreaking
Venue: arXiv DOI: 10.48550/arXiv.2412.03556 ISBN: URL: https://arxiv.org/abs/2412.03556
Verified via arXiv id 2412.03556.
|
|
|
Verified
Edited
|
Uncensor any llm with abliteration. Maxime Labonne (2024) |
Raw: Maxime Labonne. Uncensor any llm with abliteration. https:// huggingface.co/blog/mlabonne/abliteration
Match: Uncensor any LLM with abliteration
Venue: Hugging Face Blog DOI: ISBN: URL: https://huggingface.co/blog/mlabonne/abliteration
The raw citation contained a formatting error (a space in the URL: 'https:// huggingface.co/...'). This is treated as a verified citation because the intended URL is clear and points to the correct work.
|
|
|
Verified
Edited
|
Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction. Tong Liu; Yingjie Zhang; Zhe Zhao; Yinpeng Dong; Guozhu Meng; Kai Chen (2024) USENIX Security Symposium |
Raw: Tong Liu, Yingjie Zhang, Zhe Zhao, Yinpeng Dong, Guozhu Meng, and Kai Chen. Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction. In USENIX Security Symposium, pages 4711-4728, 2024.
Verified against 33rd USENIX Security Symposium (USENIX Security 24) proceedings.
|
|
|
Verified
Edited
|
AdaPPA: Adaptive position pre-fill jailbreak attack approach targeting LLMs. Lijia Lv; Weigang Zhang; Xuehai Tang; Jie Wen; Feng Liu; Jizhong Han; Songlin Hu (2025) IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) |
Raw: Lijia Lv, Weigang Zhang, Xuehai Tang, Jie Wen, Feng Liu, Jizhong Han, and Songlin Hu. AdaPPA: Adaptive position pre-fill jailbreak attack approach targeting LLMs. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), page 1-5, 2025.
Match: AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
Venue: ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) DOI: 10.1109/icassp49660.2025.10890715 ISBN: URL: https://doi.org/10.1109/icassp49660.2025.10890715
Verified via static CrossRef title search (score: 1.00)
|
|
|
Verified
Edited
|
Artificial intelligence index report 2025. Nestor Maslej; Loredana Fattorini; Raymond Perrault; Yolanda Gil; Vanessa Parli; Njenga Kariuki; Emily Capstick; Anka Reuel; Erik Brynjolfsson; John Etchemendy (2025) arXiv |
Raw: Nestor Maslej, Loredana Fattorini, Raymond Perrault, Yolanda Gil, Vanessa Parli, Njenga Kariuki, Emily Capstick, Anka Reuel, Erik Brynjolfsson, John Etchemendy, et al. Artificial intelligence index report 2025. arXiv preprint arXiv:2504.07139, 2025.
Match: Artificial Intelligence Index Report 2025
Venue: arXiv DOI: 10.48550/arXiv.2504.07139 ISBN: URL: https://arxiv.org/abs/2504.07139
Verified via arXiv id 2504.07139.
|
|
|
Verified
Edited
|
The llama 4 herd. Meta (2025) |
||
|
Verified
Edited
|
Kimi-k2-thinking. Moonshot AI (2025) |
Raw: Moonshot AI. Kimi-k2-thinking. https://moonshotai.github.io/Kimi-K2/thinking.html
Match: Introducing Kimi K2 Thinking
Venue: Moonshot AI DOI: ISBN: URL: https://moonshotai.github.io/Kimi-K2/thinking.html
The citation refers to the official technical blog post and announcement page for the Kimi K2 Thinking model, which is hosted on the official Moonshot AI GitHub Pages domain.
|
|
|
Verified
Edited
|
Introducing gpt-oss-safeguard. OpenAI (2025) |
||
|
Verified
Edited
|
Openai harmony response format. OpenAI (2025) |
Raw: OpenAI. Openai harmony response format. https://developers.openai.com/cookbook/articles/openai-harmony
Match: OpenAI Harmony Response Format
Venue: OpenAI Developer Cookbook DOI: ISBN: URL: https://developers.openai.com/cookbook/articles/openai-harmony
Document confirmed as an official article in the OpenAI developer cookbook.
|
|
|
Verified
Edited
|
Fine-tuning aligned language models compromises safety, even when users do not intend to! Xiangyu Qi; Yi Zeng; Tinghao Xie; Pin-Yu Chen; Ruoxi Jia; Prateek Mittal; Peter Henderson (2023) arXiv |
Raw: Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv preprint arXiv:2310.03693, 2023.
Match: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Venue: arXiv DOI: 10.48550/arXiv.2310.03693 ISBN: URL: https://arxiv.org/abs/2310.03693
Verified via arXiv id 2310.03693.
|
|
|
Verified
Edited
|
Safety alignment should be made more than just a few tokens deep. Xiangyu Qi; Ashwinee Panda; Kaifeng Lyu; Xiao Ma; Subhrajit Roy; Ahmad Beirami; Prateek Mittal; Peter Henderson (2025) International Conference on Learning Representations (ICLR) |
Raw: Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, and Peter Henderson. Safety alignment should be made more than just a few tokens deep. In International Conference on Learning Representations (ICLR), 2025.
The paper was indeed an Outstanding Paper Award recipient at ICLR 2025.
|
|
|
Verified
Edited
|
Gpqa: A graduate-level google-proof q&a benchmark. David Rein; Betty Li Hou; Asa Cooper Stickland; Jackson Petty; Richard Yuanzhe Pang; Julien Dirani; Julian Michael; Samuel R Bowman (2024) Conference on Language Modeling (COLM) |
Raw: David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level google-proof q&a benchmark. In Conference on Language Modeling (COLM), 2024.
Match: GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Venue: COLM 2024 DOI: ISBN: URL: https://openreview.net/submissions?page=8&venue=colmweb.org%2FCOLM%2F2024%2FConference
Upstream report states the citation matches a real COLM 2024 paper and cites OpenReview’s COLM 2024 submissions page and accepted-papers page as evidence.
|
|
|
Verified
Edited
|
An approach to technical agi safety and security. Rohin Shah; Alex Irpan; Alexander Matt Turner; Anna Wang; Arthur Conmy; David Lindner; Jonah Brown-Cohen; Lewis Ho; Neel Nanda; Raluca Ada Popa (2025) arXiv |
Raw: Rohin Shah, Alex Irpan, Alexander Matt Turner, Anna Wang, Arthur Conmy, David Lindner, Jonah Brown-Cohen, Lewis Ho, Neel Nanda, Raluca Ada Popa, et al. An approach to technical agi safety and security. arXiv preprint arXiv:2504.01849, 2025.
Match: An Approach to Technical AGI Safety and Security
Venue: arXiv DOI: 10.48550/arXiv.2504.01849 ISBN: URL: https://arxiv.org/abs/2504.01849
Verified via arXiv id 2504.01849.
|
|
|
Verified
Edited
|
Constitutional classifiers: Defending against universal jailbreaks across thousands of hours of red teaming. Mrinank Sharma; Meg Tong; Jesse Mu; Jerry Wei; Jorrit Kruthoff; Scott Goodfriend; Euan Ong; Alwin Peng; Raj Agarwal; Cem Anil (2025) arXiv |
Raw: Mrinank Sharma, Meg Tong, Jesse Mu, Jerry Wei, Jorrit Kruthoff, Scott Goodfriend, Euan Ong, Alwin Peng, Raj Agarwal, Cem Anil, et al. Constitutional classifiers: Defending against universal jailbreaks across thousands of hours of red teaming. arXiv preprint arXiv:2501.18837, 2025.
Match: Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Venue: arXiv DOI: 10.48550/arXiv.2501.18837 ISBN: URL: https://arxiv.org/abs/2501.18837
Verified via arXiv id 2501.18837.
|
|
|
Verified
Edited
|
” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models. Xinyue Shen; Zeyuan Chen; Michael Backes; Yun Shen; Yang Zhang (2024) ACM SIGSAC Conference on Computer and Communications Security |
Raw: Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. ” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models. In ACM SIGSAC Conference on Computer and Communications Security, pages 1671-1685, 2024.
Match: "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Venue: Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security DOI: 10.1145/3658644.3670388 ISBN: URL: https://doi.org/10.1145/3658644.3670388
Verified via static CrossRef title search (score: 1.00)
|
|
|
Verified
Edited
|
Bypassing the safety training of open-source llms with priming attacks. Jason Vega; Isha Chaudhary; Changming Xu; Gagandeep Singh (2023) arXiv |
Raw: Jason Vega, Isha Chaudhary, Changming Xu, and Gagandeep Singh. Bypassing the safety training of open-source llms with priming attacks. arXiv preprint arXiv:2312.12321, 2023.
Match: Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
Venue: arXiv DOI: 10.48550/arXiv.2312.12321 ISBN: URL: https://arxiv.org/abs/2312.12321
Verified via arXiv id 2312.12321.
|
|
|
Verified
Edited
|
No free lunch for defending against prefilling attack by in-context learning. Zhiyu Xue; Guangliang Liu; Bocheng Chen; Kristen Marie Johnson; Ramtin Pedarsani (2024) arXiv |
Raw: Zhiyu Xue, Guangliang Liu, Bocheng Chen, Kristen Marie Johnson, and Ramtin Pedarsani. No free lunch for defending against prefilling attack by in-context learning. arXiv preprint arXiv:2412.12192, 2024.
Match: No Free Lunch for Defending Against Prefilling Attack by In-Context Learning
Venue: arXiv DOI: 10.48550/arXiv.2412.12192 ISBN: URL: https://arxiv.org/abs/2412.12192
Verified via arXiv id 2412.12192.
|
|
|
Verified
Edited
|
Qwen3 technical report. An Yang; Anfeng Li; Baosong Yang; Beichen Zhang; Binyuan Hui; Bo Zheng; Bowen Yu; Chang Gao; Chengen Huang; Chenxu Lv (2025) arXiv |
Raw: An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025.
Match: Qwen3 Technical Report
Venue: arXiv DOI: 10.48550/arXiv.2505.09388 ISBN: URL: https://arxiv.org/abs/2505.09388
Verified via arXiv id 2505.09388.
|
|
|
Verified
Edited
|
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher. Youliang Yuan; Wenxiang Jiao; Wenxuan Wang; Jen-tse Huang; Pinjia He; Shuming Shi; Zhaopeng Tu (2023) arXiv |
Raw: Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher. arXiv preprint arXiv:2308.06463, 2023.
Match: GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
Venue: arXiv DOI: 10.48550/arXiv.2308.06463 ISBN: URL: https://arxiv.org/abs/2308.06463
Verified via arXiv id 2308.06463.
|
|
|
Verified
Edited
|
Glm-4.7: Advancing the coding capability. Z.ai (2025) |
Raw: Z.ai. Glm-4.7: Advancing the coding capability. https://z.ai/blog/glm-4.7
Match: Glm-4.7: Advancing the coding capability
Venue: Z.ai DOI: ISBN: URL: https://z.ai/blog/glm-4.7
The cited work was verified as a real publication by the organization at the provided URL.
|
|
|
Verified
Edited
|
Removing rlhf protections in gpt-4 via fine-tuning. Qiusi Zhan; Richard Fang; Rohan Bindu; Akul Gupta; Tatsunori B Hashimoto; Daniel Kang (2024) Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies |
Raw: Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori B Hashimoto, and Daniel Kang. Removing rlhf protections in gpt-4 via fine-tuning. In Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 681-687, 2024.
Match: Removing RLHF Protections in GPT-4 via Fine-Tuning
Venue: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers) DOI: 10.18653/v1/2024.naacl-short.59 ISBN: URL: https://doi.org/10.18653/v1/2024.naacl-short.59
Verified via static CrossRef title search (score: 1.00)
|
|
|
Verified
Edited
|
Qwen3guard technical report. Haiquan Zhao; Chenhan Yuan; Fei Huang; Xiaomeng Hu; Yichang Zhang; An Yang; Bowen Yu; Dayiheng Liu; Jingren Zhou; Junyang Lin (2025) arXiv |
Raw: Haiquan Zhao, Chenhan Yuan, Fei Huang, Xiaomeng Hu, Yichang Zhang, An Yang, Bowen Yu, Dayiheng Liu, Jingren Zhou, Junyang Lin, et al. Qwen3guard technical report. arXiv preprint arXiv:2510.14276, 2025.
Match: Qwen3Guard Technical Report
Venue: arXiv DOI: 10.48550/arXiv.2510.14276 ISBN: URL: https://arxiv.org/abs/2510.14276
Verified via arXiv id 2510.14276.
|
|
|
Verified
Edited
|
Autodan: interpretable gradient-based adversarial attacks on large language models. Sicheng Zhu; Ruiyi Zhang; Bang An; Gang Wu; Joe Barrow; Zichao Wang; Furong Huang; Ani Nenkova; Tong Sun (2023) arXiv |
Raw: Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun. Autodan: interpretable gradient-based adversarial attacks on large language models. arXiv preprint arXiv:2310.15140, 2023.
Match: AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models
Venue: arXiv DOI: 10.48550/arXiv.2310.15140 ISBN: URL: https://arxiv.org/abs/2310.15140
Verified via arXiv id 2310.15140.
|
|
|
Verified
Edited
|
Universal and transferable adversarial attacks on aligned language models. Andy Zou; Zifan Wang; Nicholas Carlini; Milad Nasr; J Zico Kolter; Matt Fredrikson (2023) arXiv |
Raw: Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023.
Match: Universal and Transferable Adversarial Attacks on Aligned Language Models
Venue: arXiv DOI: 10.48550/arXiv.2307.15043 ISBN: URL: https://arxiv.org/abs/2307.15043
Verified via arXiv id 2307.15043.
|
|
|
Verified
Edited
|
Chemical Weapons Convention None (1997) ISBN: 978-1-280-28774-9 |
Raw: Chemical Weapons Convention (CWC) (1997)
The citation 'Chemical Weapons Convention (CWC) (1997)' refers to the treaty that entered into force on April 29, 1997. Citing by common name, acronym, and EIF year is standard. The provided ISBN (9781280287749) is a valid identifier for an edition of this text.
|
|
|
Verified
Edited
|
Rome Statute of the International Criminal Court None ISBN: 978-1-849-46996-8 |
Raw: Rome Statute of the International Criminal Court
The raw citation correctly identifies the treaty. The provided ISBN for a commentary book was an external metadata error not present in the citation text, which is ignored.
|
|
|
Verified
Edited
|
Norm-preserving biprojected abliteration. Jim Lai (2025) |
Raw: Jim Lai. Norm-preserving biprojected abliteration. https://huggingface.co/blog/grimjim/norm-preserving-biprojected-abliteration
Match: Norm-Preserving Biprojected Abliteration
Venue: Hugging Face Blog DOI: ISBN: URL: https://huggingface.co/blog/grimjim/norm-preserving-biprojected-abliteration |