2504.21038v1.pdf

Status: Completed
13
Pages
32
Total Citations
0
Verified
0
Hallucinations / Not Found
0
Duplicates
Notes
It is this paper, https://arxiv.org/abs/2504.21038v1
Citation Details
Filter:
Status Citation (Found) Matched Data / Notes Actions
Minor Error
Edited
Prefill Claude’s response for greater output control.
Anthropic (2025)
Raw: 4. Anthropic: Prefill Claude’s response for greater output control. https:// web.archive.org/web/20250221204158/https://docs.anthropic.com/en/docs/12AuthorsSuppressedDue
Match: Prefill Claude’s response for greater output control
Authors: Anthropic
Venue: Anthropic Documentation
DOI:

ISBN:

URL: https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/prefill-claudes-response
The citation identifies a real Anthropic documentation page. However, the provided URL was broken due to an apparent copy-paste or extraction error (the string '12AuthorsSuppressedDuetoExcessiveLength' was inserted into the URL path). The intended URL is provided.
Minor Error
Edited
Guard: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models.
Jin, H.; Chen, R.; Zhou, A.; Zhang, Y.; Wang, H. (2024)
arXiv
Raw: 15. Jin, H., Chen, R., Zhou, A., Zhang, Y., Wang, H.: Guard: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models. arXiv preprint arXiv:2402.03299 (2024)
Match: GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models
Authors: Haibo Jin; Ruoxi Chen; Peiyan Zhang; Andy Zhou; Haohan Wang
Venue: arXiv
DOI: 10.48550/arXiv.2402.03299

ISBN:

URL: https://arxiv.org/abs/2402.03299
CrossRef arxiv_static matches title/DOI (score: 1.00), but cited author identities disagree with the official list. Unmatched cited author(s): Zhang, Y.. Official authors: Haibo Jin, Ruoxi Chen, Peiyan Zhang, Andy Zhou, Haohan Wang.
Verified
Edited
Gpt-4 technical report
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F.L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. (2023)
arXiv
Raw: 1. Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al.: Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
Match: GPT-4 Technical Report
Authors: OpenAI; Josh Achiam; Steven Adler; Sandhini Agarwal; Lama Ahmad; Ilge Akkaya; Florencia Leoni Aleman; Diogo Almeida; Janko Altenschmidt; Sam Altman; Shyamal Anadkat; Red Avila; Igor Babuschkin; Suchir Balaji; Valerie Balcom; Paul Baltescu; Haiming Bao; Mohammad Bavarian; Jeff Belgum; Irwan Bello; Jake Berdine; Gabriel Bernadett-Shapiro; Christopher Berner; Lenny Bogdonoff; Oleg Boiko; Madelaine Boyd; Anna-Luisa Brakman; Greg Brockman; Tim Brooks; Miles Brundage; Kevin Button; Trevor Cai; Rosie Campbell; Andrew Cann; Brittany Carey; Chelsea Carlson; Rory Carmichael; Brooke Chan; Che Chang; Fotis Chantzis; Derek Chen; Sully Chen; Ruby Chen; Jason Chen; Mark Chen; Ben Chess; Chester Cho; Casey Chu; Hyung Won Chung; Dave Cummings; Jeremiah Currier; Yunxing Dai; Cory Decareaux; Thomas Degry; Noah Deutsch; Damien Deville; Arka Dhar; David Dohan; Steve Dowling; Sheila Dunning; Adrien Ecoffet; Atty Eleti; Tyna Eloundou; David Farhi; Liam Fedus; Niko Felix; Simón Posada Fishman; Juston Forte; Isabella Fulford; Leo Gao; Elie Georges; Christian Gibson; Vik Goel; Tarun Gogineni; Gabriel Goh; Rapha Gontijo-Lopes; Jonathan Gordon; Morgan Grafstein; Scott Gray; Ryan Greene; Joshua Gross; Shixiang Shane Gu; Yufei Guo; Chris Hallacy; Jesse Han; Jeff Harris; Yuchen He; Mike Heaton; Johannes Heidecke; Chris Hesse; Alan Hickey; Wade Hickey; Peter Hoeschele; Brandon Houghton; Kenny Hsu; Shengli Hu; Xin Hu; Joost Huizinga; Shantanu Jain; Shawn Jain; Joanne Jang; Angela Jiang; Roger Jiang; Haozhun Jin; Denny Jin; Shino Jomoto; Billie Jonn; Heewoo Jun; Tomer Kaftan; Łukasz Kaiser; Ali Kamali; Ingmar Kanitscheider; Nitish Shirish Keskar; Tabarak Khan; Logan Kilpatrick; Jong Wook Kim; Christina Kim; Yongjik Kim; Jan Hendrik Kirchner; Jamie Kiros; Matt Knight; Daniel Kokotajlo; Łukasz Kondraciuk; Andrew Kondrich; Aris Konstantinidis; Kyle Kosic; Gretchen Krueger; Vishal Kuo; Michael Lampe; Ikai Lan; Teddy Lee; Jan Leike; Jade Leung; Daniel Levy; Chak Ming Li; Rachel Lim; Molly Lin; Stephanie Lin; Mateusz Litwin; Theresa Lopez; Ryan Lowe; Patricia Lue; Anna Makanju; Kim Malfacini; Sam Manning; Todor Markov; Yaniv Markovski; Bianca Martin; Katie Mayer; Andrew Mayne; Bob McGrew; Scott Mayer McKinney; Christine McLeavey; Paul McMillan; Jake McNeil; David Medina; Aalok Mehta; Jacob Menick; Luke Metz; Andrey Mishchenko; Pamela Mishkin; Vinnie Monaco; Evan Morikawa; Daniel Mossing; Tong Mu; Mira Murati; Oleg Murk; David Mély; Ashvin Nair; Reiichiro Nakano; Rajeev Nayak; Arvind Neelakantan; Richard Ngo; Hyeonwoo Noh; Long Ouyang; Cullen O'Keefe; Jakub Pachocki; Alex Paino; Joe Palermo; Ashley Pantuliano; Giambattista Parascandolo; Joel Parish; Emy Parparita; Alex Passos; Mikhail Pavlov; Andrew Peng; Adam Perelman; Filipe de Avila Belbute Peres; Michael Petrov; Henrique Ponde de Oliveira Pinto; Michael; Pokorny; Michelle Pokrass; Vitchyr H. Pong; Tolly Powell; Alethea Power; Boris Power; Elizabeth Proehl; Raul Puri; Alec Radford; Jack Rae; Aditya Ramesh; Cameron Raymond; Francis Real; Kendra Rimbach; Carl Ross; Bob Rotsted; Henri Roussez; Nick Ryder; Mario Saltarelli; Ted Sanders; Shibani Santurkar; Girish Sastry; Heather Schmidt; David Schnurr; John Schulman; Daniel Selsam; Kyla Sheppard; Toki Sherbakov; Jessica Shieh; Sarah Shoker; Pranav Shyam; Szymon Sidor; Eric Sigler; Maddie Simens; Jordan Sitkin; Katarina Slama; Ian Sohl; Benjamin Sokolowsky; Yang Song; Natalie Staudacher; Felipe Petroski Such; Natalie Summers; Ilya Sutskever; Jie Tang; Nikolas Tezak; Madeleine B. Thompson; Phil Tillet; Amin Tootoonchian; Elizabeth Tseng; Preston Tuggle; Nick Turley; Jerry Tworek; Juan Felipe Cerón Uribe; Andrea Vallone; Arun Vijayvergiya; Chelsea Voss; Carroll Wainwright; Justin Jay Wang; Alvin Wang; Ben Wang; Jonathan Ward; Jason Wei; CJ Weinmann; Akila Welihinda; Peter Welinder; Jiayi Weng; Lilian Weng; Matt Wiethoff; Dave Willner; Clemens Winter; Samuel Wolrich; Hannah Wong; Lauren Workman; Sherwin Wu; Jeff Wu; Michael Wu; Kai Xiao; Tao Xu; Sarah Yoo; Kevin Yu; Qiming Yuan; Wojciech Zaremba; Rowan Zellers; Chong Zhang; Marvin Zhang; Shengjia Zhao; Tianhao Zheng; Juntang Zhuang; William Zhuk; Barret Zoph
Venue: arXiv
DOI: 10.48550/arXiv.2303.08774

ISBN:

URL: https://arxiv.org/abs/2303.08774
Verified via arXiv id 2303.08774.
Verified
Edited
Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills.
Agrawal, A.; Panwar, A.; Mohan, J.; Kwatra, N.; Gulavani, B.S.; Ramjee, R. (2023)
arXiv
Raw: 2. Agrawal, A., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B.S., Ramjee, R.: Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills. arXiv preprint arXiv:2308.16369 (2023)
Match: SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Authors: Amey Agrawal; Ashish Panwar; Jayashree Mohan; Nipun Kwatra; Bhargav S. Gulavani; Ramachandran Ramjee
Venue: arXiv
DOI: 10.48550/arXiv.2308.16369

ISBN:

URL: https://arxiv.org/abs/2308.16369
Verified via arXiv id 2308.16369.
Verified
Edited
Claude 3.7 sonnet system card
Anthropic (2025)
Anthropic
Raw: 3. Anthropic: Claude 3.7 sonnet system card. System card, Anthropic (2025), https://assets.anthropic.com/m/785e231869ea8b3b/original/claude-3-7-sonnet-system-card.pdf
Match: Claude 3.7 Sonnet System Card
Authors: Anthropic
Venue: Anthropic
DOI:

ISBN:

URL:
Verified as an official technical document published by Anthropic.
Verified
Edited
Jailbreaking black box large language models in twenty queries.
Chao, P.; Robey, A.; Dobriban, E.; Hassani, H.; Pappas, G.J.; Wong, E. (2023)
arXiv
Raw: 5. Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G.J., Wong, E.: Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419 (2023)
Match: Jailbreaking Black Box Large Language Models in Twenty Queries
Authors: Patrick Chao; Alexander Robey; Edgar Dobriban; Hamed Hassani; George J. Pappas; Eric Wong
Venue: arXiv
DOI: 10.48550/arXiv.2310.08419

ISBN:

URL: https://arxiv.org/abs/2310.08419
Verified via arXiv id 2310.08419.
Verified
Edited
The sifo benchmark: Investigating the sequential instruction following ability of large language models.
Chen, X.; Liao, B.; Qi, J.; Eustratiadis, P.; Monz, C.; Bisazza, A.; de Rijke, M. (2024)
arXiv
Raw: 7. Chen, X., Liao, B., Qi, J., Eustratiadis, P., Monz, C., Bisazza, A., de Rijke, M.: The sifo benchmark: Investigating the sequential instruction following ability of large language models. arXiv preprint arXiv:2406.19999 (2024)
Match: The SIFo Benchmark: Investigating the Sequential Instruction Following Ability of Large Language Models
Authors: Xinyi Chen; Baohao Liao; Jirui Qi; Panagiotis Eustratiadis; Christof Monz; Arianna Bisazza; Maarten de Rijke
Venue: Findings of the Association for Computational Linguistics: EMNLP 2024
DOI: 10.18653/v1/2024.findings-emnlp.92

ISBN:

URL: https://doi.org/10.18653/v1/2024.findings-emnlp.92
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
Chat Prefix Completion (Beta).
DeepSeek (2025)
Raw: 8. DeepSeek: Chat Prefix Completion (Beta). https://web.archive.org/web
Match: Chat Prefix Completion (Beta)
Authors: DeepSeek
Venue: DeepSeek API Documentation
DOI:

ISBN:

URL: https://api-docs.deepseek.com/guides/chat_prefix_completion
The citation references a real, identifiable technical guide from DeepSeek's official documentation.
Verified
Edited
Masterkey: Automated jailbreaking of large language model chatbots.
Deng, G.; Liu, Y.; Li, Y.; Wang, K.; Zhang, Y.; Li, Z.; Wang, H.; Zhang, T.; Liu, Y. (2024)
NDSS
Raw: 9. Deng, G., Liu, Y., Li, Y., Wang, K., Zhang, Y., Li, Z., Wang, H., Zhang, T., Liu, Y.: Masterkey: Automated jailbreaking of large language model chatbots. In: NDSS (2024)
Match: MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots
Authors: Gelei Deng; Yi Liu; Yuekang Li; Kailong Wang; Ying Zhang; Zefeng Li; Haoyu Wang; Tianwei Zhang; Yang Liu
Venue: Proceedings 2024 Network and Distributed System Security Symposium
DOI: 10.14722/ndss.2024.24188

ISBN:

URL: https://doi.org/10.14722/ndss.2024.24188
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily.
Ding, P.; Kuang, J.; Ma, D.; Cao, X.; Xian, Y.; Chen, J.; Huang, S. (2023)
arXiv
Raw: 10. Ding, P., Kuang, J., Ma, D., Cao, X., Xian, Y., Chen, J., Huang, S.: A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily. arXiv preprint arXiv:2311.08268 (2023)
Match: A Wolf in Sheep’s Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
Authors: Peng Ding; Jun Kuang; Dan Ma; Xuezhi Cao; Yunsen Xian; Jiajun Chen; Shujian Huang
Venue: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
DOI: 10.18653/v1/2024.naacl-long.118

ISBN:

URL: https://doi.org/10.18653/v1/2024.naacl-long.118
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
Do llms" know" internally when they follow instructions?
Heo, J.; Heinze-Deml, C.; Elachqar, O.; Ren, S.; Nallasamy, U.; Miller, A.; Chan, K.H.R.; Narain, J. (2024)
arXiv
Raw: 11. Heo, J., Heinze-Deml, C., Elachqar, O., Ren, S., Nallasamy, U., Miller, A., Chan, K.H.R., Narain, J.: Do llms" know" internally when they follow instructions? arXiv preprint arXiv:2410.14516 (2024)
Match: Do LLMs "know" internally when they follow instructions?
Authors: Juyeon Heo; Christina Heinze-Deml; Oussama Elachqar; Kwan Ho Ryan Chan; Shirley Ren; Udhay Nallasamy; Andy Miller; Jaya Narain
Venue: arXiv
DOI: 10.48550/arXiv.2410.14516

ISBN:

URL: https://arxiv.org/abs/2410.14516
Verified via arXiv id 2410.14516.
Verified
Edited
Efficient llm jailbreak via adaptive dense-to-sparse constrained optimization.
Hu, K.; Yu, W.; Li, Y.; Yao, T.; Li, X.; Liu, W.; Yu, L.; Shen, Z.; Chen, K.; Fredrikson, M. (2024)
Advances in Neural Information Processing Systems
Raw: 12. Hu, K., Yu, W., Li, Y., Yao, T., Li, X., Liu, W., Yu, L., Shen, Z., Chen, K., Fredrikson, M.: Efficient llm jailbreak via adaptive dense-to-sparse constrained optimization. Advances in Neural Information Processing Systems 37, 23224-23245 (2024)
Match: Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained Optimization
Authors: Kai Hu; Weichen Yu; Yining Li; Tianjun Yao; Xiang Li; Wenhe Liu; Lijun Yu; Zhiqiang Shen; Kai Chen; Matt Fredrikson
Venue: Advances in Neural Information Processing Systems 37
DOI: 10.52202/079017-0731

ISBN:

URL: https://doi.org/10.52202/079017-0731
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
Artprompt: Ascii art-based jailbreak attacks against aligned llms.
Jiang, F.; Xu, Z.; Niu, L.; Xiang, Z.; Ramasubramanian, B.; Li, B.; Poovendran, R. (2024)
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Raw: 13. Jiang, F., Xu, Z., Niu, L., Xiang, Z., Ramasubramanian, B., Li, B., Poovendran, R.: Artprompt: Ascii art-based jailbreak attacks against aligned llms. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 15157-15173 (2024)
Match: ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
Authors: Fengqing Jiang; Zhangchen Xu; Luyao Niu; Zhen Xiang; Bhaskar Ramasubramanian; Bo Li; Radha Poovendran
Venue: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
DOI: 10.18653/v1/2024.acl-long.809

ISBN:

URL: https://doi.org/10.18653/v1/2024.acl-long.809
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
Followbench: A multi-level fine-grained constraints following benchmark for large language models.
Jiang, Y.; Wang, Y.; Zeng, X.; Zhong, W.; Li, L.; Mi, F.; Shang, L.; Jiang, X.; Liu, Q.; Wang, W. (2023)
arXiv
Raw: 14. Jiang, Y., Wang, Y., Zeng, X., Zhong, W., Li, L., Mi, F., Shang, L., Jiang, X., Liu, Q., Wang, W.: Followbench: A multi-level fine-grained constraints following benchmark for large language models. arXiv preprint arXiv:2310.20410 (2023)
Match: FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
Authors: Yuxin Jiang; Yufei Wang; Xingshan Zeng; Wanjun Zhong; Liangyou Li; Fei Mi; Lifeng Shang; Xin Jiang; Qun Liu; Wei Wang
Venue: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
DOI: 10.18653/v1/2024.acl-long.257

ISBN:

URL: https://doi.org/10.18653/v1/2024.acl-long.257
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
Exploiting programmatic behavior of llms: Dual-use through standard security attacks.
Kang, D.; Li, X.; Stoica, I.; Guestrin, C.; Zaharia, M.; Hashimoto, T. (2024)
IEEE
Raw: 16. Kang, D., Li, X., Stoica, I., Guestrin, C., Zaharia, M., Hashimoto, T.: Exploiting programmatic behavior of llms: Dual-use through standard security attacks. In: 2024 IEEE Security and Privacy Workshops (SPW). pp. 132-143. IEEE (2024)
Match: Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks
Authors: Daniel Kang; Xuechen Li; Ion Stoica; Carlos Guestrin; Matei Zaharia; Tatsunori Hashimoto
Venue: 2024 IEEE Security and Privacy Workshops (SPW)
DOI: 10.1109/spw63631.2024.00018

ISBN:

URL: https://doi.org/10.1109/spw63631.2024.00018
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
False sense of security: A study on the effectivity of jailbreak detection in banking apps.
Kellner, A.; Horlboge, M.; Rieck, K.; Wressnegger, C. (2019)
IEEE
Raw: 17. Kellner, A., Horlboge, M., Rieck, K., Wressnegger, C.: False sense of security: A study on the effectivity of jailbreak detection in banking apps. In: 2019 IEEE European Symposium on Security and Privacy (EuroS&P). pp. 1-14. IEEE (2019)
Match: False Sense of Security: A Study on the Effectivity of Jailbreak Detection in Banking Apps
Authors: Ansgar Kellner; Micha Horlboge; Konrad Rieck; Christian Wressnegger
Venue: 2019 IEEE European Symposium on Security and Privacy (EuroS&P)
DOI: 10.1109/eurosp.2019.00011

ISBN:

URL: https://doi.org/10.1109/eurosp.2019.00011
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
Multi-step jailbreaking privacy attacks on chatgpt.
Li, H.; Guo, D.; Fan, W.; Xu, M.; Huang, J.; Meng, F.; Song, Y. (2023)
arXiv
Raw: 18. Li, H., Guo, D., Fan, W., Xu, M., Huang, J., Meng, F., Song, Y.: Multi-step jailbreaking privacy attacks on chatgpt. arXiv preprint arXiv:2304.05197 (2023)
Match: Multi-step Jailbreaking Privacy Attacks on ChatGPT
Authors: Haoran Li; Dadi Guo; Wei Fan; Mingshi Xu; Jie Huang; Fanpu Meng; Yangqiu Song
Venue: Findings of the Association for Computational Linguistics: EMNLP 2023
DOI: 10.18653/v1/2023.findings-emnlp.272

ISBN:

URL: https://doi.org/10.18653/v1/2023.findings-emnlp.272
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
Deepseek-v3 technical report.
Liu, A.; Feng, B.; Xue, B.; Wang, B.; Wu, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.; et al. (2024)
arXiv
Raw: 19. Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., et al.: Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437 (2024)
Match: DeepSeek-V3 Technical Report
Authors: DeepSeek-AI; Aixin Liu; Bei Feng; Bing Xue; Bingxuan Wang; Bochao Wu; Chengda Lu; Chenggang Zhao; Chengqi Deng; Chenyu Zhang; Chong Ruan; Damai Dai; Daya Guo; Dejian Yang; Deli Chen; Dongjie Ji; Erhang Li; Fangyun Lin; Fucong Dai; Fuli Luo; Guangbo Hao; Guanting Chen; Guowei Li; H. Zhang; Han Bao; Hanwei Xu; Haocheng Wang; Haowei Zhang; Honghui Ding; Huajian Xin; Huazuo Gao; Hui Li; Hui Qu; J. L. Cai; Jian Liang; Jianzhong Guo; Jiaqi Ni; Jiashi Li; Jiawei Wang; Jin Chen; Jingchang Chen; Jingyang Yuan; Junjie Qiu; Junlong Li; Junxiao Song; Kai Dong; Kai Hu; Kaige Gao; Kang Guan; Kexin Huang; Kuai Yu; Lean Wang; Lecong Zhang; Lei Xu; Leyi Xia; Liang Zhao; Litong Wang; Liyue Zhang; Meng Li; Miaojun Wang; Mingchuan Zhang; Minghua Zhang; Minghui Tang; Mingming Li; Ning Tian; Panpan Huang; Peiyi Wang; Peng Zhang; Qiancheng Wang; Qihao Zhu; Qinyu Chen; Qiushi Du; R. J. Chen; R. L. Jin; Ruiqi Ge; Ruisong Zhang; Ruizhe Pan; Runji Wang; Runxin Xu; Ruoyu Zhang; Ruyi Chen; S. S. Li; Shanghao Lu; Shangyan Zhou; Shanhuang Chen; Shaoqing Wu; Shengfeng Ye; Shengfeng Ye; Shirong Ma; Shiyu Wang; Shuang Zhou; Shuiping Yu; Shunfeng Zhou; Shuting Pan; T. Wang; Tao Yun; Tian Pei; Tianyu Sun; W. L. Xiao; Wangding Zeng; Wanjia Zhao; Wei An; Wen Liu; Wenfeng Liang; Wenjun Gao; Wenqin Yu; Wentao Zhang; X. Q. Li; Xiangyue Jin; Xianzu Wang; Xiao Bi; Xiaodong Liu; Xiaohan Wang; Xiaojin Shen; Xiaokang Chen; Xiaokang Zhang; Xiaosha Chen; Xiaotao Nie; Xiaowen Sun; Xiaoxiang Wang; Xin Cheng; Xin Liu; Xin Xie; Xingchao Liu; Xingkai Yu; Xinnan Song; Xinxia Shan; Xinyi Zhou; Xinyu Yang; Xinyuan Li; Xuecheng Su; Xuheng Lin; Y. K. Li; Y. Q. Wang; Y. X. Wei; Y. X. Zhu; Yang Zhang; Yanhong Xu; Yanhong Xu; Yanping Huang; Yao Li; Yao Zhao; Yaofeng Sun; Yaohui Li; Yaohui Wang; Yi Yu; Yi Zheng; Yichao Zhang; Yifan Shi; Yiliang Xiong; Ying He; Ying Tang; Yishi Piao; Yisong Wang; Yixuan Tan; Yiyang Ma; Yiyuan Liu; Yongqiang Guo; Yu Wu; Yuan Ou; Yuchen Zhu; Yuduan Wang; Yue Gong; Yuheng Zou; Yujia He; Yukun Zha; Yunfan Xiong; Yunxian Ma; Yuting Yan; Yuxiang Luo; Yuxiang You; Yuxuan Liu; Yuyang Zhou; Z. F. Wu; Z. Z. Ren; Zehui Ren; Zhangli Sha; Zhe Fu; Zhean Xu; Zhen Huang; Zhen Zhang; Zhenda Xie; Zhengyan Zhang; Zhewen Hao; Zhibin Gou; Zhicheng Ma; Zhigang Yan; Zhihong Shao; Zhipeng Xu; Zhiyu Wu; Zhongyu Zhang; Zhuoshu Li; Zihui Gu; Zijia Zhu; Zijun Liu; Zilin Li; Ziwei Xie; Ziyang Song; Ziyi Gao; Zizheng Pan
Venue: arXiv
DOI: 10.48550/arXiv.2412.19437

ISBN:

URL: https://arxiv.org/abs/2412.19437
Verified via arXiv id 2412.19437.
Verified
Edited
Chatcounselor: A large language models for mental health support.
Liu, J.M.; Li, D.; Cao, H.; Ren, T.; Liao, Z.; Wu, J. (2023)
arXiv
Raw: 20. Liu, J.M., Li, D., Cao, H., Ren, T., Liao, Z., Wu, J.: Chatcounselor: A large language models for mental health support. arXiv preprint arXiv:2309.15461 (2023)
Match: ChatCounselor: A Large Language Models for Mental Health Support
Authors: June M. Liu; Donghao Li; He Cao; Tianhe Ren; Zeyi Liao; Jiamin Wu
Venue: arXiv
DOI: 10.48550/arXiv.2309.15461

ISBN:

URL: https://arxiv.org/abs/2309.15461
Verified via arXiv id 2309.15461.
Verified
Edited
Autodan: Generating stealthy jailbreak prompts on aligned large language models.
Liu, X.; Xu, N.; Chen, M.; Xiao, C. (2023)
arXiv
Raw: 21. Liu, X., Xu, N., Chen, M., Xiao, C.: Autodan: Generating stealthy jailbreak prompts on aligned large language models. arXiv preprint arXiv:2310.04451 (2023)
Match: AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Authors: Xiaogeng Liu; Nan Xu; Muhao Chen; Chaowei Xiao
Venue: arXiv
DOI: 10.48550/arXiv.2310.04451

ISBN:

URL: https://arxiv.org/abs/2310.04451
Verified via arXiv id 2310.04451.
Verified
Edited
Flipattack: Jailbreak llms via flipping.
Liu, Y.; He, X.; Xiong, M.; Fu, J.; Deng, S.; Hooi, B. (2024)
arXiv
Raw: 22. Liu, Y., He, X., Xiong, M., Fu, J., Deng, S., Hooi, B.: Flipattack: Jailbreak llms via flipping. arXiv preprint arXiv:2410.02832 (2024)
Match: FlipAttack: Jailbreak LLMs via Flipping
Authors: Yue Liu; Xiaoxin He; Miao Xiong; Jinlan Fu; Shumin Deng; Yingwei Ma; Jiaheng Zhang; Bryan Hooi
Venue: arXiv
DOI: 10.48550/arXiv.2410.02832

ISBN:

URL: https://arxiv.org/abs/2410.02832
Verified via arXiv id 2410.02832.
Verified
Edited
Autoregressive text generation beyond feedback loops
Schmidt, F.; Mandt, S.; Hofmann, T. (2019)
arXiv
Raw: 23. Schmidt, F., Mandt, S., Hofmann, T.: Autoregressive text generation beyond feedback loops. arXiv preprint arXiv:1908.11658 (2019)
Match: Autoregressive Text Generation Beyond Feedback Loops
Authors: Florian Schmidt; Stephan Mandt; Thomas Hofmann
Venue: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)
DOI: 10.18653/v1/d19-1338

ISBN:

URL: https://doi.org/10.18653/v1/d19-1338
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
Scalable and transferable black-box jailbreaks for language models via persona modulation
Shah, R.; Pour, S.; Tagade, A.; Casper, S.; Rando, J.; et al. (2023)
arXiv
Raw: 24. Shah, R., Pour, S., Tagade, A., Casper, S., Rando, J., et al.: Scalable and transferable black-box jailbreaks for language models via persona modulation. arXiv preprint arXiv:2311.03348 (2023)
Match: Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
Authors: Rusheb Shah; Quentin Feuillade--Montixi; Soroush Pour; Arush Tagade; Stephen Casper; Javier Rando
Venue: arXiv
DOI: 10.48550/arXiv.2311.03348

ISBN:

URL: https://arxiv.org/abs/2311.03348
Verified via arXiv id 2311.03348.
Verified
Edited
Multilingual instruction tuning with just a pinch of multilinguality
Shaham, U.; Herzig, J.; Aharoni, R.; Szpektor, I.; Tsarfaty, R.; Eyal, M. (2024)
arXiv
Raw: 25. Shaham, U., Herzig, J., Aharoni, R., Szpektor, I., Tsarfaty, R., Eyal, M.: Multilingual instruction tuning with just a pinch of multilinguality. arXiv preprint arXiv:2401.01854 (2024)
Match: Multilingual Instruction Tuning With Just a Pinch of Multilinguality
Authors: Uri Shaham; Jonathan Herzig; Roee Aharoni; Idan Szpektor; Reut Tsarfaty; Matan Eyal
Venue: Findings of the Association for Computational Linguistics ACL 2024
DOI: 10.18653/v1/2024.findings-acl.136

ISBN:

URL: https://doi.org/10.18653/v1/2024.findings-acl.136
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
" do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models.
Shen, X.; Chen, Z.; Backes, M.; Shen, Y.; Zhang, Y. (2024)
Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security
Raw: 26. Shen, X., Chen, Z., Backes, M., Shen, Y., Zhang, Y.: " do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. In: Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. pp. 1671-1685 (2024)
Match: "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Authors: Xinyue Shen; Zeyuan Chen; Michael Backes; Yun Shen; Yang Zhang
Venue: Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security
DOI: 10.1145/3658644.3670388

ISBN:

URL: https://doi.org/10.1145/3658644.3670388
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
None
user4262 (2022)
Raw: 27. user4262: Dan is my new friend (2022), https://old.reddit.com/r/ChatGPT/comments/zlcyr9/dan_is_my_new_friend/
Match: DAN is my new friend
Authors:
Venue: r/ChatGPT
DOI:

ISBN:

URL:
Verified as a non-traditional/web source based on the provided identification of the Reddit thread.
Verified
Edited
Fundamental limitations of alignment in large language models
Wolf, Y.; Wies, N.; Avnery, O.; Levine, Y.; Shashua, A. (2023)
arXiv
Raw: 28. Wolf, Y., Wies, N., Avnery, O., Levine, Y., Shashua, A.: Fundamental limitations of alignment in large language models. arXiv preprint arXiv:2304.11082 (2023)
Match: Fundamental Limitations of Alignment in Large Language Models
Authors: Yotam Wolf; Noam Wies; Oshri Avnery; Yoav Levine; Amnon Shashua
Venue: arXiv
DOI: 10.48550/arXiv.2304.11082

ISBN:

URL: https://arxiv.org/abs/2304.11082
Verified via arXiv id 2304.11082.
Verified
Edited
A systematic study of cross-layer kv sharing for efficient llm inference
Wu, Y.; Wu, H.; Tu, K. (2024)
arXiv
Raw: 29. Wu, Y., Wu, H., Tu, K.: A systematic study of cross-layer kv sharing for efficient llm inference. arXiv preprint arXiv:2410.14442 (2024)
Match: A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference
Authors: You Wu; Haoyi Wu; Kewei Tu
Venue: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers)
DOI: 10.18653/v1/2025.naacl-short.34

ISBN:

URL: https://doi.org/10.18653/v1/2025.naacl-short.34
Verified via static CrossRef title search (score: 1.00)
Verified
Edited
Don’t listen to me: Understanding and exploring jailbreak prompts of large language models
Yu, Z.; Liu, X.; Liang, S.; Cameron, Z.; Xiao, C.; Zhang, N. (2024)
USENIX Association
Raw: 30. Yu, Z., Liu, X., Liang, S., Cameron, Z., Xiao, C., Zhang, N.: Don’t listen to me: Understanding and exploring jailbreak prompts of large language models. In: 33rd USENIX Security Symposium (USENIX Security 24). pp. 4675-4692. USENIX Association, Philadelphia, PA (Aug 2024)
Match: Don’t listen to me: Understanding and exploring jailbreak prompts of large language models
Authors: Zhiyuan Yu; Xiaogeng Liu; Shunning Liang; Zach Cameron; Chaowei Xiao; Ning Zhang
Venue: 33rd USENIX Security Symposium
DOI:

ISBN:

URL:
Verified based on the provided report confirming the accuracy of all bibliographic details.
Verified
Edited
Prepacking: A simple method for fast prefilling and increased throughput in large language models
Zhao, S.; Israel, D.; Broeck, G.V.d.; Grover, A. (2024)
arXiv
Raw: 31. Zhao, S., Israel, D., Broeck, G.V.d., Grover, A.: Prepacking: A simple method for fast prefilling and increased throughput in large language models. arXiv preprint arXiv:2404.09529 (2024)
Match: Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models
Authors: Siyan Zhao; Daniel Israel; Guy Van den Broeck; Aditya Grover
Venue: arXiv
DOI: 10.48550/arXiv.2404.09529

ISBN:

URL: https://arxiv.org/abs/2404.09529
Verified via arXiv id 2404.09529.
Verified
Edited
Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x
Zheng, Q.; Xia, X.; Zou, X.; Dong, Y.; Wang, S.; Xue, Y.; Shen, L.; Wang, Z.; Wang, A.; Li, Y.; et al. (2023)
ACM
Raw: 32. Zheng, Q., Xia, X., Zou, X., Dong, Y., Wang, S., Xue, Y., Shen, L., Wang, Z., Wang, A., Li, Y., et al.: Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x. In: Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. pp. 5673-5684 (2023)
Match: CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
Authors: Qinkai Zheng; Xiao Xia; Xu Zou; Yuxiao Dong; Shan Wang; Yufei Xue; Lei Shen; Zihan Wang; Andi Wang; Yang Li; Teng Su; Zhilin Yang; Jie Tang
Venue: Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
DOI: 10.1145/3580305.3599790

ISBN:

URL: https://doi.org/10.1145/3580305.3599790
Verified via static CrossRef title search (score: 0.99)
Verified
Edited
Universal and transferable adversarial attacks on aligned language models
Zou, A.; Wang, Z.; Carlini, N.; Nasr, M.; Kolter, J.Z.; Fredrikson, M. (2023)
arXiv
Raw: 33. Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J.Z., Fredrikson, M.: Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043 (2023)
Match: Universal and Transferable Adversarial Attacks on Aligned Language Models
Authors: Andy Zou; Zifan Wang; Nicholas Carlini; Milad Nasr; J. Zico Kolter; Matt Fredrikson
Venue: arXiv
DOI: 10.48550/arXiv.2307.15043

ISBN:

URL: https://arxiv.org/abs/2307.15043
Verified via arXiv id 2307.15043.