References
[1]
S.
Russell and P. Norvig, Artificial intelligence: A modern
approach, 4th ed. Pearson, 2021.
[2]
M.
Wooldridge and N. R. Jennings, “Intelligent agents: Theory and
practice,” The Knowledge Engineering Review, vol. 10,
no. 2, pp. 115–152, 1995.
[3]
A.
S. Rao and M. P. Georgeff, “BDI agents: From theory to
practice,” in Proceedings of the first international
conference on multi-agent systems (ICMAS), 1995.
[4]
R.
S. Sutton and A. G. Barto, Reinforcement learning: An
introduction, 2nd ed. MIT Press, 2018.
[5]
R.
A. Brooks, “Intelligence without representation,”
Artificial Intelligence, vol. 47, no. 1–3, pp. 139–159,
1991.
[6]
A.
Newell and H. A. Simon, “Computer science as empirical inquiry:
Symbols and search,” Communications of the ACM, vol. 19,
no. 3, pp. 113–126, 1976.
[7]
J.
McCarthy, M. L. Minsky, N. Rochester, and C. E. Shannon, “A
proposal for the Dartmouth summer research project on
artificial intelligence,” Dartmouth College, 1955.
[8]
R.
E. Fikes and N. J. Nilsson, “STRIPS: A new approach
to the application of theorem proving to problem solving,”
Artificial Intelligence, vol. 2, no. 3–4, pp. 189–208,
1971.
[9]
R.
A. Brooks, “A robust layered control system for a mobile
robot,” IEEE Journal of Robotics and Automation, vol.
RA–2, no. 1, pp. 14–23, 1986.
[10]
C.
J. C. H. Watkins, “Learning from delayed rewards,” PhD
thesis, King’s College, University of Cambridge, 1989.
[11]
J.
E. Laird, The Soar cognitive architecture. MIT
Press, 2012.
[12]
J.
R. Anderson, How can the human mind occur in the physical
universe? Oxford University Press, 2007.
[13]
J.
Wei et al., “Chain-of-thought prompting elicits reasoning
in large language models,” in Advances in neural information
processing systems (NeurIPS), 2022.
[14]
S.
Yao et al., “ReAct: Synergizing reasoning
and acting in language models,” in International conference
on learning representations (ICLR), 2023.
[15]
S.
Yao et al., “Tree of thoughts: Deliberate problem solving
with large language models,” in Advances in neural
information processing systems (NeurIPS), 2023.
[16]
N.
Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao,
“Reflexion: Language agents with verbal reinforcement
learning,” in Advances in neural information processing
systems (NeurIPS), 2023.
[17]
T.
Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large
language models are zero-shot reasoners,” in Advances in
neural information processing systems (NeurIPS), 2022.
[18]
X.
Wang et al., “Self-consistency improves chain of thought
reasoning in language models,” in International conference on
learning representations (ICLR), 2023.
[19]
L.
Wang et al., “Plan-and-solve prompting: Improving
zero-shot chain-of-thought reasoning by large language models,”
in Proceedings of the 61st annual meeting of the association for
computational linguistics (ACL), 2023.
[20]
S.
Kim et al., “An LLM compiler for parallel
function calling,” in International conference on machine
learning (ICML), 2024.
[21]
H.
Lightman et al., “Let’s verify step by step,”
arXiv preprint arXiv:2305.20050, 2023.
[22]
M.
Besta et al., “Demystifying chains, trees, and graphs of
thoughts,” arXiv preprint arXiv:2401.14295, 2024.
[23]
I.
Arcuschin, J. Janiak, R. Krzyzanowski, S. Rajamanoharan, N. Nanda, and
A. Conmy, “Chain-of-thought reasoning in the wild is not always
faithful,” in International conference on machine learning
(ICML), 2025.
[24]
Y.
Chen et al., “Reasoning models don’t always say what they
think,” arXiv preprint arXiv:2505.05410, 2025.
[25]
F.
Haji, M. Bethany, M. Tabar, J. Chiang, A. Rios, and P. Najafirad,
“Improving LLM reasoning with multi-agent tree-of-thought
validator agent,” arXiv preprint arXiv:2409.11527,
2024.
[26]
T.
B. Brown et al., “Language models are few-shot
learners,” in Advances in neural information processing
systems (NeurIPS), 2020.
[27]
L.
Ouyang et al., “Training language models to follow
instructions with human feedback,” in Advances in neural
information processing systems (NeurIPS), 2022.
[28]
J.
Wei et al., “Emergent abilities of large language
models,” Transactions on Machine Learning Research
(TMLR), 2022.
[29]
Z.
Ji et al., “Survey of hallucination in natural language
generation,” ACM Computing Surveys, vol. 55, no. 12, pp.
1–38, 2023.
[30]
N.
F. Liu et al., “Lost in the middle: How language models
use long contexts,” Transactions of the Association for
Computational Linguistics (TACL), vol. 12, 2024.
[31]
L.
Berglund et al., “The reversal curse: Language models
trained on ‘a is b’ fail to learn ‘b is
a’,” arXiv preprint arXiv:2309.12288, 2024.
[32]
I.
Mirzadeh, K. Alizadeh, H. Shahrokhi, O. Tuzel, S. Bengio, and M.
Farajtabar, “GSM-Symbolic: Understanding the
limitations of mathematical reasoning in large language models,”
arXiv preprint arXiv:2410.05229, 2024.
[33]
T.
Schick et al., “Toolformer: Language models can teach
themselves to use tools,” in Advances in neural information
processing systems (NeurIPS), 2023.
[34]
M.
Li et al., “API-Bank: A comprehensive
benchmark for tool-augmented LLMs,” in
Proceedings of the 2023 conference on empirical methods in natural
language processing (EMNLP), 2023.
[35]
S.
G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez, “Gorilla: Large
language model connected with massive APIs,”
arXiv preprint arXiv:2305.15334, 2023.
[36]
Y.
Qin et al., “ToolLLM: Facilitating large
language models to master 16000+ real-world APIs,”
arXiv preprint arXiv:2307.16789, 2023.
[37]
F.
Yan et al., “Berkeley function-calling leaderboard
(BFCL).” https://gorilla.cs.berkeley.edu/blogs/8_berkeley_function_calling_leaderboard.html,
2024.
[38]
Y.
Ruan et al., “Identifying the risks of LM
agents with an LM-emulated sandbox,”
International Conference on Learning Representations (ICLR),
2024.
[39]
P.
Lewis et al., “Retrieval-augmented generation for
knowledge-intensive NLP tasks,” in Advances in neural
information processing systems (NeurIPS), 2020.
[40]
G.
Mialon et al., “Augmented language models: A
survey,” arXiv preprint arXiv:2302.07842, 2023,
Available: https://arxiv.org/abs/2302.07842
[41]
A.
Asai, Z. Wu, Y. Wang, A. Sil, and H. Hajishirzi,
“Self-RAG: Learning to retrieve, generate, and
critique through self-reflection,” in International
conference on learning representations (ICLR), 2024. Available: https://arxiv.org/abs/2310.11511
[42]
S.
Borgeaud et al., “Improving language models by retrieving
from trillions of tokens,” in International conference on
machine learning (ICML), 2022. Available: https://arxiv.org/abs/2112.04426
[43]
J.
S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S.
Bernstein, “Generative agents: Interactive simulacra of human
behavior,” in ACM symposium on user interface software and
technology (UIST), 2023.
[44]
G.
Wang et al., “Voyager: An open-ended embodied agent with
large language models,” arXiv preprint arXiv:2305.16291,
2023.
[45]
Z.
Xi et al., “The rise and potential of large language
model based agents: A survey,” arXiv preprint
arXiv:2309.07864, 2023.
[46]
T.
Kwa, B. West, J. Becker, et al., “Measuring AI ability to
complete long tasks,” arXiv preprint arXiv:2503.14499,
Mar. 2025, Available: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
[47]
Y.
Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch,
“Improving factuality and reasoning in language models through
multiagent debate,” arXiv preprint arXiv:2305.14325,
2023.
[48]
M.
Cemri et al., “Why do multi-agent LLM systems
fail?” arXiv preprint arXiv:2503.13657, 2025.
[49]
K.-T. Tran, D. Dao, M.-D. Nguyen, Q.-V. Pham,
B. O’Sullivan, and H. D. Nguyen, “Multi-agent collaboration
mechanisms: A survey of LLMs,” arXiv preprint
arXiv:2501.06322, 2025.
[50]
Q.
Wu et al., “AutoGen: Enabling next-gen
LLM applications via multi-agent conversation,”
arXiv preprint arXiv:2308.08155, 2023.
[51]
S.
Hong et al., “MetaGPT: Meta programming for
a multi-agent collaborative framework,” arXiv preprint
arXiv:2308.00352, 2023.
[52]
G.
Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem,
“CAMEL: Communicative agents for ‘mind’
exploration of large language model society,” in Advances in
neural information processing systems (NeurIPS), 2023.
[53]
A.
Smit, P. Duckworth, N. Grinsztajn, T. D. Barrett, and A. Pretorius,
“Should we be going MAD? A look at multi-agent debate
strategies for LLMs,” arXiv preprint
arXiv:2311.17371, 2024.
[54]
X.
Liu, H. Yu, H. Zhang, et al., “AgentBench: Evaluating
LLMs as agents,” in International conference on learning
representations (ICLR), 2024.
[55]
C.
E. Jimenez et al., “SWE-bench: Can language models resolve real-world
GitHub issues?” International Conference on Learning
Representations (ICLR), 2024.
[56]
T.
Xie et al., “OSWorld: Benchmarking
multimodal agents for open-ended tasks in real computer
environments,” in Advances in neural information processing
systems (NeurIPS), datasets and benchmarks track, 2024.
[57]
S.
Zhou et al., “WebArena: A realistic web environment for
building autonomous agents,” arXiv preprint
arXiv:2307.13854, 2023.
[58]
X.
Deng et al., “Mind2Web: Towards a generalist agent for
the web,” in Advances in neural information processing
systems (NeurIPS), 2023.
[59]
B.
Zheng, B. Gou, J. Kil, H. Sun, and Y. Su, “GPT-4V(ision) is a generalist web agent, if
grounded,” in International conference on machine learning
(ICML), 2024.
[60]
A.
Brohan et al., “RT-2: Vision-language-action
models transfer web knowledge to robotic control,” in
Conference on robot learning (CoRL), 2023.
[61]
F.
F. Xu et al., “TheAgentCompany: Benchmarking
LLM agents on consequential real world tasks.” 2024.
Available: https://arxiv.org/abs/2412.14161
[62]
X.
Deng et al., “SWE-Bench Pro: Can
AI agents solve long-horizon software engineering
tasks?” 2025. Available: https://arxiv.org/abs/2509.16941
[63]
Anthropic, “Introducing computer use, a
new claude 3.5 sonnet, and claude 3.5 haiku.” https://www.anthropic.com/news/3-5-models-and-computer-use,
Oct. 2024.
[64]
OpenAI, “Introducing operator.” https://openai.com/index/introducing-operator/, Jan.
2025.
[65]
S.
Yao, N. Shinn, P. Razavi, and K. Narasimhan, “τ-bench:
A benchmark for tool-agent-user interaction in real-world
domains,” arXiv preprint arXiv:2406.12045, 2024,
Available: https://arxiv.org/abs/2406.12045
[66]
L.
Zheng et al., “Judging LLM-as-a-judge with
MT-Bench and chatbot arena,” in Advances in
neural information processing systems (NeurIPS), datasets and benchmarks
track, 2023. Available: https://arxiv.org/abs/2306.05685
[67]
S.
Es, J. James, L. Espinosa-Anke, and S. Schockaert,
“RAGAS: Automated evaluation of retrieval augmented
generation,” arXiv preprint arXiv:2309.15217, 2023,
Available: https://arxiv.org/abs/2309.15217
[68]
OpenAI, “Introducing SWE-bench verified.” https://openai.com/index/introducing-swe-bench-verified/,
Aug. 2024.
[69]
Anthropic, M. Grace, J. Hadfield, R. Olivares,
and J. De Jonghe, “Demystifying evals for AI agents.” https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents,
Jan. 2026.
[70]
Anthropic, E. Schluntz, and B. Zhang,
“Building effective agents.” https://www.anthropic.com/engineering/building-effective-agents,
Dec. 2024.
[71]
Anthropic, “Prompt caching with
claude.” https://www.anthropic.com/news/prompt-caching, Aug.
2024.
[72]
OpenAI, “Batch API.” https://platform.openai.com/docs/guides/batch,
2024.
[73]
Stripe, “Idempotent requests.” https://docs.stripe.com/api/idempotent_requests,
2024.
[74]
L.
Chen, M. Zaharia, and J. Zou, “FrugalGPT: How to use large
language models while reducing cost and improving performance,”
arXiv preprint arXiv:2305.05176, 2023.
[75]
I.
Ong et al., “RouteLLM: Learning to route LLMs with
preference data,” arXiv preprint arXiv:2406.18665,
2024.
[76]
Q.
Zhang, M. Wornow, G. Wan, and K. Olukotun, “Agentic plan caching:
Test-time memory for fast and cost-efficient LLM agents,”
arXiv preprint arXiv:2506.14852, 2025.
[77]
H.
Xia et al., “Unlocking efficiency in large language model
inference: A comprehensive survey of speculative decoding,” in
Findings of the association for computational linguistics: ACL
2024, Bangkok, Thailand: Association for Computational Linguistics,
2024, pp. 7655–7671.
[78]
Anthropic et al., “How we built
our multi-agent research system.” https://www.anthropic.com/engineering/multi-agent-research-system,
Jun. 2025.
[79]
W.
Yan and Cognition, “Don’t build multi-agents.” https://cognition.ai/blog/dont-build-multi-agents, Jun.
2025.
[80]
Microsoft, “Agentic application
patterns.” https://learn.microsoft.com/en-us/azure/durable-task/sdks/durable-agents-patterns,
2026.
[81]
OpenAI, “New tools for building
agents.” https://openai.com/index/new-tools-for-building-agents/,
Mar. 2025.
[82]
Anthropic, “Model context
protocol.” https://modelcontextprotocol.io, 2024.
[83]
Model Context Protocol, “Model context
protocol specification (2025-06-18).” https://modelcontextprotocol.io/specification/2025-06-18,
2025.
[84]
X.
Hou, Y. Zhao, S. Wang, and H. Wang, “Model context protocol
(MCP): Landscape, security threats, and future research
directions,” arXiv preprint arXiv:2503.23278,
2025.
[85]
Z.
Wang et al., “MCPTox: A benchmark for tool
poisoning attack on real-world MCP servers,”
arXiv preprint arXiv:2508.14925, 2025.
[86]
A.
RoyChowdhury, M. Luo, P. Sahu, S. Banerjee, and M. Tiwari,
“ConfusedPilot: Confused deputy risks in
RAG-based LLMs,” arXiv preprint
arXiv:2408.04870, 2024.
[87]
LangChain, “LangGraph
documentation.” https://langchain-ai.github.io/langgraph/, 2024.
[88]
LangChain, “LangGraph
persistence.” https://docs.langchain.com/oss/python/langgraph/persistence,
2025.
[89]
LangChain, “Human-in-the-loop.” https://docs.langchain.com/oss/python/langchain/human-in-the-loop,
2025.
[90]
OpenTelemetry Authors, “Semantic
conventions for generative AI systems.” https://opentelemetry.io/blog/2026/genai-observability/,
2026.
[91]
OWASP, “OWASP top 10 for large language
model applications.” https://owasp.org/www-project-top-10-for-large-language-model-applications/,
2025.
[92]
K.
Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz,
“Not what you’ve signed up for: Compromising real-world
LLM-integrated applications with indirect prompt
injection,” in Proceedings of the 16th ACM workshop on
artificial intelligence and security (AISec), 2023. Available: https://arxiv.org/abs/2302.12173
[93]
S.
Willison, “Prompt injection attacks against
GPT-3.” https://simonwillison.net/2022/Sep/12/prompt-injection/,
Sep. 2022.
[94]
S.
Willison, “The lethal trifecta for AI agents: Private
data, untrusted content, and external communication.” https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/,
Jun. 2025.
[95]
B.
H. Sigelman et al., “Dapper, a large-scale distributed
systems tracing infrastructure,” Google, Inc., 2010. Available:
https://research.google/pubs/dapper-a-large-scale-distributed-systems-tracing-infrastructure/
[96]
S.
Kanzhelev, M. McLean, A. Reitbauer, B. Drutu, N. Molnar, and Y. Shkuro,
“Trace context.” https://www.w3.org/TR/trace-context/, Nov. 2021.
[97]
Arize AI, “OpenInference
semantic conventions.” https://github.com/Arize-ai/openinference/blob/main/spec/semantic_conventions.md,
2024.
[98]
C.
Packer et al., “MemGPT: Towards
LLMs as operating systems,” arXiv preprint
arXiv:2310.08560, 2023, Available: https://arxiv.org/abs/2310.08560
[99]
W.
Zhong, L. Guo, Q. Gao, H. Ye, and Y. Wang,
“MemoryBank: Enhancing large language models with
long-term memory,” arXiv preprint arXiv:2305.10250,
2023, Available: https://arxiv.org/abs/2305.10250
[100]
P. Chhikara, D. Khant, S. Aryan, T. Singh, and
D. Yadav, “Mem0: Building production-ready
AI agents with scalable long-term memory,” arXiv
preprint arXiv:2504.19413, 2025, Available: https://arxiv.org/abs/2504.19413
[101]
D. Wu, H. Wang, W. Yu, Y. Zhang, K.-W. Chang,
and D. Yu, “LongMemEval: Benchmarking chat assistants
on long-term interactive memory,” in International conference
on learning representations (ICLR), 2025. Available: https://arxiv.org/abs/2410.10813
[102]
A. Maharana, D.-H. Lee, S. Tulyakov, M. Bansal,
F. Barbieri, and Y. Fang, “Evaluating very long-term
conversational memory of LLM agents,” arXiv
preprint arXiv:2402.17753, 2024, Available: https://arxiv.org/abs/2402.17753
[103]
H. Tan, Z. Zhang, C. Ma, X. Chen, Q. Dai, and
Z. Dong, “MemBench: Towards more comprehensive
evaluation on the memory of LLM-based agents,” in
Findings of the association for computational linguistics: ACL
2025, 2025, pp. 19336–19352. Available: https://aclanthology.org/2025.findings-acl.989/
[104]
W. Xu, Z. Liang, K. Mei, H. Gao, J. Tan, and Y.
Zhang, “A-MEM: Agentic memory for LLM
agents,” in Advances in neural information processing systems
(NeurIPS), 2025. Available: https://arxiv.org/abs/2502.12110
[105]
Y. Hu et al., “Memory in the age
of AI agents,” arXiv preprint
arXiv:2512.13564, 2025, Available: https://arxiv.org/abs/2512.13564
[106]
A. Vaswani et al., “Attention is
all you need,” in Advances in neural information processing
systems (NeurIPS), 2017. Available: https://arxiv.org/abs/1706.03762
[107]
R. Sennrich, B. Haddow, and A. Birch,
“Neural machine translation of rare words with subword
units,” in Proceedings of the 54th annual meeting of the
association for computational linguistics (ACL), 2016. Available:
https://arxiv.org/abs/1508.07909
[108]
A. Radford, J. Wu, R. Child, D. Luan, D.
Amodei, and I. Sutskever, “Language models are unsupervised
multitask learners,” OpenAI, 2019. Available: https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf
[109]
J. Kaplan et al., “Scaling laws
for neural language models,” arXiv preprint
arXiv:2001.08361, 2020, Available: https://arxiv.org/abs/2001.08361
[110]
J. Hoffmann et al., “Training
compute-optimal large language models,” in Advances in neural
information processing systems (NeurIPS), 2022. Available: https://arxiv.org/abs/2203.15556
[111]
P. F. Christiano, J. Leike, T. B. Brown, M.
Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from
human preferences,” in Advances in neural information
processing systems (NeurIPS), 2017. Available: https://arxiv.org/abs/1706.03741
[112]
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon,
C. D. Manning, and C. Finn, “Direct preference optimization: Your
language model is secretly a reward model,” in Advances in
neural information processing systems (NeurIPS), 2023. Available:
https://arxiv.org/abs/2305.18290
[113]
Z. Shao et al.,
“DeepSeekMath: Pushing the limits of mathematical
reasoning in open language models,” arXiv preprint
arXiv:2402.03300, 2024, Available: https://arxiv.org/abs/2402.03300
[114]
D. Guo et al.,
“DeepSeek-R1: Incentivizing reasoning capability in
LLMs via reinforcement learning,” Nature,
vol. 645, pp. 633–638, 2025, doi: 10.1038/s41586-025-09422-z.
[115]
W. Fedus, B. Zoph, and N. Shazeer,
“Switch transformers: Scaling to trillion parameter models with
simple and efficient sparsity,” Journal of Machine Learning
Research, vol. 23, no. 120, pp. 1–39, 2022, Available: https://arxiv.org/abs/2101.03961
[116]
W. Kwon et al., “Efficient
memory management for large language model serving with
PagedAttention,” in Proceedings of the 29th
symposium on operating systems principles (SOSP), 2023. Available:
https://arxiv.org/abs/2309.06180
[117]
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y.
Choi, “The curious case of neural text degeneration,” in
International conference on learning representations (ICLR),
2020. Available: https://arxiv.org/abs/1904.09751
[118]
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y.
Liu, “RoFormer: Enhanced transformer with rotary
position embedding,” Neurocomputing, vol. 568, p.
127063, 2024, doi: 10.1016/j.neucom.2023.127063.
[119]
R. Schaeffer, B. Miranda, and S. Koyejo,
“Are emergent abilities of large language models a mirage?”
in Advances in neural information processing systems (NeurIPS),
2023. Available: https://arxiv.org/abs/2304.15004
[120]
M. Besta et al., “Graph of
thoughts: Solving elaborate problems with large language models,”
in Proceedings of the AAAI conference on artificial intelligence
(AAAI), 2024, pp. 17682–17690. doi: 10.1609/aaai.v38i16.29720.
[121]
L. Weng, “LLM powered
autonomous agents.” https://lilianweng.github.io/posts/2023-06-23-agent/,
2023.
[122]
OpenAI, “A practical guide to building
agents.” https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf,
2025.
[123]
H. Chase, “What is a "cognitive
architecture"?” https://blog.langchain.com/what-is-a-cognitive-architecture/,
2024.
[124]
crewAI Inc., “CrewAI
flows.” https://docs.crewai.com/en/concepts/flows, 2025.
[125]
Letta, “Letta: The stateful
agents framework.” https://docs.letta.com/, 2024.
[126]
P. Rasmussen, P. Paliychuk, T. Beauvais, J.
Ryan, and D. Chalef, “Zep: A temporal knowledge graph
architecture for agent memory,” in arXiv preprint
arXiv:2501.13956, 2025. Available: https://arxiv.org/abs/2501.13956
[127]
LangChain, “LangGraph
streaming.” https://docs.langchain.com/oss/python/langgraph/streaming,
2025.
[128]
OpenAI, “Swarm: Educational framework
exploring ergonomic, lightweight multi-agent orchestration.” https://github.com/openai/swarm, 2024.
[129]
OpenAI, “OpenAI Agents SDK
documentation.” https://openai.github.io/openai-agents-python/,
2025.
[130]
OpenAI, “OpenAI Responses
API: New tools for building agents.” https://openai.com/index/new-tools-for-building-agents/,
2025.
[131]
L. D. Erman, F. Hayes-Roth, V. R. Lesser, and
D. R. Reddy, “The Hearsay-II speech-understanding
system: Integrating knowledge to resolve uncertainty,” ACM
Computing Surveys, vol. 12, no. 2, pp. 213–253, 1980, doi: 10.1145/356810.356816.
[132]
Google, “Agent2Agent (A2A)
Protocol.” https://google.github.io/A2A/, 2025.
[133]
IBM, “Agent Communication Protocol
(ACP).” https://agentcommunicationprotocol.dev/, 2025.
[134]
B. Efron and R. J. Tibshirani, An
introduction to the bootstrap. Chapman & Hall/CRC, 1993.
[135]
E. Debenedetti et al.,
“Defeating prompt injections by design,” arXiv preprint
arXiv:2503.18813, 2025, Available: https://arxiv.org/abs/2503.18813
[136]
OpenAI, “OpenAI
API pricing.” https://openai.com/api/pricing/, 2025.
[137]
Anthropic, “Claude
pricing.” https://www.anthropic.com/pricing, 2025.
[138]
E. Debenedetti, J. Zhang, M. Balunović, L.
Beurer-Kellner, M. Fischer, and F. Tramr, “AgentDojo:
A dynamic environment to evaluate prompt injection attacks and defenses
for LLM agents,” in Advances in neural
information processing systems (NeurIPS), 2024. Available: https://arxiv.org/abs/2406.13352
[139]
S. Zhou et al.,
“WebArena: A realistic web environment for building
autonomous agents,” in International conference on learning
representations (ICLR), 2024. Available: https://arxiv.org/abs/2307.13854
[140]
S. Willison, “The dual LLM
pattern for building AI assistants that can resist prompt
injection.” https://simonwillison.net/2023/Apr/25/dual-llm-pattern/,
2023.
[141]
G. Mialon, C. Fourrier, T. Wolf, Y. LeCun, and
T. Scialom, “GAIA: A benchmark for general
AI assistants,” in International conference on
learning representations (ICLR), 2024. Available: https://arxiv.org/abs/2311.12983
[142]
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J.
Gao, “Set-of-mark prompting unleashes extraordinary visual
grounding in GPT-4V,” in arXiv preprint
arXiv:2310.11441, 2023. Available: https://arxiv.org/abs/2310.11441
[143]
Google, “Agent2Agent
(A2A) protocol.” https://google.github.io/A2A/, 2024.
[144]
J. Kirkpatrick et al.,
“Overcoming catastrophic forgetting in neural networks,”
Proceedings of the National Academy of Sciences, vol. 114, no.
13, pp. 3521–3526, 2017, doi: 10.1073/pnas.1611835114.
[145]
Google DeepMind, “Project
Mariner: An early research prototype for agentic
browsing.” https://deepmind.google/technologies/project-mariner/,
2024.
[146]
Anthropic, “Anthropic’s responsible
scaling policy.” https://www.anthropic.com/rsp, 2024.
[147]
European Union, “Regulation
(EU) 2024/1689 of the european parliament and of the
council laying down harmonised rules on artificial intelligence
(AI act).” https://eur-lex.europa.eu/eli/reg/2024/1689/oj,
2024.
[148]
M. McCloskey and N. J. Cohen,
“Catastrophic interference in connectionist networks: The
sequential learning problem,” Psychology of Learning and
Motivation, vol. 24, pp. 109–165, 1989, doi: 10.1016/S0079-7421(08)60536-8.
[149]
National Institute of Standards and Technology,
“Artificial intelligence risk management framework (AI RMF
1.0),” U.S. Department of Commerce, NIST AI 100-1, 2023.
Available: https://doi.org/10.6028/NIST.AI.100-1