Comparison of AI agents for creating SQL queries
Article Sidebar
Issue Vol. 40 (2026)
-
Analysis of the capabilities of predictive artificial intelligence models in corporate risk management
Kacper Ziemski188-192
-
Usability and availability of selected e-commerce services
Marcin Kozicki, Maria Skublewska-Paszkowska193-200
-
Comparison of C++ and Python performance based on selected algorithms
Szymon Bogucki, Kacper Burda201-205
-
Security analysis of selected web applications using vulnerability scanners
Mariusz Choroś, Marta Dziuba-Kozieł206-212
-
Comparison of Java and .NET reflection mechanisms for dynamic module loading: a performance benchmark study
Michał Mazur, Sebastian Maruszak, Marek Miłosz213-217
-
Comparative analysis of network vulnerability detection tools
Mateusz Zdunek218-225
-
Evaluation of mobile applications for personal finance management using the MARS scale
Łukasz Nikiel, Artsiom Patskevich, Marek Miłosz226-231
-
Comparison of the effectiveness of roulette betting strategies using Monte Carlo simulation
Marek Sarnecki232-238
-
Analysis of optimization capabilities of selected database management systems
Paweł Tarkiewicz, Małgorzata Plechawska-Wójcik239-246
-
Comparative analysis of Espresso and Appium frameworks for automated UI testing of Android mobile applications
Jakub Derkacz247-254
-
Comparative analysis of the applicability of artificial intelligence models for code generation
Patryk Warchoł, Małgorzata Plechawska-Wójcik255-262
-
Comparison of the effectiveness of selected tools for detecting texts generated by artificial intelligence
Marcin Brodacki, Małgorzata Plechawska-Wójcik263-269
-
Comparative analysis of selected containerization tools in terms of MCP
Paweł Jan Tłusty, Maciej Pańczyk270-276
-
SpikeCliff effect: empirical analysis of deterministic timing discontinuities in sponge-based XOF functions
Łukasz Wójcik, Stanisław Lota277-282
-
Comparison of AI agents for creating SQL queries
Julia Sierpień, Maria Skublewska-Paszkowska283-288
-
Comparative analysis of the performance of PostgreSQL and Neo4j databases in the context of genealogical queries
Michał Muzyka, Mateusz Niedźwiedź, Marek Miłosz289-296
-
Evaluation of the effectiveness of static and dynamic methods in malware analysis
Dominik Tracz, Daniel Sawicki, Konrad Gromaszek297-303
-
Comparison of classical machine learning methods in the task of obesity level classification
Paweł Biesaga, Paweł Powroźnik304-312
Main Article Content
Authors
Abstract
Large language models have been increasingly applied for Text-to-SQL tasks recently, due to the fact that generating correct SQL queries is the most important factor. This study presents a comparison of AI agents based on open-source and closed-source large language models for generating SQL queries. GPT-4o, Claude 3.7 Sonnet, and LLaMA 3 8B were evaluated using two widely known datasets: Spider and WikiSQL. The chosen architectures were compared using Exact Match, F1-score, BERTScore, Execution Accuracy and processing time. Obtained results show that all agents perform well on simple queries, while agents based on closed-source models achieve better performance on complex SQL generation tasks.
Keywords:
Sustainable Development Goal (SDG)
- Industry, Innovation, Technology and Infrastructure
References
[1] A. Singh, A. Shetty, A. Ehtesham, S. Kumar, T. T. Khoei, A survey of large language model-based generative AI for text-to-SQL: Benchmarks, applications, use cases, and challenges, In 2025 IEEE 15th Annual Computing and Communication Workshop and Conference (CCWC) (2025) 15–21, https://doi.org/10.1109/CCWC62904.2025.10903689
[2] A. Mohammadjafari, A. S. Maida, R. Gottumukkala, From natural language to SQL: Review of LLM-based text-to-SQL systems, arXiv preprint arXiv:2410.01066 (2025), https://doi.org/10.48550/arXiv.2410.01066.
[3] G. Katsogiannis-Meimarakis, G. Koutrika, A survey on deep learning approaches for text-to-SQL, The VLDB Journal 32(4) (2023) 905–936, https://doi.org/10.1007/s00778-022-00776-8.
[4] A. B. Kanburoğlu, F. B. Tek, Text-to-SQL: A methodical review of challenges and models, Turkish Journal of Electrical Engineering and Computer Sciences 32(3) (2024) 403–419, https://doi.org/10.55730/1300-0632.4077.
[5] Z. Hong, Z. Yuan, Q. Zhang, H. Chen, J. Dong, F. Huang, X. Huang, Next-generation database interfaces: A survey of LLM-based text-to-SQL, arXiv preprint arXiv:2406.08426 (2024), https://doi.org/10.48550/arXiv.2406.08426.
[6] D. Gao, H. Wang, Y. Li, X. Sun, Y. Qian, B. Ding, J. Zhou, Text-to-SQL empowered by large language models: A benchmark evaluation, arXiv preprint arXiv:2308.15363 (2023), https://doi.org/10.48550/arXiv.2308.15363.
[7] S. Sun, Y. Zhang, J. Yan, Y. Gao, D. Ong, B. Chen, J. Su, Battle of the large language models: Dolly vs LLaMA vs Vicuna vs Guanaco vs Bard vs ChatGPT – A text-to-SQL parsing comparison, In Findings of the Association for Computational Linguistics: EMNLP 2023 (2023) 11225–11238, https://doi.org/10.18653/v1/2023.findings-emnlp.750
[8] R. Roberson, G. Kaki, A. Trivedi, Analyzing the effectiveness of large language models on text-to-SQL synthesis, arXiv preprint arXiv:2401.12379 (2024), https://doi.org/10.48550/arXiv.2401.12379.
[9] C.-M. Rosca, A. Stancu, Quality assessment of GPT-3.5 and Gemini 1.0 Pro for SQL syntax, Computer Standards & Interfaces 95 (2026) 104041, https://doi.org/10.1016/j.csi.2025.104041.
[10] L. Nan, Y. Zhao, W. Zou, N. Ri, J. Tae, E. Zhang, A. Cohan, D. Radev, Enhancing text-to-SQL capabilities of large language models: A study on prompt design strategies, In Findings of the Association for Computational Linguistics: EMNLP 2023 (2023) 14935–14956, https://doi.org/10.18653/v1/2023.findings-emnlp.996.
[11] K. Zhang, X. Lin, Y. Wang, X. Zhang, F. Sun, C. Jianhe, H. Tan, X. Jiang, H. Shen, ReFSQL: A retrieval-augmentation framework for text-to-SQL generation, In Findings of the Association for Computational Linguistics: EMNLP 2023 (2023) 664–673, https://doi.org/10.18653/v1/2023.findings-emnlp.48.
[12] Z. Shen, P. Vougiouklis, C. Diao, K. Vyas, Y. Ji, J. Z. Pan, Improving retrieval-augmented text-to-SQL with AST-based ranking and schema pruning, arXiv preprint arXiv:2407.03227 (2024), https://doi.org/10.48550/arXiv.2407.03227.
[13] Z. Tan, X. Liu, Q. Shu, X. Li, C. Wan, D. Liu, Q. Wan, G. Liao, Enhancing text-to-SQL capabilities of large language models through tailored promptings, Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (2024) 6091–6109, https://aclanthology.org/2024.lrec-main.539/.
[14] G. M. C. Coelho, E. R. S. Nascimento, Y. T. Izquierdo, G. M. García, L. Feijó, M. Lemos, R. L. S. Garcia, A. R. de Oliveira, J. P. Pinheiro, M. A. Casanova, Improving the accuracy of text-to-SQL tools based on large language models for real-world relational databases, Database and Expert Systems Applications: 35th International Conference, DEXA 2024 (2024) 93–107, https://doi.org/10.1007/978-3-031-68309-1_8.
[15] OpenAI, J. Achiam, S. Adler, S. Agarwal et al., GPT-4 technical report, arXiv preprint arXiv:2303.08774 (2024), https://doi.org/10.48550/arXiv.2303.08774.
[16] J. Yang, B. Hui, M. Yang, J. Yang, J. Lin, C. Zhou, Synthesizing text-to-SQL data from weak and strong LLMs, arXiv preprint arXiv:2408.03256 (2024), https://doi.org/10.48550/arXiv.2408.03256.
[17] Y. Mellah, V. Kocaman, H. U. Haq, D. Talby, Efficient schema-less text-to-SQL conversion using large language models, Artificial Intelligence in Health 1(2) (2024) 96–106, https://doi.org/10.36922/aih.2661.
[18] W. Wen, Y. Zhang, S. Pan, Y. Sun, P. Lu, C. Ding, LR-SQL: A supervised fine-tuning method for Text2SQL tasks under low-resource scenarios, Electronics 14(17) (2025) 3489, https://doi.org/10.3390/electronics14173489.
[19] I. Azurmendi, E. Zulueta, G. García, N. Uriarte-Arrazola, J. M. Lopez-Guede, Beyond standard losses: Redefining text-to-SQL with task-specific optimization, Mathematics 13(14) (2025) 2315, https://doi.org/10.3390/math13142315.
[20] Spider dataset, https://yale-lily.github.io/spider, [05.06.2026].
[21] WikiSQL dataset, https://github.com/salesforce/WikiSQL, [05.06.2026].
[22] E. Kasireddy, C. Chow, J. Collet, M.-M. Pourrahmat, M. S. Fazeli, Evaluating the performance of Claude 3.7 Sonnet in data extraction automation for systematic literature reviews, Value in Health Regional Issues 53 (2026) 101539, https://doi.org/10.1016/j.vhri.2025.101539.
[23] A. Grattafiori et al., The Llama 3 herd of models, arXiv preprint arXiv:2407.21783 (2024), https://doi.org/10.48550/arXiv.2407.21783.
[24] Exact Match, Execution Accuracy, https://deepwiki.com/taoyds/spider/3.1-evaluation-metrics, [05.06.2026].
[25] F1 Score, https://www.v7labs.com/blog/f1-score-guide, [05.06.2026].
[26] BERTScore, https://www.analyticsvidhya.com/blog/2025/04/bertscore-a-contextual-metric-for-llm-evaluation/, [05.06.2026].

