Design and implementation of a pedagogical layer for large language models for mathematics education
Issue
Section
Keywords:
Artificial intelligence, large language models, prompt engineering, mathematics education, personalized education
Published
Abstract
This paper presents the design and theoretical validation of an intelligent agent aimed at improving mathematics learning in secondary education in Colombia. Despite the new models and opportunities introduced with the use of digital technologies, student performance in mathematics remains critical, particularly in public institutions in Cartagena, as evidenced by Saber 11 and PISA results. Large Language Models (LLMs) offer scalable tutoring capabilities, but off-the-shelf generalist models often lack pedagogical context, risking “overhelp” by providing direct solutions rather than instructional guidance. We introduce a novel intelligent agent that integrates a “Pedagogical Value Layer” (PVL) on top of a general-purpose LLM that constrains its generative behavior through an explicit pedagogical control layer designed to guide the learning process. The proposed system utilizes a four-level prompt architecture (general principles, local curricular context, instructor-defined strategies, and student preferences) to transform the model into an adaptive tutor without the need for expensive and complex fine-tuning. We implemented a prototype platform and conducted a comparative benchmark against a baseline model to evaluate this proposal. Results demonstrate that the PVL successfully suppresses direct answer-giving behaviors, ensuring the agent adopts a questioning approach to encourage deliberate reasoning and conceptual development, while incorporating local and content information for better contextualized interaction.Supporting agencias
- The authors express their gratitude for the support provided for carrying out this research to the Generalitat de Catalunya (2021SGR-633) and Universitat Rovira i Virgili (2023PFR-URV-00633).
License
Copyright (c) 2026 Fabio E. García Ramírez, Jordi Duch

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Authors publishing in this journal agree to the following terms:
- Authors retain copyright and grant in the journal the right of first publication of the work, registered under a Creative Commons acknowledgement licence (CC BY-NC-SA), which allows dissemination with acknowledgement of authorship and first publication in this journal.
- Authors may independently enter into other contractual arrangements to allow publication of the version published in this journal in other media (e.g., in an institutional repository or in a book), with acknowledgement of initial publication in this journal.
- Authors have permission to publish their work online and are encouraged to do so (e.g., in institutional repositories or on their website) before and during the submission process, because it can produce good results and lead to the published work receiving more citations (see The Effect of Open Access).
PRIVACY STATEMENT
Names and email addresses on the journal website will only be used for the uses indicated in this journal and will not be made available for any other use or to third parties.
Downloads
References
Angel-Urdinola, D. F., Avitabile, C., & Chinen, M. H. (2023). Can digital personalized learning for mathematics remediation level the playing field in higher education? Experimental evidence from Ecuador (Policy Research Working Paper No. 10483). World Bank. https://doi.org/10.1596/1813-9450-10483
Avella, B. (2025, June). Socratic AI tutoring in primary school mathematics: A case study on the development of problem-solving and digital competence. Paper presented at the 2025 MIT AI and Education Summit, Cambridge, MA. https://dspace.mit.edu/handle/1721.1/163131
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Marber, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences (PNAS). https://doi.org/10.1073/pnas.2422633122
Borchers, C., & Shou, T. (2025). Can large language models match tutoring system adaptivity? A benchmarking study. Proceedings of AIED 2025. https://doi.org/10.1007/978-3-031-98417-4_29
Cartagena Cómo Vamos. (2023, December 13). Lo que está detrás de los nuevos resultados de las Pruebas Saber 11: 2023-4. https://cartagenacomovamos.org/lo-que-esta-detras-de-los-nuevos-resultados-de-las-pruebas-saber-11/
Cartagena Cómo Vamos. (2024, December 9). Cartagena logra avances en las pruebas Saber 11, pero persisten desafíos en el desempeño. https://cartagenacomovamos.org/cartagena-logra-avances-en-las-pruebas-saber-11-pero-persisten-desafios-en-el-desempeno/
Cartagena Cómo Vamos. (2025, May 23). Cartagena valora a sus docentes, pero enfrenta grandes retos en infraestructura y calidad educativa. https://cartagenacomovamos.org/cartagena-valora-docentes-pero-tiene-retos-educativos-cartagena-2025/
Chan, K. K., & Leung, S. W. (2014). Dynamic geometry software improves mathematical achievement: Systematic review and meta-analysis. Journal of Educational Computing Research, 51(3), 311–325. https://doi.org/10.2190/EC.51.3.c
Chi, M. T. H., Siler, S. A., Jeong, H., Yamauchi, T., & Hausmann, R. G. (2001). Learning from human tutoring. Cognitive Science, 25(4), 471–533. https://doi.org/10.1207/s15516709cog2504_1
Cohn, C., Rayala, S., Srivastava, N., Fonteles, J. H., Jain, S., Luo, X., Mereddy, D., Mohammed, N., & Biswas, G. (2025). A theory of adaptive scaffolding for LLM-based pedagogical agents (arXiv:2508.01503). arXiv. https://doi.org/10.48550/arXiv.2508.01503
Collins, A., & Halverson, R. (2018). Rethinking education in the age of technology: The digital revolution and schooling in America (2nd ed.). Teachers College Press.
Cosyn, E., Uzun, H., Doble, C., & Matayoshi, J. (2021). A practical perspective on knowledge space theory: ALEKS and its data. Journal of Mathematical Psychology, 101, 102512. https://doi.org/10.1016/j.jmp.2021.102512
Doignon, J. P., & Falmagne, J. C. (1985). Spaces for the assessment of knowledge. International Journal of Man-Machine Studies, 23(2), 175–196. https://doi.org/10.1016/S0020-7373(85)80008-8
Drijvers, P., & Sinclair, N. (2024). The role of digital technologies in mathematics education: Purposes and perspectives. ZDM – Mathematics Education, 56(2), 235–248. https://doi.org/10.1007/s11858-023-01535-x
Giner, R. (2024, April 11). Productive struggle: The imperative of friction in AI-driven learning. Kaplan. https://kaplan.com/about/trends-insights/productive-struggle-friction-ai-education
Godsk, M., & Møller, K. L. (2025). Engaging students in higher education with educational technology. Education and Information Technologies, 30(3), 2941–2976. https://doi.org/10.1007/s10639-024-12901-x
Hanushek, E. A., & Woessmann, L. (2015). The economic impact of educational quality. In P. Dixon, S. Humble, & C. Counihan (Eds.), Handbook of international development and education (pp. 6–19). Edward Elgar Publishing. https://hanushek.stanford.edu/publications/economic-impact-educational-quality
Hillmayr, D., Ziernwald, L., Reinhold, F., Hofer, S. I., & Reiss, K. M. (2020). The potential of digital tools to enhance mathematics and science learning in secondary schools: A context-specific meta-analysis. Computers & Education, 153, 103897. https://doi.org/10.1016/j.compedu.2020.103897
Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), 1–55. https://doi.org/10.1145/3703155
Jermakowicz, E. K. (2023). The coming transformative impact of large language models and artificial intelligence on global business and education. Journal of Global Awareness, 4(2), 3. https://doi.org/10.24073/jga/4/02/03
Kamoi, R., Zhang, Y., Zhang, N., Han, J., & Zhang, R. (2024). When can LLMs actually correct their own mistakes? A critical survey of self-correction of LLMs. Transactions of the Association for Computational Linguistics, 12, 1417–1440. https://doi.org/10.1162/tacl_a_00713
Kilpatrick, J., Swafford, J., & Findell, B. (Eds.) (2001). Adding it up: Helping children learn mathematics. National Academy Press.
Koedinger, K. R., & Aleven, V. (2007). Exploring the assistance dilemma in experiments with cognitive tutors. Educational Psychology Review, 19(3), 239–264. https://doi.org/10.1007/s10648-007-9049-0
Laboratorio de Economía de la Educación (LEE). (2025). Informe N° 114: Pruebas Saber 11: cerrando brechas de sector, más no de género y de zona. Pontificia Universidad Javeriana. https://www.javeriana.edu.co/recursosdb/d/lee/inf-114-informe-saber-11-2025-lee
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9471. https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html
Liu, J., Huang, Z., Xiao, T., Sha, J., Wu, J., Liu, Q., Wang, S., & Chen, E. (2024). SocraticLM: Exploring Socratic personalized teaching with large language models. Advances in Neural Information Processing Systems, 37, 85693–85721. https://proceedings.neurips.cc/paper_files/paper/2024/hash/9bae399d1f34b8650351c1bd3692aeae-Abstract-Conference.html
Liu, Z., Agrawal, P., Singhal, S., Madaan, V., Kumar, M., & Verma, P. K. (2025). LPITutor: An LLM based personalized intelligent tutoring system using RAG and prompt engineering. PeerJ Computer Science, 11, e2991. https://doi.org/10.7717/peerj-cs.2991
Major, L., Francis, G. A., & Tsapali, M. (2021). The effectiveness of technology-supported personalised learning in low- and middle-income countries: A meta-analysis. British Journal of Educational Technology, 52(5), 1935–1964. https://doi.org/10.1111/bjet.13116
McKenney, S., & Reeves, T. C. (2018). Conducting educational design research (2nd ed.). Routledge. https://doi.org/10.4324/9781315105642
Ni, S., Bi, K., Yu, L., & Guo, J. (2024). Are large language models more honest in their probabilistic or verbalized confidence? (arXiv:2408.09773). arXiv. https://doi.org/10.48550/arXiv.2408.09773
OECD. (2019). PISA 2018 results (Volume I): What students know and can do. OECD Publishing. https://doi.org/10.1787/5f07c754-en
Oreopoulos, P., Gibbs, C., Jensen, M., & Price, J. (2024). Teaching teachers to use computer assisted learning effectively: Experimental and quasi-experimental evidence (Working Paper No. 32388). National Bureau of Economic Research. https://www.nber.org/papers/w32388
Pane, J. F., Griffin, B. A., McCaffrey, D. F., & Karam, R. (2014). Effectiveness of Cognitive Tutor Algebra I at scale. Educational Evaluation and Policy Analysis, 36(2), 127–144. https://doi.org/10.3102/0162373713507480
Prensky, M. (2001). Digital natives, digital immigrants part 1. On the Horizon, 9(5), 1–6. https://doi.org/10.1108/10748120110424816
Roschelle, J., Feng, M., Murphy, R. F., & Mason, C. A. (2016). Online mathematics homework increases student achievement. AERA Open, 2(4), 1–12. https://doi.org/10.1177/2332858416673968
Selwyn, N. (2022). Education and technology: Key issues and debates (3rd ed.). Bloomsbury Academic.
Shute, V. J. (2008). Focus on formative feedback. Review of Educational Research, 78(1), 153–189. https://doi.org/10.3102/0034654307313795
STEM Education Coalition. (2019). 2019 annual report. https://www.stemedcoalition.org/
Tapscott, D. (2009). Grown up digital: How the net generation is changing your world. McGraw-Hill.
Zhang, Y., Wang, P., Jia, W., Zhang, A., & Chen, G. (2025). Dynamic visualization by GeoGebra for mathematics learning: A meta-analysis of 20 years of research. Journal of Research on Technology in Education, 57(2), 437–458. https://doi.org/10.1080/15391523.2023.2250886

