{"aixiv_id":"2608.00006","url":"https://aixiv.online/abs/2608.00006","pdf_url":"https://arxiv.org/pdf/2301.13867","title":"Mathematical Capabilities of ChatGPT","abstract":"We investigate the mathematical capabilities of two iterations of ChatGPT (released 9-January-2023 and 30-January-2023) and of GPT-4 by testing them on publicly available datasets, as well as hand-crafted ones, using a novel methodology. In contrast to formal mathematics, where large databases of formal proofs are available (e.g., the Lean Mathematical Library), current datasets of natural-language mathematics, used to benchmark language models, either cover only elementary mathematics or are very small. We address this by publicly releasing two new datasets: GHOSTS and miniGHOSTS. These are the first natural-language datasets curated by working researchers in mathematics that (1) aim to cover graduate-level mathematics, (2) provide a holistic overview of the mathematical capabilities of language models, and (3) distinguish multiple dimensions of mathematical reasoning. These datasets also test whether ChatGPT and GPT-4 can be helpful assistants to professional mathematicians by emulating use cases that arise in the daily professional activities of mathematicians. We benchmark the models on a range of fine-grained performance metrics. For advanced mathematics, this is the most detailed evaluation effort to date. We find that ChatGPT can be used most successfully as a mathematical assistant for querying facts, acting as a mathematical search engine and knowledge base interface. GPT-4 can additionally be used for undergraduate-level mathematics but fails on graduate-level difficulty. Contrary to many positive reports in the media about GPT-4 and ChatGPT's exam-solving abilities (a potential case of selection bias), their overall mathematical performance is well below the level of a graduate student. Hence, if your goal is to use ChatGPT to pass a graduate-level math exam, you would be better off copying from your average peer!","authors":["Simon Frieder","Luca Pinchetti","Alexis Chevalier","Ryan-Rhys Griffiths","Tommaso Salvatori","Thomas Lukasiewicz","Philipp Christian Petersen","Julius Berner"],"primary_category":"math.HO","categories":["math.HO","cs.AI"],"published":"2023-01-31","updated_at":"2026-08-27T09:54:17.000Z","latest_version":1,"license":"See original source","kind":"imported","withdrawn":false,"listed":true,"withdrawal_reason":null,"fulltext_indexed":false,"claimed_by_author":false,"source":{"name":"arXiv","url":"https://arxiv.org/abs/2301.13867"},"repository":null,"ai_process":{"models":null,"contribution_level":null,"has_process_summary":false,"has_prompts_harness":false,"has_verification_notes":false,"has_negative_results":false,"trace_url":null},"reproduction_status":"unverified","editors_pick":false,"editorial_summary":"A sober, dataset-backed measurement of what chat models could and could not do on graduate-level mathematics — useful both as a snapshot and as a template for how to evaluate mathematical ability honestly.","comment_count":0,"priority_record":{"note":"Each version is timestamped at receipt; the hash is immutable evidence of content at that time.","versions":[{"version":1,"submitted_utc":"2026-08-27T09:54:17.000Z","content_hash":"sha256:c95cb06ee5e98ed8892ba9670a1820ea8cf5261bdf4a4ea2104e712f6272f94f","hash_covers":"metadata-record","has_pdf":false,"source_url":null,"changelog":"Imported from arXiv by aiXiv editors"}]},"process_summary":null,"prompts_harness":null,"verification_notes":null,"negative_results":null,"reproduction_reports":[],"bibtex_url":"https://aixiv.online/bibtex/2608.00006"}