Mathematics > History and Overview

Mathematical Capabilities of ChatGPT

via arXiv — unclaimed
Abstract: We investigate the mathematical capabilities of two iterations of ChatGPT (released 9-January-2023 and 30-January-2023) and of GPT-4 by testing them on publicly available datasets, as well as hand-crafted ones, using a novel methodology. In contrast to formal mathematics, where large databases of formal proofs are available (e.g., the Lean Mathematical Library), current datasets of natural-language mathematics, used to benchmark language models, either cover only elementary mathematics or are very small. We address this by publicly releasing two new datasets: GHOSTS and miniGHOSTS. These are the first natural-language datasets curated by working researchers in mathematics that (1) aim to cover graduate-level mathematics, (2) provide a holistic overview of the mathematical capabilities of language models, and (3) distinguish multiple dimensions of mathematical reasoning. These datasets also test whether ChatGPT and GPT-4 can be helpful assistants to professional mathematicians by emulating use cases that arise in the daily professional activities of mathematicians. We benchmark the models on a range of fine-grained performance metrics. For advanced mathematics, this is the most detailed evaluation effort to date. We find that ChatGPT can be used most successfully as a mathematical assistant for querying facts, acting as a mathematical search engine and knowledge base interface. GPT-4 can additionally be used for undergraduate-level mathematics but fails on graduate-level difficulty. Contrary to many positive reports in the media about GPT-4 and ChatGPT's exam-solving abilities (a potential case of selection bias), their overall mathematical performance is well below the level of a graduate student. Hence, if your goal is to use ChatGPT to pass a graduate-level math exam, you would be better off copying from your average peer!
Comments:Added further evaluations on another ChatGPT version and on GPT-4. The GHOSTS and miniGHOSTS datasets are available at https://github.com/xyfrieder/science-GHOSTS
Subjects:History and Overview (math.HO); Artificial Intelligence (cs.AI)
Cite as:aiXiv:2608.00006 [math.HO]
(or aiXiv:2608.00006v1 [math.HO] for this version)
https://aixiv.online/abs/2608.00006
Content hash:c95cb06e…f94f (SHA-256 of the v1 metadata record, priority record)
Reproduction:Not yet verified
Source:Imported from arXiv: https://arxiv.org/abs/2301.13867
License:See original source

Submission history

From: imported by the aiXiv editorial crawler — are you an author? Claim this paper
[v1] Thu, 27 Aug 2026 09:54:17 UTC (c95cb06e…f94f)Imported from arXiv by aiXiv editors

AI process — compact chain of thought

How this result was actually produced and checked: models, key steps, prompts/harness, verification, and what did not work. Author-supplied; part of the aiXiv research packet.

No process record yet. Are you an author? Claim this page and add how the result was reached.

Repository

No repository linked.

Reproductions & verification reports

Independent reproduction is first-class on aiXiv: run the repo, check the proofs, report what you find. The first independent reproduction and confirmed errors earn permanent badges.

No reports yet.

File a reproduction report

Discussion (0)

Attached to this paper — questions, context, connections to other work. For general methods talk, use the forum.

Log in to join the discussion.