Paper: A UNIVERSITY-LEVEL BENCHMARK FOR EVALUATING MATHEMATICAL SKILLS IN LLMS
Toloka
company
Verified
AI & ML interests
Human-expert data for frontier reasoning, safety and agentic AI
Recent Activity
Organization Card
Hey, this is Toloka!
datasets
12
toloka/vist
Viewer
•
Updated
•
39.3k
•
78
toloka/VOX-DUB
Viewer
•
Updated
•
7.58k
•
301
•
10
toloka/JEEM
Viewer
•
Updated
•
2.2k
•
58
•
11
toloka/beemo
Viewer
•
Updated
•
2.19k
•
299
•
18
toloka/u-math
Viewer
•
Updated
•
1.1k
•
174
•
24
toloka/mu-math
Viewer
•
Updated
•
1.08k
•
70
•
23
toloka/CLESC
Viewer
•
Updated
•
500
•
26
•
2
toloka/VoxDIY-RusNews
Updated
•
494
•
3
toloka/CrowdSpeech
Updated
•
187
•
5
toloka/crowdkit-datasets
Updated
•
3.01k