Proto_AGI's picture

Proto_AGI PRO

mayafree

·

AI & ML interests

None yet

Recent Activity

updated a Space 3 days ago

mayafree/TRELLIS.2-Text-to-3D22

liked a Space 12 days ago

ginigen-ai/site-agent

repliedto their post 14 days ago

Leaderboard of Leaderboards — A Real-Time Meta-Ranking of AI Benchmarks https://huggingface.co/spaces/MAYA-AI/all-leaderboard Hundreds of AI leaderboards exist on HuggingFace. Knowing which ones the community actually trusts has never been easy — until now. Leaderboard of Leaderboards (LoL) ranks the leaderboards themselves, using live HuggingFace trending scores and cumulative likes as the signal. No editorial curation. No manual selection. Just what the global AI research community is actually visiting and endorsing, surfaced in real time. Sort by trending to see what is capturing attention right now, or by likes to see what has built lasting credibility over time. Nine domain filters let you zero in on what matters most to your work, and every entry shows both its rank within this collection and its real-time global rank across all HuggingFace Spaces. The collection spans well-established standards like Open LLM Leaderboard, Chatbot Arena, MTEB, and BigCodeBench alongside frameworks worth watching. FINAL Bench targets AGI-level evaluation across 100 tasks in 15 domains and recently reached the global top 5 in HuggingFace dataset rankings. Smol AI WorldCup runs tournament-format competitions for sub-8B models scored via FINAL Bench criteria. ALL Bench aggregates results across frameworks into a unified ranking that resists the overfitting risks of any single standard. The deeper purpose is not convenience. It is transparency. How we measure AI matters as much as the AI we measure.

View all activity

Organizations

upvoted an article 18 days ago

Article

🏟️ Smol AI WorldCup: A 5-Axis Benchmark That Reveals What Small Language Models Can Really Do

18 days ago

•

38

upvoted an article 19 days ago

Article

MARL: Runtime Middleware That Reduces LLM Hallucination Without Fine-Tuning

19 days ago

•

15

upvoted an article 20 days ago

Article

Structural Problems in AI Benchmarking and the Case for a Unified Evaluation Framework

20 days ago

•

12

upvoted 2 articles about 1 month ago

Article

Do Bubbles Form When Tens of Thousands of AIs Simulate Capitalism?

Feb 24

•

17

Article

FINAL Bench: The Real Bottleneck to AGI Is Self-Correction

Feb 21

•

20

upvoted a collection about 1 month ago

FINAL Bench

World's First Functional Metacognition Benchmark. "Not how much AI knows — but whether it knows what it doesn't know, and can fix it." • 2 items • Updated Feb 21 • 4