Hugging Face
The central hub for AI models, datasets, Spaces, libraries, and open-source ML collaboration.
LLM APIs, model hosting, inference platforms, and local runtimes for running and deploying AI models.
Models and infrastructure is the builder's category: where AI models get hosted, served, connected, and measured. If you're shipping an AI product rather than using one, these are your suppliers.
Hugging Face is the gravity well — the hub where open models, datasets, and demos live, plus hosting on top. Replicate and fal.ai sell serverless inference: call a model via API, pay per second or per generation, never touch a GPU. Modal and Baseten are deployment platforms for when you outgrow serverless and want your own endpoints with real scaling controls. Pinecone is the vector database most retrieval-augmented apps started on. LiteLLM is the open-source gateway that normalizes a hundred model APIs behind one interface. Heurist and x402 round out the infrastructure edge.
Unusually, this category also includes the scorekeepers: LMArena runs the crowdsourced model leaderboard, Artificial Analysis benchmarks price and performance across providers, and Stanford HELM and SWE-bench publish rigorous academic and coding evaluations. That mix is deliberate — choosing infrastructure without benchmarks is guessing.
The differentiators are practical: cold-start latency, GPU pricing, developer experience, and lock-in. Costs range from free (the benchmarks) to enterprise contracts. The sound pattern: prototype on serverless inference, validate with public benchmarks plus your own evals, and only then commit to dedicated deployment.
The central hub for AI models, datasets, Spaces, libraries, and open-source ML collaboration.
Community-powered model leaderboard for comparing AI systems through real user battles.
Software engineering benchmark and leaderboard for evaluating AI coding agents on real GitHub issues.
Open-source LLM gateway for routing, logging, and cost control
Serverless AI infrastructure for running code, jobs, containers, and GPUs from Python.
Production AI inference platform for deploying, optimizing, and scaling models.
Independent AI model benchmarks for intelligence, speed, pricing, context, and modalities.
Managed vector database for semantic search, RAG, recommendations, and AI retrieval.
Fast generative media APIs for images, video, audio, and creative model workflows.
Open framework for holistic, reproducible evaluation of language and multimodal models.
Run open and community AI models from a web playground or API.
Open payment protocol for agentic and API-based machine-to-machine commerce.
1–12 of 13 tools