Hugging Face Inference Endpoints logo

Hugging Face Inference Endpoints

Functions 4.8/10

Dedicated autoscaling inference billed per minute of uptime. CPU from $0.032/core/hr; GPUs T4 $0.50/hr, L4 $0.80/hr, A100 and up to ~$45/hr; can scale to zero.

Hugging Face Inference Endpoints scores 4.8/10 for free value in Functions. The main constraint: Paid — From $0.032/CPU-core-hr; GPUs from $0.50/hr (T4). Dedicated autoscaling inference billed per minute of uptime. CPU from $0.032/core/hr; GPUs T4 $0.50/hr, L4 $0.80/hr, A100 and up to ~$45/hr; can scale to zero.

Category Functions Kind Paid plan

Free value score

4.8 / 10
Show the math
Overall 4.8/10
value = 4.8/100
weighted_components_v1
Invocation value 1/10
value = 0.01 ×
fixed_log_v1
Runtime fit 9/10
value = 9/100
policy_score_v1
Deploy friction 8/10
value = 8/100
policy_score_v1

Key metrics

$.032/coreh
T4 $0.50/hr
per-min

What you get

Paid — From $0.032/CPU-core-hr; GPUs from $0.50/hr (T4). Dedicated autoscaling inference billed per minute of uptime. CPU from $0.032/core/hr; GPUs T4 $0.50/hr, L4 $0.80/hr, A100 and up to ~$45/hr; can scale to zero.

~2 T4 GPU-hours per $1 ($0.50/hr); per-minute billing and scale-to-zero make it efficient for spiky inference. CPU endpoints extremely cheap for light

Sources

paid value serverless-functions