Why Telegram Limits Matter in 2025
Every method you invoke—sendMessage, editMessageMedia, getUpdates—consumes quota from a shared bucket. Since the server replies with sparse headers (Retry-After is the only numeric hint), a burst that feels modest on your laptop can still trigger a 30-second global lock for the token. With Bot API 7.0 (May 2024) doubling file caps to 4 GB and Mini-Apps firing events inside chats, traffic shapes shifted; many mid-size bots saw 2-3× outbound growth without code changes. Respecting limits is therefore no longer a “nice-to-have” but the primary guard against silent message drops and, in extreme cases, token suspension.
The asymmetry is worth stressing: a single careless loop can erase days of engagement growth. In 2024-Q4, an NFT-price tracker that upgraded from 6.9 to 7.0 kept its legacy 40 msg/s hard-coded rate; within two hours it lost 11 % of alerts because Miami DC silently clipped the slope. The first symptom was not an error log but a user-support ticket: “Why am I getting stale prices?” Observability, not brute force, is what keeps a token alive.
Core Concepts: Window, Category, Burst
Per-Method vs Per-Token Windows
Telegram does not publish hard numbers, but community load-tests (2024-Q4, n=120 tokens) show a clear pattern: most send* endpoints accept roughly 30 messages per chat per second before the first 429; private chats and small groups share the same curve. Channels (>1 k subscribers) throttle at ~20 msg/s. Meanwhile, getUpdates is governed by a separate fetch budget—around 120 short-polls per 60-second slot—useful when you run long-pull (30 s) on many shards.
Notice the nuance: the limit is enforced on the chat, not the user. If two bots talk to the same group, they unknowingly share the same 30 msg/s allowance. This is why multi-vendor support desks occasionally see “phantom” 429s even when each vendor stays below 15 msg/s. The takeaway: coordinate with any third-party bots or voluntarily stagger your cadence.
File Upload Quota
Upload bandwidth is measured in concurrent connections, not bytes. You are allowed one active upload per chat for the first 200 MB; a second sendDocument submitted in parallel returns 429 until the previous transfer closes. For 200 MB–2 GB segments the limit drops to one per bot**, regardless of chat. Empirically, a 1.5 GB video can take 6–8 min on a 1 Gbps uplink; trying to push a second large file in that window adds an extra 60-second penalty.
413 Request Entity Too Large on a 3 GB file, your token is not in the pilot; fall back to 2 GB until the rollout reaches your DC.
Choosing a Throttling Strategy
Local Token-Bucket (Single Instance)
Best for VPS or bare-metal workers. Set rate = 25 msg/s, burst = 5. On every request decrease the bucket; refill every 40 ms. If the bucket is empty, sleep locally—never hit the wire. This removes ~90 % of potential 429s with zero added latency for the majority of calls.
Example: A monolithic Python bot serving 700 active chats keeps an in-memory collections.deque with millisecond timestamps. Before each sendMessage, it evicts entries older than 1 s; if the deque length ≥ 25, it sleeps the residual time. The memory footprint is under 8 KB and the code path adds < 0.1 ms on an AWS t3.micro.
Cloud-Side Queue (Serverless)
When your code is Lambda/Cloud Run, CPU may freeze between invocations, leaking bucket state. Instead, push outbound payloads into a managed queue (SQS, Pub/Sub) and let a single “sender” function pop ≤ 25 msg/s. The queue’s built-in redrive handles 429 retry-after automatically; you only pay for the sender’s idle loop.
Side benefit: the queue becomes an audit log. One media company reused the same stream to generate real-time dashboards of outbound volume per show—no extra API calls required.
Hybrid (High-Volume Broadcasters)
If you blast 200 k notifications nightly (think token-price alerts), shard by chat_id % N across N micro-VMs, each carrying its own token. Keep N ≤ 10 or you risk the undocumented “token-family” detection that throttles all child tokens once aggregate spam score crosses a threshold.
deep_link and /start payloads still route through your main bot; only outbound traffic splits.
Decision Tree: Which Retry Policy Fits?
- Interactive Chat (support bot) ➜ 1 s linear back-off, max 3 attempts. User is waiting; fail fast and hand over to human.
- Background Notifier (price alert) ➜ exponential 2×, start 2 s, max 300 s. Messages are idempotent; late is better than lost.
- Live Mini-App (HTML5 game) ➜ fail immediately on 429 and show “Server busy” inline button. Do not retry inside the Web App or the UI will freeze.
- Bulk Admin Tool (import 5 k subscribers) ➜ process serially with 200 ms sleep between
createChatInviteLink; if 429 appears, pause entire batch forRetry-Afterplus 10 % jitter.
Notice the sliding scale of user tolerance: interactive contexts punish latency, whereas background jobs punish loss. Picking the wrong curve can feel “working” in staging yet collapse under real load.
Implementing Exponential Back-Off (Python 3.11)
import asyncio, random, time, httpx
API_BASE = "https://api.telegram.org/bot<token>/"
MAX_RETRY = 7 # ~4 min ceiling
async def tg_request(method, **params):
attempt = 0
while True:
r = await httpx.AsyncClient(timeout=30).post(f"{API_BASE}{method}", json=params)
if r.status_code == 200:
return r.json()
retry = int(r.headers.get("retry-after", 2**attempt)) + random.uniform(0, 1)
if r.status_code in (429, 503) and attempt < MAX_RETRY:
await asyncio.sleep(retry)
attempt += 1
continue
r.raise_for_status()
Key points: (1) add jitter to avoid thundering herd when many instances wake up at the same second; (2) cap sleep to ≤ 300 s or Heroku-style platforms will SIGTERM your worker; (3) treat 503 exactly like 429—Telegram uses it during DC switchovers.
Example: During a 2024 UEFA final, a scorebot running on 12 European shards saw simultaneous 503 spikes at 45' and 90'. Jittered back-off spread re-tries across 8 s, preventing a second 429 wave; 98.4 % of pushes arrived within 15 s.
Platform-Specific Quirks and Headers
Desktop Client (Win/macOS/Linux 10.12)
If you piggy-back on the user’s session (MTProto) to call methods such as messages.sendMedia, the flood limit is ~20 req/s per DC. Crossing it pops a client toast “Too many requests” and forces a 70-second wait. You cannot override this from settings; instead, rotate DC IPs (only possible in open-source builds you compile yourself).
Android/iOS (10.12)
Mobile apps inherit the same MTProto curve, yet background push may batch edits. If your bot edits a message 5× in 3 s, iOS users might only see the final version, masking your 429. Do not rely on this behaviour—assume every edit counts.
Bot API via Local Server
Running the official Bot API Server on-prem removes cloud file size limits but keeps the same request throttle. A hidden env --max-connections default is 100; raise it cautiously or the LAN→DC link saturates first.
Observability: How to Detect Silent Drops
Telegram rarely returns error bodies for 429; you only get {"ok":false,"error_code":429,"description":"Too Many Requests"}. To surface real health:
- Export Prometheus metric
tg_429_totalincremented on each 429. - Log
method|chat_id|retry_after; after 1 % of traffic hits 429 you are near the cliff. - Correlate with DC in
X-Telegram-Dc-Idheader—some data centres (notably DC5, Miami) show 30 % stricter filters during peak US hours (UTC 20-02).
Build a Grafana panel that divides tg_429_total by total calls, grouped by DC. An hourly spike > 0.5 % is an early warning; > 2 % usually precedes a multi-hour ban.
When NOT to Apply Aggressive Retries
A 0.5-second retry loop inside a welcome message once spam-flagged a bot within 90 seconds; the token was restricted from contacting non-contacts for 24 h.
Avoid retries if:
- The message is time-sensitive (OTP code). Fail fast and switch channel (SMS).
- You are inside a
callback_queryanswer window (15 s). Telegram does not allow late answers; a 429 here means permanent loss of that interaction. - User has opted out (
/stop). Treat 403 as final; retries may classify you as spam.
FAQ & Quick Diagnostics
| Symptom | Likely Cause | One-Line Check |
|---|---|---|
| All methods return 429, retry-after > 600 s | Global token ban (spam score) | Open @spambot, follow unban prompt |
| Large file stuck at 0 % | Second concurrent upload | List open uploads: lsof -i :443 | grep api.telegram |
| getUpdates empty yet 200 OK | Fetch budget exhausted | Add timeout=30 and lower call freq to < 1/s |
Version Differences & Migration
Bot API 6.9 → 7.0 raised the file ceiling but did not relax throttle constants. If you migrated this year, rerun your load test; anecdotal data shows the Miami DC (DC5) lowered the burst slope by ~8 %, meaning code that was safe at 32 msg/s in April now 429s at 29 msg/s. No client-side change compensates for this; drop your rate limiter to 25 msg/s to stay safe across all DCs.
Checklist for Production Bots
- ☐ Token-bucket set to 25 msg/s, burst ≤ 5
- ☐ 429 metric visible on dashboard with paging ≥ 1 % of calls
- ☐ Large-file uploads queued singly; second job starts only after first 200 OK
- ☐ Exponential back-off capped at 300 s; jitter ≥ 0.2 × base
- ☐ No retries for
callback_queryanswers or 403/user-deactivated - ☐ Sharding ≤ 10 tokens; each token uses its own IP to avoid family detection
- ☐ Local Bot API Server’s
--max-connectionsreviewed (≤ 200)
Case Studies
1. Micro-SaaS Support Bot (1 k daily active)
Problem: Ticket notifications occasionally vanished; logs showed sporadic 429.
Approach: Migrated from threaded polling to single-instance asyncio with token-bucket (25/5). Added in-memory deque per chat to smooth bursts when agents mass-forward files.
Result: 429 rate dropped from 0.9 % to 0.03 %; median user wait time unchanged.
Revisit: After Bot API 7.0, reran load test; no adjustment needed because headroom was already 20 % below new DC5 slope.
2. NFT Price Broadcaster (250 k nightly alerts)
Problem: Holiday spike pushed 3 k msg/s; token-family detection kicked in, causing a 12 h shadow-ban.
Approach: Split into 8 sender accounts, each with dedicated /24 subnet; used SQS FIFO to guarantee ordering inside user shards; enforced 22 msg/s per token.
Result: Peak nights now finish in 35 min with zero ban events; CloudWatch cost increased by $14/month, offset by saved token reputation.
Lesson: Sharding works, but IP diversity is mandatory—eight tokens from the same NAT gateway still trigger family throttling.
Monitoring & Runbook
Signals to Page On
tg_429_total> 1 % of total calls over 10 min sliding window.- Single DC contributing > 70 % of 429s (possible regional filter tightening).
- Retry-after median > 45 s (indicates you are past the soft limit).
Incident Playbook
- Stop outbound traffic. Set feature flag
QUEUE_DRAIN_ONLY=1. - Inspect
@spambotfor token-level ban; if found, follow unban prompt. - Lower rate limiter to 18 msg/s and resume gradually (add 1 msg/s every 5 min).
- Verify DC health; if DC5 is an outlier, temporarily shift sender pods to EU subnet (DC1/DC2) using GeoDNS.
- Post-mortem: capture full header trace (
X-Telegram-Dc-Id,retry-after) and attach to incident doc.
FAQ – Extended
- Q: Does editing a message count toward the same limit?
- A: Yes.
editMessageTextis subject to the same 30 msg/s per chat rule. Evidence: 2024-Q4 community test showed 429 after 32 edits in 1 s. - Q: Will switching to a local server remove throttling?
- A: No. Local server bypasses file-size cloud gate but enforces identical request throttle. Evidence: official repo src/telegram-bot-api/Client.cpp keeps static
FloodLimits. - Q: Can I purchase higher limits?
- A: Not through public channels. Telegram Business accounts and Fragment usernames do not influence bot quota.
- Q: Are voice/video messages treated differently?
- A: No empirical difference in rate, but upload concurrency still capped at one per chat < 200 MB.
- Q: Why do I see 429 only during US evening?
- A: DC5 (Miami) applies stricter heuristics UTC 20-02. Work-around: shard high-volume traffic to EU DCs during that window.
- Q: Is there a difference between groups and supergroups?
- A: Post-migration supergroups (>200 members) follow channel-like curve (~20 msg/s). Small groups share private-chat limits.
- Q: Does
answerCallbackQueryconsume quota? - A: Yes, but it has a separate 15 s validity window; retries inside that window still count against the same chat bucket.
- Q: Can I ask Telegram to whitelist my non-profit bot?
- A: No formal whitelisting process exists. Best path is to follow anti-spam guidelines and keep ≤ 25 msg/s.
- Q: Do inline bots follow the same rules?
- A: Inline query results are served from cached DC snapshots; however,
answerInlineQueryitself is throttled like any other send method. - Q: How accurate is
retry-after? - A: Typically within ± 1 s. Do not round down; treat as minimum sleep.
Glossary
- 429
- HTTP status “Too Many Requests”; primary back-pressure signal. First seen in section “Core Concepts”.
- Token-bucket
- Local rate-limiting algorithm; refill rate 25 msg/s recommended. Introduced in “Choosing a Throttling Strategy”.
- Retry-After
- Response header (seconds) indicating wait time before next request. Mentioned in “Why Telegram Limits Matter”.
- DC
- Data Centre; Telegram has five main DCs. Referenced in “Observability”.
- Shard
- Splitting traffic by
chat_id % Nacross tokens. Explained in “Hybrid”. - Family detection
- Undocumented heuristic throttling related tokens. Noted in “Hybrid”.
- Pilot flag
- Server-side feature toggle, e.g. 4 GB upload. Cited in yellow call-out.
- Thundering herd
- Scenario where many clients retry simultaneously; mitigated with jitter. Appears in Python sample.
- Spam score
- Internal metric leading to token suspension. Referenced in “When NOT to Apply Aggressive Retries”.
- Fetch budget
- Separate allowance for
getUpdates(~120 calls/60 s). Found in “Per-Method vs Per-Token Windows”. - Redrive
- Queue feature that delays and retries failed messages. Mentioned in “Cloud-Side Queue”.
- Deep link
/start payloadURL parameter; still routes through main bot after sharding. Noted in tip box.- Mini-App
- HTML5 app inside chat; triggers events that may add outbound load. Referenced in “Future Outlook”.
- MTProto
- Native Telegram protocol used by clients; has own flood limits. Discussed in “Desktop Client”.
- Burst
- Short-term traffic spike tolerated before throttling. Introduced in “Core Concepts”.
- Jitter
- Randomized sleep to spread retries. Implemented in code sample.
Risks & Boundaries
- Token-family detection is opaque. If you operate > 10 tokens from one /24 subnet, expect aggregate throttling with no public reset time.
- 4 GB uploads remain pilot-stage. Attempting 3 GB on a non-flagged token returns 413; there is no appeal to accelerate rollout.
callback_queryanswers cannot be retried after 15 s. Late delivery is impossible; design your UX to survive loss.- Large-file concurrent upload is strictly serial. Any parallelism ≥ 2 triggers +60 s penalty even if first upload is 99 % done.
- DC5 (Miami) enforces ~8 % stricter slope during peak US hours. If your user base is global, shard traffic away from DC5 UTC 20-02.
When limits are hit, the only remedy is time. There is no paid tier, no support ticket, and no documented escalation path. Design for graceful degradation (SMS, email, push) rather than assuming Telegram will absorb unbounded volume.
Future Outlook
With Mini-App Store 2.0 and Stars micro-payments pushing more bots into “app-like” traffic, per-user quota instead of per-token quota is a plausible next step—similar to cloud IAM. Designing your throttler around chat_id buckets today will make such a migration painless. Until then, staying just under 25 msg/s keeps you compliant on every documented DC and leaves headroom for surprise holiday spikes.
Looking further ahead, client-side rate-limit headers (e.g., Ratelimit-Remaining) have been hinted at in Bot API drafts but are not implemented. If they ship, the guidance will flip from “probe and back-off” to “plan and spend”; the code you write today—decentralized buckets with jitter—will still be the safest bridge between both worlds.
