Google Search Console API Automation: A Python Guide to Daily Pulls, Quotas and Alerts
Automate the Google Search Console API with Python: service-account auth, daily pulls past 25,000 rows, quotas, URL Inspection and drop alerts.
Long Nguyen
Fullstack Developer · AI Engineer · Researcher
What the Search Console API can and cannot automate
The Search Console API is a read-heavy interface to four things: performance data, URL inspection, sitemaps and the list of properties you own. It is not a way to force indexing of ordinary pages, and it does not expose every report you see in the web interface. Before writing any code, map each job you want automated to the endpoint that actually backs it.
| Job | What to use | Hard limit to plan around |
|---|---|---|
| Daily clicks, impressions, CTR and position by page or query | searchAnalytics.query |
25,000 rows per request; 50,000 rows per day per search type |
| Check if a URL is indexed, its canonical and last crawl | urlInspection.index.inspect |
2,000 queries per day per site |
| Submit or list sitemaps | sitemaps.submit / sitemaps.list |
Needs the read/write scope |
| Full performance history in a warehouse | Bulk export to BigQuery | Everything except anonymized queries |
| Ask Google to crawl an ordinary page now | Not available | The Indexing API covers only job posting and livestream pages |
That last row is where most failed automation projects start: they plan a nightly script that pushes every new article to Google, then discover no supported endpoint does that for normal pages. Section 7 covers it.
Authenticate: service account or OAuth for unattended jobs
Every request must carry an OAuth 2.0 token; API keys do not unlock private property data. Google defines two scopes: webmasters.readonly for read-only access and webmasters for read/write. Use the read-only scope for reporting jobs and add write access only for the script that submits sitemaps.
For a job that runs on a schedule with nobody watching, a service account is the practical choice. A user-consent OAuth flow needs a browser and a refresh token, and refresh tokens issued while your consent screen is still in testing status expire after about a week, so the job dies silently.
- In Google Cloud, create a project and enable the Search Console API.
- Create a service account and download its JSON key. Treat the key like a password.
- In Search Console, open Settings, then Users and permissions, and add the service account email as a user on every property the job needs. Grant the lowest permission that covers the job.
- Use the property identifier exactly as Search Console shows it:
https://www.example.com/for a URL-prefix property orsc-domain:example.comfor a Domain property.
from google.oauth2 import service_account
from googleapiclient.discovery import build
SCOPES = ['https://www.googleapis.com/auth/webmasters.readonly']
creds = service_account.Credentials.from_service_account_file(
'service-account.json', scopes=SCOPES)
gsc = build('searchconsole', 'v1', credentials=creds)
print(gsc.sites().list().execute()) # properties this account can see
If sites().list() returns an empty list, authentication worked but the service account was never added to a property. Enabling the API in the cloud project is not enough; access is granted per property inside Search Console. This is the single most common cause of 403 errors in new setups.
Pull Search Console data with Python beyond the 25,000-row cap
The rowLimit parameter accepts 1 to 25,000 and defaults to 1,000, so a query that omits it silently returns a fraction of your data. Page through results by raising startRow until a response comes back empty. Google's own guidance, in its guide to retrieving all your performance data, is to run one query per day for one day of data, which keeps you inside quota and gives you a clean unit to retry.
import time, random
from googleapiclient.errors import HttpError
PAGE = 25000
def execute(request, tries=5):
for attempt in range(tries):
try:
return request.execute()
except HttpError as err:
quota = 'quota' in str(err).lower()
if attempt == tries - 1:
raise
if quota:
time.sleep(15 * 60) # short-term load quota: wait 15 minutes
elif err.resp.status in (429, 500, 503):
time.sleep(2 ** attempt + random.random())
else:
raise
def pull_day(site, day, search_type='web',
dims=('date', 'page', 'query', 'device', 'country')):
rows, start = [], 0
while True:
body = {
'startDate': day, 'endDate': day,
'dimensions': list(dims), 'type': search_type,
'rowLimit': PAGE, 'startRow': start,
}
resp = execute(gsc.searchanalytics().query(siteUrl=site, body=body))
batch = resp.get('rows', [])
if not batch:
return rows
rows.extend(batch)
start += PAGE
Why the target date is three days back
Google notes that data is typically available after two to three days, and the request dates are interpreted in Pacific Time, not your local time. A job in Hanoi or Berlin that asks for yesterday by local clock can request a Pacific date that is not finalized yet. Compute the target date in Pacific Time and subtract three days. If you need fresher numbers, pass dataState: 'all'; when you group by date the response metadata includes first_incomplete_date, and every value after it can still change.
Accurate totals and detailed rows are two different queries
When you group by page or query, Search Console may drop rows to keep the computation affordable. That means the sum of a page-and-query pull will not match your real totals. The reliable pattern is to run two queries per day and store them separately.
| Purpose | Dimensions | Trade-off |
|---|---|---|
| Accurate totals | None, or only country and device | No page or query detail |
| Detail for analysis | Page, query, plus optional country and device | Some rows are dropped |
Even the detailed pull has a ceiling: the API exposes at most 50,000 rows per day per search type, sorted by clicks. On a large site the long tail beyond that is invisible to this endpoint, which is the case for the BigQuery export covered below.
Search Console API quotas and limits
These numbers come from Google's Search Console API usage limits page, checked in . Quotas change, so confirm them before you size a job.
| Resource | Scope | Limit |
|---|---|---|
| Search Analytics | Per site | 1,200 queries per minute |
| Search Analytics | Per user | 1,200 queries per minute |
| Search Analytics | Per project | 40,000 per minute; 30,000,000 per day |
| URL Inspection | Per site | 600 per minute; 2,000 per day |
| URL Inspection | Per project | 15,000 per minute; 10,000,000 per day |
| All other resources | Per user | 20 per second; 200 per minute |
The per-minute numbers are rarely the problem. Search Analytics also has a load quota, measured in 10-minute and one-day windows, and that is what trips real pipelines. Load grows with the date range you query, and grouping or filtering by page or by query string is expensive; grouping by both is the most expensive combination. The error message is identical for every kind of quota event, so you diagnose by behavior: if a single query in a quiet 10-minute window still fails, you are over the daily load quota.
- Query one day at a time instead of one wide range.
- Do not re-query data you already stored. Keep the raw rows.
- If you hit the short-term quota, wait 15 minutes; if it persists, cut the page-and-query grouping or shrink the range.
Automate URL Inspection and sitemap submission
The URL Inspection endpoint returns the same index status you see in the web interface: whether the URL is on Google, which canonical Google chose, and when it was last crawled. It reports status only; it does not request indexing.
def inspect(site, url):
body = {'inspectionUrl': url, 'siteUrl': site}
res = execute(gsc.urlInspection().index().inspect(body=body))
s = res['inspectionResult']['indexStatusResult']
return {
'url': url,
'state': s.get('coverageState'),
'last_crawl': s.get('lastCrawlTime'),
'google_canonical': s.get('googleCanonical'),
'user_canonical': s.get('userCanonical'),
}
At 2,000 inspections per day per site, you cannot check a 50,000-URL catalog daily. Spend the budget where a change matters: URLs published in the last week, your top pages by clicks, and pages whose clicks just dropped. Rotate the rest across the month. Store every result so you can alert on a change in coverageState or a canonical mismatch, which is more useful than any single snapshot. Quota applies per property, so splitting a large site into URL-prefix properties for its main sections is a legitimate way to widen the budget.
Sitemap submission needs the read/write scope and one call, gsc.sitemaps().submit(siteUrl=site, feedpath=sitemap_url). It tells Google where the file lives; resubmitting an unchanged file after every deploy adds nothing. What moves crawling is an accurate lastmod in the sitemap itself.
Teams that would rather not maintain this plumbing can hand it over: Netalith builds this kind of reporting and monitoring as part of its SEO automation work.
When the API is not enough: bulk export to BigQuery
If the 50,000-row daily ceiling or the dropped rows from page-and-query grouping hurt your analysis, stop fighting the API. Search Console can schedule a daily export of your performance data to BigQuery. It includes all the performance data available for the property, except anonymized queries. The export lands in two main tables, searchdata_site_impression and searchdata_url_impression.
| Need | Better fit |
|---|---|
| Small or mid-size site, dashboards, alerts | API pull into your own database |
| Long-tail queries beyond 50,000 rows a day | BigQuery bulk export |
| Joining Search Console with revenue or log data | BigQuery bulk export |
| URL index status and sitemaps | API (the export does not carry them) |
One caution from practice: the export starts from the day you configure it, so history before that day is not backfilled. Turn it on early, and keep your API pull running for the period before it. The two are complements, not rivals.
Can the Indexing API request indexing for normal pages?
No. Google documents the Indexing API as a way to notify it when job posting or livestream video pages are added or removed, and it works only for pages carrying JobPosting structured data or BroadcastEvent embedded in a VideoObject. It ships with a default quota of 200 for onboarding and testing, requires approval for more, and Google warns that abusing it, including using multiple accounts to exceed quotas, can get access revoked.
Tutorials that use it to push blog posts or product pages work today by accident, not by design. Building a production workflow on unsupported behavior is a risk you carry alone. For normal pages, the supported levers are a clean sitemap with honest lastmod values, strong internal links from crawled pages, and the inspection data above to find what is stuck.
Schedule the pipeline and alert on traffic drops
A daily job needs four properties: it targets a date three days back in Pacific Time, it is idempotent so a rerun never duplicates rows, it stores raw rows, and it alerts only on changes large enough to matter. Search Console keeps roughly 16 months of data, so your own store is also the only way to compare years.
Store rows with a unique key on (date, page, query, device, country) and upsert them. Then a single query can compare the last seven days with the seven before them and flag pages that lost real traffic.
WITH w AS (
SELECT page,
SUM(CASE WHEN date >= date('now', '-10 day') THEN clicks ELSE 0 END) AS last7,
SUM(CASE WHEN date < date('now', '-10 day') THEN clicks ELSE 0 END) AS prev7
FROM gsc
WHERE date >= date('now', '-17 day')
GROUP BY page
)
SELECT page, prev7, last7
FROM w
WHERE prev7 >= 50 AND last7 < prev7 * 0.7
ORDER BY prev7 - last7 DESC;
The two thresholds are judgment calls, not Google numbers. The 50-click floor keeps tiny pages from paging you over noise, and the 30 percent drop catches real losses without firing on normal weekly variation. Tune both against a month of your own history before sending alerts to a channel people read. Run the job from cron, a CI scheduler or a container job; the schedule matters less than making the step that stores data safe to repeat.
If you want this pipeline running against your own properties without building and maintaining it, request a free quote and describe the reports you need.
FAQ
Frequently asked questions
How many rows can the Search Console API return?
A single request returns at most 25,000 rows, and you page through results with startRow. Separately, the API exposes at most 50,000 rows of data per day per search type, sorted by clicks. For anything beyond that, use the BigQuery bulk export.
Can I use a service account with the Search Console API?
Yes. Create a service account, enable the Search Console API in its cloud project, then add the service account email as a user on each property in Search Console. Without that per-property step, calls return empty lists or 403 errors even though authentication succeeded.
How often should I pull Search Console data?
Once a day, for one day of data. Google recommends this pattern because it stays inside quota, and data is typically available after two to three days, so target a date about three days back in Pacific Time.
Can the Search Console API request indexing for a URL?
No. The URL Inspection endpoint only reports index status. The Indexing API can notify Google about pages, but Google documents it for job posting and livestream video pages only, so it is not a supported way to push ordinary pages.
Should I use the API or the BigQuery bulk export?
Use the API for URL inspection, sitemaps and lightweight daily pulls into your own database. Use the BigQuery export when you need the long tail beyond 50,000 rows a day or want to join Search Console with other data. The export is not backfilled, so enable it early.