Home › Guides › Google Images Scraper
Google Images data: original image URLs, sizes, source pages, products and licenses
The Google Images Scraper returns what Google Images shows for a search term, one row per image: the full-size image URL (Google's link to the original file, not its thumbnail) with its width and height, file size, format and dominant color, the page it appears on with its title, site name and domain, and — when Google has them — the product on the page (title, brand, price, currency) and the image's license details. Google's own filters work in 51 countries and the interface language you choose. No Google account, no API key, no browser — requests go through Apify's Google SERP proxy.
Who it's for
- Machine-learning teams building image datasets: many unique original images per term, with size, format and file size, filterable by size, type and color.
- Image SEO and brand monitoring: which pages and domains rank in Google Images for your terms, by country, and at which position.
- E-commerce research: which sites show a product's images, and the product, brand and price Google attaches to them.
- Anyone looking for reusable images: Creative Commons results with the license name, creators, credit line and license link Google shows.
- Keyword research: Google's related searches for each term.
What you get
| Group | Fields |
|---|---|
| Image | Original image URL, a crawler-only flag (Instagram, Facebook and TikTok links that return a web page or an error instead of the file), width, height, aspect ratio, orientation, format, file size in bytes and as text, dominant color (RGB and hex), thumbnail URL and size |
| Source | Title, the page the image appears on, its domain and site name, Google's "About this image" link |
| Product | Product flag, title, brand, price, currency, price text, description, product ID, Google Shopping link |
| License | Licensable flag, license link, Creative Commons license name ("CC BY-NC 2.0", "CC0 1.0"…), creators, credit line, copyright notice |
| Context | Your term, the Google search that returned the image (your term or a related search), position, results page, Google's image ID (unique per term), country, language, filters, the Google page it came from, scrape time |
Product and license fields are filled only when Google shows that information for the image. A real product row from a test run on 2026-09-30, search nike air max 90 (trimmed):
{
"type": "image",
"query": "nike air max 90",
"position": 8,
"title": "Nike Air Max 90 (Green Bay Packers) Men's NFL Shoes - Vintage Green/University Gold/Black/Alabaster",
"imageUrl": "https://static.nike.com/a/images/t_web_pdp_936_v2/f_auto,u_9ddf04c7-2a9a-4d76-add1-d15af8f0263d,c_scale,fl_relative,w_1.0,h_1.0,fl_layer_apply/357105d2-f234-4916-807a-b9e312aa5bb1/NFL+RIVALRY+AIR+MAX+90+PACKERS.png",
"imageUrlCrawlerOnly": false,
"imageWidth": 1872,
"imageHeight": 2340,
"format": "png",
"fileSizeText": "266KB",
"pageUrl": "https://www.nike.com/t/air-max-90-green-bay-packers-mens-nfl-shoes-YXEcR3is",
"domain": "www.nike.com",
"siteName": "Nike",
"isProduct": true,
"productBrand": "Nike",
"productPrice": 140,
"productCurrency": "USD",
"productPriceText": "$140.00"
}
And the license part of a row for golden retriever from the same day:
{
"imageWidth": 1000,
"imageHeight": 667,
"fileSizeBytes": 44032,
"dominantColorHex": "#58554b",
"domain": "a-z-animals.com",
"isLicensable": true,
"licenseUrl": "https://www.shutterstock.com/license?utm_source=iptc&utm_medium=googleimages&utm_campaign=webstatement",
"licenseCreators": ["ConstanzaMartinez"],
"licenseCredit": "Shutterstock",
"licenseCopyright": "Copyright (c) 2023 ConstanzaMartinez/Shutterstock. No use without permission."
}
How it works
- Google's own pages first. A Google Images results page has up to 100 images (sometimes 50). The scraper reads the term's pages until it has enough images or Google has no more. How many Google has varies: "golden retriever" stopped after 237 new images (3 pages), while "mountain landscape" still had more after 400.
- Then related searches, only on your subject. When the term's own pages run out before your limit, the scraper continues with the related searches Google lists — only those that contain every word of your term, such as "golden retriever puppy". Rows found this way say so, and name the search that found them.
- The data comes from Google's page, not from the images. URL, size, file size, color, title, site, product and license details are what Google publishes with each result; the scraper does not download the images.
- No duplicates within a term. Images are de-duplicated by Google's image ID; an option delivers each image once per run across terms.
Google's filters, checked against Google
Size, color, type, file type, time and usage-rights filters are Google's own. Every value was tried against Google on 2026-09-30, most of them next to the same search without the filter:
imageSize: large→ 94% of images at least 1,000 px wide (57% without the filter).color: blue→ 77% blue-dominant images (11% without).fileType: svg→ 93%.svgfiles.usageRights: creativeCommons→ all 150 images in two runs carried a creativecommons.org link, 149 of them to a specific license — from CC0 and public-domain marks to CC BY, BY-SA and BY-NC, so check the license name before reusing an image commercially.
Aspect ratio is computed by the scraper from the image size, because Google ignores its own aspect-ratio parameter. Minimum width and height, and leaving out crawler-only links, are applied by the scraper too.
Input
| Input | What it does |
|---|---|
queries · maxImagesPerQuery | Search terms, one per line · images per term, up to 1,000 (default 100) |
relatedSearches · relatedSearchesExclude | Continue with Google's related searches when the term runs out (default) or not · words that keep a related search out ("drawing", "anime") |
country · language | Which Google to search (51 countries, each on its own Google domain) · interface language |
imageSize, color, imageType, fileType, time, usageRights | Google's own filters |
aspectRatio, minWidth, minHeight, skipCrawlerOnlyImages | Filters applied by the scraper, which reads more pages to replace the images it leaves out |
removeDuplicatesAcrossQueries · outputRelatedSearches | Each image once per run · one free row per term with Google's related searches |
A training set of large photos, at least 1,200 px wide, without crawler-only links and without related searches that change the subject:
{ "queries": ["red panda"], "maxImagesPerQuery": 500, "imageSize": "large", "imageType": "photo", "minWidth": 1200, "skipCrawlerOnlyImages": true, "relatedSearchesExclude": ["drawing", "anime", "cartoon", "turning red"] }
In a platform run (338 s, 13 Google searches) this gave 500 unique images, all at least 1,200 px wide and none crawler-only; 107 crawler-only and 115 narrower images were left out.
Images you can reuse under Creative Commons licenses:
{ "queries": ["mountain landscape"], "usageRights": "creativeCommons", "maxImagesPerQuery": 100 }
Measured data quality
Share of images with a value in four platform runs:
| Field | golden retriever (700) | nike air max 90 (200) | kedi, Turkey (200) | 3 terms × 400 |
|---|---|---|---|---|
| Image URL, width, height, page URL, title, domain, site name, file size, dominant color, thumbnail | 100% | 100% | 100% | 100% |
| Format (from the file name) | 82% | 78% | 89% | 85% |
| Product brand / price | 2% / 2% | 25% / 9% | 1% / 1% | 7% / 3% |
| License link | 9% | 0% | 4% | 10% |
| Unique images per term | 700 / 700 | 200 / 200 | 200 / 200 | 1,200 / 1,200 |
Checked against the files themselves: for 40 sampled rows from 39 sites, a QA script downloaded the linked image and read its real pixel size. 34 links returned an image file, and all 34 had exactly the listed width and height; the byte size was within 5% of the listed file size for 29 of the 31 files with a known size.
Honest limits
- Some links only work for crawlers. In earlier test runs, 369 of 4,670 image rows (7.9%) had an Instagram, Facebook or TikTok crawler link as the image URL; in checks, such links returned a web page or HTTP 403 instead of the file. They are flagged with
imageUrlCrawlerOnly, andskipCrawlerOnlyImagesleaves them out and fills the limit with other images. - A related search can change the subject. For "red panda", Google suggested "red panda drawings", "anime red panda" and "turning red red panda" (a film); without exclusions those searches brought 141 and 168 of 500 rows in two runs. Put such words in
relatedSearchesExclude, or drop rows by the search that found them. - Product and license data depend on the term. A product search had a brand on 25% and a price on 9% of rows; other terms had far fewer. Product prices are in the currency Google shows for that listing, which can differ from your country.
- No date per image. Google's time filter changes the result set, but Google shows no date per image, so its accuracy is not measured.
- Format comes from the file name. It is empty when the URL has no image extension, and the extension matched the real file type in all but 2 of 34 downloaded files.
- Country matters. For
kedi, the first 100 images in Turkey and in the US shared only 20.
Speed
On the Apify platform, 50 images took 11 seconds, 100 Creative Commons images 16 seconds, 200 product images 150 seconds, and 3 terms × 400 images (1,200 in all) 269 seconds. A 700-image run with related searches took 221–248 seconds. Run time depends mostly on how fast Google answers through the proxy.
Pricing
Pay per result: one image event per delivered image row. Status rows and related-searches rows are free, and images left out by the scraper's own filters are not delivered, so they are not charged. Current prices, including any per-run start event, are on the actor's Pricing tab on Apify.
Pull a term's images now: original URLs with size, source page and site — plus product and license details where Google shows them — exportable to JSON, CSV or Excel.
Open the Google Images Scraper →Related
To see which of your terms are gaining search interest, pair this with the Google Trends scraper guide. For e-commerce research beyond images, the e-commerce & app intelligence guide covers detecting which domains run Shopify and pulling their catalogue. All Google datasets are in the Google data section.
Frequently asked questions
Are these the original images?
imageUrl is Google's link to the original file on the site that hosts it, with the width and height Google reports. Of 40 sampled links, 34 returned an image file and all 34 had exactly the listed size; 4 sites refused the download and 2 were Instagram crawler links. thumbnailUrl is Google's small copy.
Why do I get fewer images than I asked for?
Google stops after a few pages for one term. With related searches on (the default), the scraper continues with related searches that contain every word of your term; when those run out, or 3 in a row bring no image that passes your filters, the term ends. SOURCE_REPORT shows each search made and the reason.
Can I use the images?
Check the license on the source page. The usage-rights filter and the license fields show what sites publish; they are not legal advice. Images belong to their owners.
Does the scraper download the images?
No. It returns the URLs and the information Google publishes about each image; download the files yourself if you need them.
How do I get only my term's own results?
Set relatedSearches to off, or keep only rows with fromRelatedSearch: false.
Google Images Scraper — pay per image, Google's own filters, no Google account.
Get Google Images data on Apify →