AI image tool

Image Caption Generator

Upload a photo and get a caption written for it — accessibility alt text, a social post, ad copy, or a plain description. Powered by OpenAI’s vision models.

Caption an image

JPEG, PNG, GIF or WebP. Large photos are shrunk in your browser before upload, so nothing oversized leaves your device.

Drop an image here, or click to choose JPEG, PNG, GIF or WebP — up to 8 MB
Preview of the image you selected
Caption style

The four styles

Same image, four different jobs. Pick the one that matches where the text is going.

StyleWhat you getUse it for
Alt textOne sentence under 125 characters, describing what matters visually.The alt attribute on a web image — screen readers and SEO.
Social captionOne or two engaging sentences, plus 3–5 hashtags.Instagram, Facebook, LinkedIn posts.
Ad copyA headline, a benefit-led body line, a call to action, and hashtags.Paid social and display creative.
Plain descriptionTwo or three factual sentences with no marketing language.Cataloguing, image libraries, internal notes.

Frequently asked questions

How the caption generator works.

What happens to the images I upload?

Your image is sent to the OpenAI API to be read, and the caption comes straight back to your browser. It is not saved to this website, and no copy is kept on the server after the caption is generated.

Why is my photo resized before uploading?

Photos straight off a phone are often far larger than the model needs. Your browser scales anything over 1568 pixels on its long edge down to that size first, which makes the upload faster and the request cheaper without any loss in caption quality.

Is there a usage limit?

Yes. Each visitor can generate a limited number of captions per hour. Every caption costs the site owner money in API usage, so the cap keeps that predictable.

Can I edit the caption?

Yes — the result box is editable. Tweak the wording, then use Copy. Treat every caption as a first draft and check it against the image before publishing.

Which file types work?

JPEG, PNG, GIF and WebP — the four formats the vision model accepts. The file’s real contents are checked on the server, not just its extension.