File Handling
Learning Outcomes
- Identify which file types and size limits each major AI app supports
- Upload PDFs and extract clean, structured summaries and data
- Analyse images and screenshots with precise visual prompts
- Process spreadsheets and code files for analysis and review
- Protect sensitive data when deciding what to upload and where
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Files are where these tools earn their keep |
| Limits | 6 min | What each platform supports, side by side |
| PDFs | 7 min | Summarising and extracting from documents |
| Images | 7 min | Screenshots, photos, diagrams, charts |
| Spreadsheets | 7 min | CSV and XLSX analysis without a single formula |
| Code & privacy | 8 min | Reviewing code, combining files, protecting data |
| Wrap-up | 2 min | Key takeaways, preview next lesson |
Before You Begin
Pre-work:
- Be comfortable in at least two apps from earlier lessons — Claude Projects & Artifacts, ChatGPT Canvas & Code Interpreter, and Gemini Across Workspace; skim Comparing Tools too
- Have a paid tier on at least one app — free tiers limit uploads sharply
Shopping List:
- A sample PDF (report, invoice, or paper of 5-30 pages)
- A screenshot or photo containing text (chart, receipt, or error)
- A spreadsheet (CSV or XLSX) with a few hundred rows
- A code file or small folder you are happy to share
Check the ceiling before you drag a 90 MB scan into the box. Here is the comparison as of mid-2026 — treat the numbers as a snapshot, since vendors move them.
| Capability | Claude | ChatGPT | Gemini |
|---|---|---|---|
| Documents | PDF, DOCX, CSV, TXT, HTML, RTF, EPUB | PDF, DOCX, TXT, CSV, XLSX, ZIP | DOC, DOCX, PDF, RTF, most types |
| Images | JPEG, PNG, GIF, WebP | PNG, JPG | Photos, common formats |
| Per-file size | Up to 500 MB | Large cap on paid tiers | 100 MB; video 2 GB |
| Files at once | Up to 20 per chat | Several per message | Up to 10 per prompt |
Two limits matter more than file size. Image dimensions: Claude caps images at 8000 x 8000 pixels, so a giant screenshot can be rejected even at a few megabytes. The context window: every file becomes tokens that share space with your conversation, and a dense PDF can fill it fast. When answers go vague late in a chat, you have run out of room, not hit a bug.
File Uploads FAQ in the OpenAI Help Center, not a blog post. Paid tiers raise file size, upload counts, and context, often by a large multiple.The PDF is the workhorse of file handling. Click the paperclip or plus icon at the left of the message box, pick your file, wait for the thumbnail, then write your prompt with the file attached. A vague prompt wastes the upload — compare What's in this? with:
> This is a quarterly board report. Summarise it in five bullet points an
> executive can act on, leading each with its most important number. Ignore
> the appendices, and if a figure is missing, say so rather than guess.
The second names the document type, audience, format, scope, and how to handle gaps — so you get something usable on the first try. For data extraction, demand a format and tell it to write "not found" for absent fields rather than infer.
Platform-wise: Claude reads text and visual layout (charts, stamps) for PDFs under roughly 100 pages, text-only beyond; ChatGPT can pull tables via Code Interpreter (see Lesson 4).
All three apps are multimodal — they genuinely see images. That unlocks reading a chart you cannot copy, transcribing a whiteboard, or debugging an error screenshot. Upload as you would a document, then be specific:
> This screenshot is a line chart with no underlying data. Read the values
> off the axes and rebuild them as a Markdown table with columns Month and
> Revenue. Flag any point you are unsure about.
For text trapped in an image — a receipt, slide, or sign — ask for transcription explicitly: "Transcribe all text exactly as written; do not correct or translate." For dense charts, Claude and Gemini reason most reliably; for a photo plus live web context, Gemini and ChatGPT combine vision and browsing.
Here non-programmers get the biggest lift: analyse a spreadsheet without writing a formula, just by describing the question in plain English. Two modes exist. Read-as-text has the model reason over the rows — fine for small files, but it falls apart on large datasets that exceed the context window. Compute-on-it runs actual code against the file (ChatGPT's Code Interpreter, Gemini's data analysis), accurate because it calculates rather than estimates, and it draws charts. Push the app toward computing:
> Here is a CSV of 12 months of sales. Calculate total revenue, the top
> five products, and month-over-month growth, then chart monthly revenue.
> Show me the steps you took, not just the answer.
One Claude note: it accepts XLSX, but reading the cells requires code execution enabled on your account. CSV is the safest default everywhere; if a multi-tab XLSX misbehaves, export the tab you need.
show me the steps you took) so you can sanity-check the logic.You do not need to be a programmer to get value from uploading code — the AI is the expert, not you. Upload a script, config, or component (or paste it if short), then aim the prompt at what you want to know:
> This is a Python script a colleague sent me. In plain English, explain
> what it does and what I need to run it. Flag anything risky: hard-coded
> keys or passwords, or anything that deletes data. I am not a developer.
Platform-wise: ChatGPT can upload a ZIP or run code through Code Interpreter, and Gemini connects to a code folder or GitHub repository up to 5,000 files within 100 MB.
Real work rarely involves one tidy file. Stitch the steps above into one task — a market brief from a PDF, a pricing CSV, and a screenshot of a rival's page, reconciled with each source cited. Keep it in one conversation so every file stays in context; for ongoing work, Projects (Claude) and Custom GPTs (ChatGPT) hold the files persistently — the next lesson.
The flip side is privacy. Every file is processed on a vendor's servers, so before you drag one in, ask: would I be comfortable if this leaked? The rules changed recently: on Claude's consumer plans, data training is a toggle you can turn off — check it. Anthropic states its commercial products (Claude for Work, the API, Enterprise) are not trained on by default. Across vendors, business tiers protect by default; consumer tiers do not.
| File contains | Where it belongs |
|---|---|
| Public report, your own draft | Any app |
| Internal company doc | Business tier with no-training terms |
| Others' personal data (PII) | Redact first, or an enterprise tool |
| Passwords, keys, credentials | Do not upload — ever |
| Regulated data (health, legal) | Only on a contracted, compliant plan |
Questions & Answers
Key Takeaways
- Check limits first — file size, image dimensions, and the context window are separate ceilings; the context window bites first on dense documents.
- Specific prompts unlock files — name the document type, audience, format, and how to handle gaps for a usable first-try answer.
- Compute, don't estimate — push spreadsheet work toward code so the app calculates rather than guesses, and show its working.
- Crop and clean inputs — a tight image crop and a flat, single-header spreadsheet beat clever prompting on messy files.
- Verify extracted facts — make the model cite the page, cell, or line, then spot-check; confident-but-wrong is the real risk.
- Privacy is a tier decision — never upload credentials or regulated data to a consumer chat; for confidential work, use a business plan with no-training terms.
Next Steps: Lesson 9: Building Custom GPTs & Projects