Skip to main content

📔 Lesson 2.1: The Source Landscape — Docs, PDFs, Slides, Text & Copied Text

In Module 1 you built your first notebook and asked a grounded question. Now we go deep on the single most important skill in the whole tool: getting good material in. This lesson covers the "upload" family — the files and text you feed NotebookLM directly. Master these five and you'll have a notebook that actually knows what it's talking about.

📚 What You'll Learn

By the end of this lesson, you will be able to:

  • Add the five "upload family" source types: PDFs, Google Docs, Google Slides, plain text/Markdown, and pasted copied text
  • Explain what NotebookLM actually does with each — it works from the readable text it can extract
  • Understand why scanned PDFs and image-only files can come in thin, and how to check
  • Decide when to paste text versus when to upload the whole file
  • Remember that sources are copies — editing the original later won't update the notebook

⏱️ Estimated Time: 40 minutes

🎯 Project: Add at least three different source types to your notebook and confirm each one processed successfully.

In This Lesson

The Upload Family — A Map

Think of NotebookLM as a brilliant research assistant who can only read text. You can hand this assistant almost anything — a PDF report, a slide deck, a Google Doc, a plain .txt file, or a paragraph you copied from somewhere — and the assistant's very first job is always the same: pull the readable words out and study them. Everything the notebook later does — the grounded answers, the citations, the Audio Overview, the mind map — is built on those extracted words.

That single fact explains almost everything about which sources work brilliantly and which come in disappointingly thin. A crisp, text-based document is a feast. A photo of a document — a scan with no real text underneath, just pixels — is a locked box the assistant has to squint at. Hold that image and the rest of this lesson will feel obvious.

In this course we split all sources into two families. The upload family — this lesson — is the material you feed in directly: files on your computer and text you paste. The link family — websites, YouTube, audio, Google Drive — is next lesson (2.2). Both end up in the same place, the Sources panel on the left, but they behave a little differently, so we learn them separately.

graph LR A["📄 PDF"] --> Z["📝 Readable text extracted"] B["📘 Google Doc"] --> Z C["📊 Google Slides"] --> Z D["📃 Text / Markdown"] --> Z E["✂️ Pasted text"] --> Z Z --> Y["🤖 NotebookLM studies the words"] Y --> X["💬 Grounded answers with citations"]

🧠 Mindset

You are not "uploading files" so much as handing over readable words. Every time you add a source, quietly ask yourself: "Is the real content here as text, or is it trapped in a picture?" That one habit will save you from the most common beginner frustration — a source that technically uploaded but that the notebook can barely read.

PDFs — The Workhorse (and the OCR Catch)

The PDF is the workhorse of NotebookLM. Research papers, reports, manuals, ebooks, government documents, lecture handouts — most of the serious reading in the world ships as a PDF, and NotebookLM handles the good ones beautifully. To add one, open your notebook, use the Add source button (in the current interface it opens a panel with an upload option and a drag-and-drop area), and either drop the file in or browse to it. Google may move or rename this button — hold onto the idea ("add a source, choose the file") rather than the exact pixels.

📄 What makes a PDF a great source

PDFs come in two flavors that look identical to your eye but are worlds apart to NotebookLM. A text-based PDF was created from a real document — exported from Word, Google Docs, LaTeX, or a "Save as PDF" — so the letters are actual selectable text living inside the file. Try this quick test: open the PDF and try to select a sentence with your cursor. If the text highlights, NotebookLM will read it cleanly. This is the ideal source.

🖼️ The catch: scanned and image-only PDFs

The other flavor is a scanned PDF — someone photographed or scanned paper pages, so each "page" is really just an image of text. Your eyes read it fine, but there may be no actual text underneath the pixels. If your select-a-sentence test fails (nothing highlights), the file is image-only.

Here's the honest, hedged truth: NotebookLM has become quite capable at OCR — Optical Character Recognition, the technology that reads letters out of images — and it can often extract text from a clean scan. But results vary with scan quality: crisp, straight, high-contrast scans usually come through well; faded, skewed, handwritten, or low-resolution ones may come in partial or garbled, or barely at all. The capability keeps improving, so don't assume the worst — but do verify.

⚠️ Watch Out — images and figures

Even in a perfect text-based PDF, remember that NotebookLM works primarily from text. Charts, diagrams, and photos may be understood loosely, partially, or not at all, and captions or labels embedded inside a figure image can be missed. If a crucial number lives only inside a chart image, consider typing it into a short text note as its own source. Don't assume the notebook "saw" a graph the way you did — ask it, and check the citation.

✅ Pro Tip — the 30-second confirmation

After any PDF finishes processing, don't just trust the green check. Open the chat and ask something that could only be answered if the text came through — "Summarize the main argument of the [source name] PDF" or "What does page 3 say?" If you get a specific, grounded answer with a working citation, you're golden. If you get a vague non-answer, the text may not have extracted — and you just caught it early instead of halfway through your project.

Google Docs & Slides

Because NotebookLM is a Google product, it has a natural affinity for Google's own formats. When you add a source, alongside file upload you'll find options to bring in a Google Doc or Google Slides deck straight from your Drive (we'll cover the full Drive picker in Lesson 2.2). These are born-digital, text-native formats, so they import cleanly — no OCR guessing games.

📘 Google Docs

A Google Doc is about as friendly a source as exists: it's structured, selectable text, often with clear headings NotebookLM can use to navigate. Meeting notes, essays, project briefs, research write-ups — if it lives in Docs, it comes in beautifully. This is a great reason to draft your own summaries in Docs and feed them back in as sources, a trick we return to in Module 3.

📊 Google Slides

Slides are a little more nuanced. NotebookLM reads the text on your slides — titles, bullet points, speaker text in boxes — which is often exactly the skeleton of an argument or a course. But be aware of two things. First, anything that lives only as an image on a slide (a screenshot, a diagram with no real text, a photo) faces the same limits as any image. Second, the speaker notes beneath each slide — where the real substance often hides — may or may not be pulled in depending on the current behavior, so don't assume they came along. If your notes matter, consider copying them into a Doc or a text source as well.

Source type How NotebookLM reads it Watch for
Text-based PDF Extracts the real text — excellent Figures/charts understood loosely
Scanned PDF OCR reads text from the image — quality varies Faded/skewed/handwritten scans may come in thin
Google Doc Clean, structured text — excellent Embedded images are still images
Google Slides Reads on-slide text Image-only slides & speaker notes may be missed
Text / Markdown Reads it verbatim — perfectly reliable You must have the text to begin with

Plain Text, Markdown & Pasted Text

The humblest sources are the most reliable of all. A plain text file (.txt) or a Markdown file (.md) is nothing but readable text, so there is no extraction to get wrong — NotebookLM reads every character exactly as written. If you have material as plain text, it will never surprise you.

✂️ Pasting copied text directly

You don't even need a file. NotebookLM lets you paste copied text as a source: when you add a source, look for a "copied text" or "paste text" option, drop your text into the box, give it a name, and it becomes a source just like any file. This is one of the most underrated features in the whole tool, and it's the answer to a surprising number of "how do I get this in?" problems.

🤔 When to paste versus upload

Here's a simple rule of thumb. Upload the file when you want the whole document and it's already a clean file — a full PDF report, a complete Doc. Paste the text when you only want part of something, when the source has no tidy file (an email, a forum post, a chunk of a web page that won't import well), or when you need to rescue content from a stubborn source — for example, a scanned PDF that isn't extracting: open it, select and copy the text your own reader can get, and paste that in as a clean text source.

💡 The paste-to-rescue move

Pasting is your universal fallback. Any time a "smart" source type disappoints — a page that won't import, a PDF that scans thin, a slide whose notes got dropped — you can almost always select the real words with your own eyes, copy them, and paste them in as clean text. It's a little manual, but it's bulletproof: what you paste is exactly what NotebookLM reads. Keep this trick in your back pocket for the whole course.

Every source is really just text in the end. The fancy formats are conveniences; the paste box is the plain truth underneath them all.

⚠️ A note on formatting

When you paste, you're usually pasting the words, not the layout. Tables can come in as run-together text, and complex formatting may flatten. For most reading that's completely fine — the meaning survives. But if a table's structure is load-bearing (rows and columns that must line up), a clean file upload, or a quick tidy-up of the pasted text, will serve you better than a mangled paste.

Sources Are Copies — What That Means

Here is a subtle point that trips up nearly everyone eventually, so let's make it stick now. When you add a source to NotebookLM — whether you upload a PDF, import a Doc, or paste text — NotebookLM takes a snapshot copy of that material as it was at that moment. The notebook now studies its own copy. It is not a live, linked, always-syncing connection to your original file.

The practical consequence: if you edit the original document later, the source in your notebook does not automatically update. Fix a typo in your Google Doc, add three pages to the PDF, revise your slides — the notebook is still working from the older snapshot it captured. This is true even for Google Docs and Slides, which feel live because they live in Drive, but which NotebookLM read at import time. (Behavior can evolve, and there may be a way to re-sync or re-import a source in the current interface — but the safe mental model is "the source is a copy, refresh it yourself when the original changes.")

graph TD A["📄 Your original document"] --> B["➕ Add to NotebookLM"] B --> C["📸 Snapshot copy stored in the notebook"] C --> D["🤖 Notebook studies the copy"] A --> E["✏️ You edit the original later"] E --> F["⚠️ Copy is now out of date"] F --> G["🔁 Re-add or re-import to refresh"]

✅ Two things this makes easy

  • You can't break your originals. NotebookLM works on copies, so nothing you do in a notebook touches or deletes your real files. Experiment freely.
  • Keeping a source fresh is on you. When a document changes in a way that matters, re-add the updated version (and remove the stale one, so the notebook isn't answering from two versions at once). We build good source hygiene into Lesson 2.3.

🎯 Project: Add Three Source Types

Time to make the notebook you started in Module 1 genuinely useful — and to feel the differences between source types with your own hands. The goal isn't quantity; it's variety and confirmation. By the end you'll have added at least three different kinds of source and personally verified that each one processed.

🏋️ Build out your source set

Objective: Add three or more different source types to your notebook and confirm each one is readable, using the 30-second confirmation trick.

Instructions (about 20 minutes):

  1. (5 min) Add a PDF related to your notebook's topic. Before adding, do the select-a-sentence test to note whether it's text-based or scanned.
  2. (4 min) Add a second type — a Google Doc or Slides deck, or a plain text/Markdown file. If you don't have one handy, jot a few paragraphs of your own notes into a Doc or text file first.
  3. (4 min) Add a third via paste copied text — grab a relevant chunk from anywhere (an email, an article, your own notes), paste it in, and give it a clear name.
  4. (5 min) For each source, ask the chat one question that could only be answered from that source, and confirm you get a specific, cited answer.
  5. (2 min) Rename any vaguely-titled source to something you'll recognize later (we go deep on this in Lesson 2.3).
💡 Hint — good confirmation questions
For a PDF:      "Summarize the main point of [source name]."
For a Doc:      "What are the key takeaways in [source name]?"
For pasted text:"According to the pasted [topic] text, what is X?"

A good sign: a specific answer + a clickable citation.
A bad sign:  "I don't have information about that" for a
             source you know contains it -> the text may
             not have extracted. Try the paste-to-rescue move.

If a scanned PDF came in thin, this is the perfect moment to practice the paste-to-rescue trick: copy the text you can select and paste it as a clean fourth source.

✅ Project Completion Checklist

  • You added at least three different source types to one notebook
  • At least one is an uploaded file and at least one is pasted text
  • You ran the 30-second confirmation on each and got a specific, cited answer
  • You noted whether your PDF was text-based or scanned
  • Every source has a name you'll recognize next week

🎯 Quick Quiz

Question 1: You upload a scanned PDF — a photographed document — and NotebookLM gives vague, thin answers about it. What's the most likely cause?

Question 2: You add a Google Doc as a source, then later fix several typos in that Doc. What happens to the source in your notebook?

Best Practices for Uploading Sources

✅ Do's

  • Prefer text-based files. A born-digital PDF, Doc, or text file gives the cleanest, most complete extraction every time.
  • Confirm every source. Ask one question you know the answer to; a specific cited reply proves the text came through.
  • Keep the paste box in mind. When a fancy format disappoints, copying the real words and pasting them almost always works.

❌ Don'ts

  • Don't assume a scan read perfectly. OCR is good and improving, but faded or skewed scans can come in thin. Verify.
  • Don't expect charts to be "seen." NotebookLM works mainly from text; numbers trapped in a figure image may be missed.
  • Don't forget sources are snapshots. Edit the original all you like — the notebook won't know until you re-add it.

💡 Pro Tips

  • Draft your own summaries in a Google Doc and feed them back in — clean, structured, perfectly readable sources.
  • Name sources as you add them. "Q3 Sales Report 2025" beats "document (7).pdf" when you have fifteen of them.

📓 Learning Journal

Keep adding to the learning journal you started in Module 1 — a document, a note, or a page in your notebook. After this lesson, take a few minutes to write down:

  • Key concepts you learned
  • Techniques that clicked for you
  • Questions or confusion points to revisit
  • Ideas you want to try
  • Your progress and feelings about learning this — including any source that surprised you, good or bad

✍️ This lesson's prompt: Which of your sources came in cleanly, and did any come in thinner than you expected? What will you do differently — better files, the paste-to-rescue move, keeping an eye on scanned documents — now that you know NotebookLM reads text, not pictures?

📝 Lesson Summary

🎓 Key Takeaways

  • The upload family is PDFs, Google Docs, Google Slides, plain text/Markdown, and pasted copied text — everything you feed in directly.
  • NotebookLM works from readable text. Text-based files are ideal; scanned/image-only PDFs depend on OCR, whose quality varies — always verify.
  • Pasting copied text is your universal fallback: when a smart format disappoints, copy the real words and paste them in clean.
  • Sources are snapshot copies. Editing the original later doesn't update the notebook — re-add it to refresh, and know your originals are never harmed.

🎉 What You've Accomplished

You've moved from "I have a notebook" to "I can feed it well." You now know the five direct source types, what NotebookLM does with each, how to catch a source that didn't read cleanly, and the paste-to-rescue move that saves the day. That's the difference between a notebook that guesses and one that genuinely knows your material.

❓ Common Questions at This Stage

Can I upload a Word document or a PowerPoint file directly?

Supported file types shift over time, and NotebookLM has broadened what it accepts. PDF and Google's own formats are the safest bets. If a format isn't accepted in the current interface, the reliable workaround is to convert it (e.g. "Save as PDF", or open it in Google Docs/Slides) or simply paste the text — check Google's help center for the current list.

How big can a single source be?

Individual sources can be large — commonly cited as hundreds of thousands of words each — so a full book-length PDF is usually fine. The exact caps are approximate and evolving, and there's also a number-of-sources limit per notebook (we cover that in Lesson 2.3). Treat any specific figure as "at the time of writing" and check Google's page if you're near a limit.

Should I split a huge document into several sources?

Usually no need for size reasons alone — but splitting can help focus. If a giant PDF really contains three different topics, separate sources let you select just the relevant one in the chat. That's a curation decision, and it's exactly what Lesson 2.3 is about.

🔭 Looking Ahead

Next up is Lesson 2.2: Web, YouTube & Audio Sources (and Google Drive) — the "link family." You'll pull text from web pages, bring in YouTube videos via their transcripts, add audio files, use the "Discover sources" helper, and grab material straight from Google Drive — each with its own honest limits.

✅ Before the Next Lesson

  • Finish the project: three different source types added and each confirmed
  • Have a web page URL and a YouTube video (ideally one with captions) in mind to add next lesson
  • Write your Learning Journal entry for this lesson

📚 Additional Resources

🌟 Encouragement for the Journey

Getting good material in is quietly the whole game — and you just learned to do it well, with your eyes open to the gotchas. A notebook is only as smart as its sources, and yours are getting sharper. Next we open up the whole web and every video you've been meaning to watch. 📔