How to Extract Text from Images Using OCR in Google Keep

When you are at a conference and someone hands you a printed agenda, or when you are reading a physical book and want to save a brilliant quote, retyping the entire page into your phone is a massive waste of time. Instead of manually typing, you can use the hidden Optical Character Recognition (OCR) engine built directly into Google Keep. This powerful tool can instantly analyze a photograph of a physical document and perfectly extract all the written words, converting them into standard, editable digital text.

How to Capture the Initial Image

To use the OCR feature, you must first get the image into Google Keep. The easiest way to do this is using the mobile app on your smartphone.

  1. Open the Google Keep app.
  2. Look at the toolbar at the bottom of the screen. Tap the Camera icon on the far left.
  3. Select Take photo from the popup menu.
  4. Hold your phone steady over the physical document. Ensure the room is well-lit and the text is in sharp focus, then snap the picture.

Google Keep will immediately embed the photograph at the top of a brand-new note.

How to Extract the Text (Grab Image Text)

Once the image is inside the note, the extraction process takes exactly three seconds.

  1. Tap the photograph directly to open it in full-screen mode.
  2. In the top right corner of the screen, tap the three-dot menu icon (More).
  3. Select Grab image text from the drop-down list.

The moment you tap that button, Google’s advanced cloud AI analyzes the geometry of the image, identifies the letters, and extracts them. The full-screen image will close, and you will see the complete, perfectly formatted block of text dropped directly into the body of your note, just below the original photograph.

Best Practices for Perfect Transcription

While the AI is incredibly intelligent—it can even transcribe messy handwriting and cursive—its accuracy is heavily dependent on the quality of your photograph. To ensure the OCR engine does not make typos, ensure the paper is lying perfectly flat without any shadows cast across it. Furthermore, if you are photographing a restaurant menu with complex, multi-column layouts, the AI might jumble the prices into the wrong paragraphs. For complex documents, it is highly recommended to take multiple, zoomed-in photos of single paragraphs rather than trying to transcribe the entire page at once.

Get the best tech tips delivered straight to your inbox.

Join thousands of readers mastering Apple, Google, Microsoft, and Linux.