Back to home

GeoGPT - Advanced Geological Document Intelligence

Sixty years of geological knowledge, usable for the first time.

Introduction

Extracting and organising data from geological survey documents at scale - since the 1960s.

The client was sitting on thousands of scanned PDFs from Norwegian regional surveys going back to the 1960s. The data inside them was valuable. The problem was that nobody could search it, use it, or even read it efficiently. By the time they came to us, the backlog would have taken years to process by hand. Four things made it harder than it sounds.

The Problem

  • Document Formats Kept Changing: Survey PDFs from the 1960s look completely different from ones made in the 2000s - formats shifted every ten to fifteen years. One extraction approach was never going to work across all eras. The system had to recognise which era a document came from and adjust accordingly.
  • Each Document Is Its Own Puzzle: A single geological PDF contains a cover page, an index, survey reports, drilling summaries, maps at different scales, and method-specific graphs. Every page type needs a different extraction approach - miss one and you lose data that might be critical for a researcher.
  • Norwegian + Highly Technical Content: The documents are in Norwegian, written with specialised geological terminology that standard language models were not built to handle. General-purpose OCR and AI tools got the words but missed the meaning - the domain knowledge had to be built in from scratch.
  • Data Hidden Inside Visuals: Survey documents use varied symbols, multiple drilling methods, and complex graphs that encode data visually - not as text. You cannot extract numbers from a graph with a text reader. The system needed to see what was on the page and interpret it correctly.

The Solution

We built an end-to-end AI pipeline that classifies each page, extracts the right data using the method that fits that page type, geo-references drilling point locations, and makes everything searchable through a web portal. Documents that required expert eyes and hours of reading are now a database anyone can search in seconds.

  • Every Page Gets Classified First: Before extraction, custom-trained classification models identify what type of page the system is dealing with - cover, map, graph, or drilling summary. The right method gets applied to each page automatically.
  • Text and Visual Extraction Together: Azure Form Recognition handles text extraction; GPT-4o Vision processes complex visual layouts and structured tables. Each page type has its own custom prompt built to pull out exactly the fields that matter.
  • Drilling Points Placed on a Real Map: For borplan pages, the system detects printed XY coordinates and uses a grid-based architecture to locate each drilling point precisely. When coordinates are absent, it falls back to street names or building references, then uses ArcGIS to align the site.
  • Graphs Become Numbers: Detection models find each graph, extract the grid, assign values to each point, and convert the image into actual numerical data. Data previously readable only by eye is now structured, stored, and fully queryable.
  • Everything in a Searchable Portal: The web portal provides full-text search across all documents, filters by person, location, and date, and an interactive map showing every drilling point from every survey. Users can view the original PDF alongside extracted text, with highlighting showing exactly where each piece of data appeared in the source document.

The Impact

60 yrs of survey data now searchable.

From years of manual backlog to hours. Historical documents that used to take weeks of manual review can now be processed in hours. Researchers can search across sixty years of Norwegian geological survey data in seconds - filtered by location, drilling method, date, or personnel.

  • Geologists Do Analysis, Not Admin: Manual processing time dropped significantly, freeing geologists to focus on analysis rather than document handling - the work they were actually hired to do.
  • Structured, Consistent Data: All extracted data is structured and consistent, making it straightforward to compare findings across different surveys and time periods - something previously impossible at scale.
  • Spatial View of All Drilling Activity: The interactive map gives organisations a complete spatial view of all drilling activity - speeding up site selection and infrastructure planning that used to require manual cross-referencing of dozens of documents.
  • Seconds, Not Weeks: Researchers can search across sixty years of survey data in seconds, filtered by location, drilling method, date, or personnel - with the source PDF alongside extracted text for full verification.
A Platform That Keeps Growing. New documents can be uploaded at any time, so the platform keeps expanding as new surveys come in. GeoGPT did not just digitise old documents - it made sixty years of geological knowledge usable for the first time, and keeps making new knowledge usable from day one.
Promact team

We are a family of Promactians

We are an excellence-driven company passionate about technology where people love what they do.

Get opportunities to co-create, connect and celebrate!

Join Us

Vadodara

Headquarter

B-301, Monalisa Business Center, Manjalpur, Vadodara, Gujarat, India - 390011

+91 (932)-703-1275

Pune

46 Downtown, 805+806, Pashan-Sus Link Road, Near Audi Showroom, Baner, Pune, Maharashtra, India - 411045

USA

4056, 1207 Delaware Ave, Wilmington, DE, United States America, US, 19806

+1 (765)-305-4030
Promact global office locations on world map