Blog

 5 minute read.

How Grobid uses Crossref metadata to enrich bibliographic information

Crossref members register around 37,000 records daily. All this metadata is openly shared with the global scholarly community through our free APIs. We receive over two billion calls to our API each month, signalling its use by hundreds of tools and services worldwide, supporting the discovery and assessment of scholarly metadata. As part of a new series in our blog, we will spotlight the diversity of tools and integrations that rely on Crossref metadata to serve the scholarly community. In this initial post, we chat to Luca Foppiano, the developer and maintainer of Grobid1, to learn about how Grobid uses Crossref metadata to improve the extraction and enrichment of bibliographic information from scholarly PDFs.

Extracting metadata from PDFs; easier said than done

If you’ve ever tried to extract bibliographic data from a PDF, you know this is not always straightforward. Copying author names, titles, and references by hand is tedious. Often, the layout includes invisible characters that someone has to clean, and even automated extraction often fails because the PDF format was conceived for reading documents across platforms; not for clean, structured metadata. Then, if you are considering doing this in bulk, with long lists of files, the problem grows exponentially. This is the core problem that Grobid was conceived to solve.

Enter Grobid

Grobid is an open-source software library, written by Patrice Lopez, that helps users extract and parse technical and academic text from PDF files. This means you can use a scholarly PDF as input, and Grobid will extract structured bibliographic metadata, including title, abstract, authors, affiliations, keywords, references, and more. Integration with Crossref’s metadata allows Grobid to further augment the quality of that metadata and improve the overall accuracy of the process.

Grobid is used by integration services like ResearchGate, Academia.edu, HAL Research Archive, the European Patent Office, The Institute of Scientific and Technical Information - French National Centre for Scientific Research (INIST-CNRS), Mendeley, CERN (Invenio), Internet Archive1, and research projects2.

How does Grobid use Crossref systems?

During the first quarter of 2026, our REST API received more than 100 million requests from Grobid software clients. Grobid queries the REST API and retrieves the metadata for the full reference, adding missing elements and improving the overall accuracy of the process. Additionally, this integration makes full use of Crossref’s different API pools2 by letting users make polite requests or include their Metadata Plus API keys in the software if they are already Metadata Plus subscribers3.

For example, an unstructured reference that appears at the end of a scholarly document might look as follows:

Rada, F, et al. “Osmotic and turgor relations of three mangrove ecosystem species.” Australian Journal of Plant Physiology, vol. 16, no. 6, 1 Dec. 1989, pp. 477–486, https://doi-org.ezproxy.csu.edu.au/10.1071/pp9890477.

As we can see, this specific citation style includes only the leading author. Other styles might, for example, use the journal abbreviated name or even skip it entirely. After being processed by Grobid, this reference is reorganized into a specific standard format called Text Encoding Initiative or TEI. Don’t be alarmed if it looks complex, these formats aren’t intended for humans to read, they’re intended for machines and interactions between software. What we want to show is how two lines describing references are tagged and structured (for which having a DOI with rich metadata definitely helps):

<biblStruct status="consolidated" source="crossref" xml:id="b1">
                        <analytic>
                            <title level="a" type="main">Osmotic and Turgor Relations of Three Mangrove Ecosystem Species</title>
                            <author>
                                <persName>
                                    <forename type="first">F</forename>
                                    <surname>Rada</surname>
                                </persName>
                            </author>
                            <author>
                                <persName>
                                    <forename type="first">G</forename>
                                    <surname>Goldstein</surname>
                                </persName>
                            </author>
                            <author>
                                <persName>
                                    <forename type="first">A</forename>
                                    <surname>Orozco</surname>
                                </persName>
                            </author>
                            <author>
                                <persName>
                                    <forename type="first">M</forename>
                                    <surname>Montilla</surname>
                                </persName>
                            </author>
                            <author>
                                <persName>
                                    <forename type="first">O</forename>
                                    <surname>Zabala</surname>
                                </persName>
                            </author>
                            <author>
                                <persName>
                                    <forename type="first">A</forename>
                                    <surname>Azocar</surname>
                                </persName>
                            </author>
                            <idno type="DOI">10.1071/pp9890477</idno>
                        </analytic>
                        <monogr>
                            <title level="j">Australian Journal of Plant Physiology</title>
                            <idno type="ISSN">0310-7841</idno>
                            <idno type="ISSNe">1446-5655</idno>
                            <imprint>
                                <biblScope unit="volume">16</biblScope>
                                <biblScope unit="issue">6</biblScope>
                                <biblScope unit="page" from="477" to="486"/>
                                <date type="published" when="1989-12-01"/>
                                <publisher>CSIRO Publishing</publisher>
                            </imprint>
                        </monogr>
                        <note type="raw_reference">Rada, F., Goldstein, G., Orozco, A., Montilla, M., Zabala,O. &amp; AzĂłcar, A. Relaciones OsmĂłticas</note>
                    </biblStruct>

In addition, there is another tool that complements API queries with Crossref bulk metadata options. The team behind Grobid also developed Biblio-glutton4, a utility “dedicated to scientific bibliographic information” that serves as a local, high-performance cache of Crossref (and other) scholarly metadata. Being local first means you can use it with the public data file5, deploy your own copy of the Crossref database, make updates as needed, and query as many times as you require.

Grobid has recently improved its integration with Crossref to make use of the REST API polite pool6 and added more metadata statements, such as the CRediT taxonomy and conflict-of-interest statements, which were requested by users.

Beyond bibliographic metadata

Another synergistic aspect of this Crossref integration is that users’ outputs can be enriched with funding metadata, identify mentions of datasets (data citation endpoint) and software, and through the use of modules, including other discipline-specific elements like astronomical entities, superconductor materials, and entities identified with Wikidata IDs.

Have you used Grobid with the Crossref API? We’d love to hear about your experience or use cases in the comments.

Crossref holds over 180 mln records contributed by more that 25,000 member organisations from across the world. We make all the metadata openly available via our REST API. If your tool can also benefit from integrating our open metadata, you can also explore our documentation and test different metadata retrieval options.

References


  1. Houten, S., Hendricks, G., & Polischuk, P. (2024). Rebalancing our REST API traffic. https://doi-org.ezproxy.csu.edu.au/10.64000/pq6bm-0qs92 ↩︎

  2. Besançon, L., Cabanac, G., LabbĂ©, C., Magazinov, A., di Scala, J., Tkaczyk, D., & Weber-Boer, K. (2025). Detection of metadata manipulations: Finding sneaked references in the scholarly literature (Version 1). arXiv. https://doi-org.ezproxy.csu.edu.au/10.48550/ARXIV.2501.03771 ↩︎

  3. Grobid documentation. https://grobid.readthedocs.io/ ↩︎

  4. Biblio-glutton. https://github.com/kermitt2/biblio-glutton ↩︎

  5. Rittman, M., del Ojo ElĂ­as, C., & Montilla, L. (2026). 2026 public data file now available. https://doi-org.ezproxy.csu.edu.au/10.64000/7s70g-drz77 ↩︎

  6. Rittman, M., & Montilla, L. (2025). Announcing changes to REST API rate limits. Crossref. https://doi-org.ezproxy.csu.edu.au/10.64000/wadve-3tj60 ↩︎