Persée (UAR 3602) – Use Case 2

Linking older and newer knowledge by mass-producing citation metadata

About the organisation

Persée is an academic research and support unit, affiliated to ENS de Lyon, CNRS, and supported by the French Ministry of Higher Education, Research and Space. Since 2003, Persée has been a leading player for the mass digitisation of French-speaking scientific literature and the broad dissemination as high-quality digital collections and data. The Persée portal, aimed at an academic audience, contains over 1,100,000 OA documents and 660 serial titles, mainly journals and books, dating from the late 19th century to the present day. Persée also provides several “perséides” (project-specific online corpora), a triplestore, and various data services, supporting Open Science.

As a massive digitisation platform with an end-to-end production chain and further automation and processing projects, Persée collaborates with over 200 partners from both the public and private sectors such as publishers and libraries, bibliodata and research data service providers, researchers and research teams, AI engineering professionals.

Since 2003, Persée has been working with editorial boards of public and commercial scientific publishers. Persée services ensure digital continuity between current publications and OA backfiles, copyright clearance and publishers’ heritage collections long-term preservation. Persée is also interoperable with other public and commercial electronic publication platforms such as Cairn, OpenEdition journals, Erudit.

Persée also plays a significant role in CollEx-Persée, the French research infrastructure that brings together academic and research libraries with French scientific information actors.

What challenges can the GRAPHIA project help solve?

The massive, error-free production of citation metadata from full-text digitised publications remains a challenging task. How can we produce citation metadata at scale while adhering to absolute standards of quality and reliability for both metadata and enrichments? Additionally, at a time when knowledge graphs serve as hubs for data aggregation and propagation, where should citation data be created, further processed, stored, and disseminated? Could GRAPHIA be a solution for citation data production? Finally, the requirements for a sustained and continuous citation metadata production are inseparable from the pursuit of robustness and sustainable processes.

What is the proposed use case?

Persée produces about 450 000 pages yearly and largely disseminates corresponding metadata into the bibliodata ecosystem and towards bibliodata stakeholders, via Crossref for example. However, for lack of reliable large-scale processes, the citation metadata stock remains incomplete. To support Open Science, it is essential to better, more comprehensively, more openly link digitised bibliographies with born-digital bibliographies. In the context of Persée digital library, the use case includes additional challenges: French-speaking content, SSH citation practices, older citation practices, OCRed textual data (OCR noise). For Persée readers, users and beyond, more citation metadata, later findable through the GRAPHIA knowledge graph or other means, would mean simplified bibliographies, greater reliability, and more comprehensive search results. Additionally, this could help readers and researchers in diachronic research and studies, building new bridges between older and more recent knowledge.

Several actions are proposed to better understand how the GRAPHIA project can support this use case:

  • Participation in model training, post-training: providing data and ground truth data, use cases, competency questions;
  • Validation, control, and value based, business-oriented quality assessment by data and domain experts;
  • Assessing, characterising, benchmarking process performance, cost and sustainability with current production operations and ongoing R&D projects;
  • Studying, prototyping resulting citation data I/O flows.

What technical aspect would be used?

Persée is seeking to develop massive detection and linking of bibliographic references and bibliographic elements, either from scratch, or by better exploiting existing elements or with mixed approaches. That is why this use case could involve the GRAPHIA SSH Citations Index and the capabilities of the GRAPHIA LLM4SSH.

Back to Industry Hub

Contact

Contact our GRAPHIA communication team via the email address:

contact@graphia-ssh.eu

Sign up for the GRAPHIA newsletter here.

Follow us on LinkedIn, Bluesky and YouTube.