Skip to content

At least some HTML entities are handled incorrectly by JATSParser #181

Description

@seasidesparrow

Describe the bug
JATSParser isn't handling some HTML entities correctly, including some in the Turkish alphabet (e.g. İ). Using the pure xml parser lxml-xml, the entities disappear; using flexible xml parser lxml (NLM-mode) doesn't convert them properly.

To Reproduce
Use JATS parser in nlm mode to parse the file .../fulltext/sources/IOJVT/0006/11052664_ft.xml. You should see that the entity İ remains in the affiliation text, whereas some other entities have clearly been converted to unicode.

Additional context
Testing shows that if you pass incoming raw text through html.unescape(rawtext), everything gets correctly converted to Unicode

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions