Describe the bug
JATSParser isn't handling some HTML entities correctly, including some in the Turkish alphabet (e.g. İ). Using the pure xml parser lxml-xml, the entities disappear; using flexible xml parser lxml (NLM-mode) doesn't convert them properly.
To Reproduce
Use JATS parser in nlm mode to parse the file .../fulltext/sources/IOJVT/0006/11052664_ft.xml. You should see that the entity İ remains in the affiliation text, whereas some other entities have clearly been converted to unicode.
Additional context
Testing shows that if you pass incoming raw text through html.unescape(rawtext), everything gets correctly converted to Unicode
Describe the bug
JATSParser isn't handling some HTML entities correctly, including some in the Turkish alphabet (e.g. İ). Using the pure xml parser
lxml-xml, the entities disappear; using flexible xml parserlxml(NLM-mode) doesn't convert them properly.To Reproduce
Use JATS parser in
nlmmode to parse the file.../fulltext/sources/IOJVT/0006/11052664_ft.xml. You should see that the entityİremains in the affiliation text, whereas some other entities have clearly been converted to unicode.Additional context
Testing shows that if you pass incoming raw text through
html.unescape(rawtext), everything gets correctly converted to Unicode