Skip to content

2.5.5 regression: strong emphasis adjacent to word/CJK chars fails when content starts or ends with punctuation #688

Description

@AndersonBY

Summary

After upgrading from markdown2 2.5.4 to 2.5.5, strong emphasis sometimes stops parsing when all of these are true:

  • the **...** span is adjacent to alphanumeric or CJK text without surrounding spaces
  • the emphasized content starts or ends with punctuation, for example Chinese quotes or parentheses

On 2.5.4 these cases render correctly. On 2.5.5 the parser leaves literal ** in the output, and in longer inputs it can also emit malformed mixed HTML.

This reproduces with the default parser configuration, no extras required.

Related to #679 because it also looks like an emphasis-regression in the recent parser changes, but this one reproduces without middle-word-em and with plain markdown2.markdown(...).

Minimal repro

import markdown2

cases = [
    'a**“b”**c',
    '**“b”**c',
    'a**(b)**c',
]

print('version:', markdown2.__version__)
for text in cases:
    print('INPUT :', text)
    print('OUTPUT:', markdown2.markdown(text).strip())
    print()

2.5.4

<p>a<strong>“b”</strong>c</p>
<p><strong>“b”</strong>c</p>
<p>a<strong>(b)</strong>c</p>

2.5.5

<p>a**“b”**c</p>
<p>**“b”**c</p>
<p>a**(b)**c</p>

Longer repro that produces malformed mixed HTML

import markdown2

text = '*   **示例**:系统会使用**(方案A)**或**(方案B)**进行处理。'
print(markdown2.markdown(text))

2.5.4

<ul>
<li><strong>示例</strong>:系统会使用<strong>(方案A)</strong>或<strong>(方案B)</strong>进行处理。</li>
</ul>

2.5.5

<ul>
<li><strong>示例</strong>:系统会使用**(方案A)<strong>或</strong>(方案B)**进行处理。</li>
</ul>

Suspected regression window

Looking at the source, Markdown._do_italics_and_bold() changed between these versions:

2.5.4

text = self._strong_re.sub(r"<strong>\\2</strong>", text)
text = self._em_re.sub(r"<em>\\2</em>", text)

2.5.5

if not self._iab_processor:
    self._iab_processor = GFMItalicAndBoldProcessor(self, None)
if self._iab_processor.test(text):
    text = self._iab_processor.run(text)

So this looks related to the switch to GFMItalicAndBoldProcessor as the default implementation for _do_italics_and_bold().

Environment

  • Python: 3.14.0
  • markdown2: 2.5.4 vs 2.5.5
  • OS: Windows 11

Activity

Crozzers commented on Mar 20, 2026

@Crozzers
Contributor

Yeah, this looks like it's down to a change in the rules as to what is a valid em/strong opening/closing delimiter run.

From the minimal repro examples:

  1. Leftmost ** is not a valid opening delimiter run as it's followed by punctuation, but not preceded by punctuation/whitespace as well
  2. Rightmost ** is not a valid closing run as it's preceeded by punctuation but not followed by punctioation/whitespace
  3. Leftmost ** invalid for same reason as 1

For the GFM implementation we tried to follow GFM's spec, and their rules on what a left/right flanking delimiter run can consist of.

In theory you could do something like this to restore the old behaviour:

class ForceOldIAB(markdown2.ItalicAndBoldProcessor):
    name = 'force-old-iab'
    order = (markdown2.Stage.ITALIC_AND_BOLD,), (markdown2.Stage.ITALIC_AND_BOLD,)

    def sub(self, match):
        syntax = match.group(1)
        tag = 'strong' if len(syntax) == 2 else 'em'
        return f'<{tag}>{match.group(2)}</{tag}>'

ForceOldIAB.register()

markdown2.markdown(text, extras=['force-old-iab']

Although testing this, the old strong/em behaviour was changed in #644 before being overwritten by the GFM changes in the same release cycle. So the old behaviour still wouldn't work.

I guess, since the updated behaviour of the old implementation never actually got used, since the GFM changes superseded them and fixed all the issues from that PR anyway, we could revert that change for the old implementation so that users have an option to stick with the old behaviour.
@nicholasserra thoughts on this?

Crozzers commented on Mar 20, 2026

@Crozzers
Contributor

Created #690 just in case with the suggested change of reverting the legacy behaviour, although open to discussion

nicholasserra commented on Mar 26, 2026

@nicholasserra
Collaborator

So basically, if we merge #690 and revert to the old behavior, they'd just need to add their own extra to disable the GFM default, and it would work as expected?

Crozzers commented on Mar 26, 2026

@Crozzers
Contributor

Yeah, pretty much

added a commit that references this issue on Mar 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions