re pattern for matching JS/CSS

Collapse
This topic is closed.
X
X
 
  • Time
  • Show
Clear All
new posts
  • i80and

    #1

    re pattern for matching JS/CSS

    I'm working on a program to remove tags from a HTML document, leaving
    just the content, but I want to do it simply. I've finished a system
    to remove simple tags, but I want all CSS and JS to be removed. What
    re pattern could I use to do that?

    I've tried
    '<script[\S\s]*/script>'
    but that didn't work properly. I'm fairly basic in my knowledge of
    Python, so I'm still trying to learn re.
    What pattern would work?

  • ina

    #2
    Re: re pattern for matching JS/CSS


    i80and wrote:
    I'm working on a program to remove tags from a HTML document, leaving
    just the content, but I want to do it simply. I've finished a system
    to remove simple tags, but I want all CSS and JS to be removed. What
    re pattern could I use to do that?
    >
    I've tried
    '<script[\S\s]*/script>'
    but that didn't work properly. I'm fairly basic in my knowledge of
    Python, so I'm still trying to learn re.
    What pattern would work?
    I use re.compile("<sc ript.*?</script>",re.DOT ALL)
    for scripts. I strip this out first since my tag stripping re will
    strip out script tags as well hope this was of help.

    Comment

    • Tim Chase

      #3
      Re: re pattern for matching JS/CSS

      >I've tried
      >'<script[\S\s]*/script>'
      >but that didn't work properly. I'm fairly basic in my knowledge of
      >Python, so I'm still trying to learn re.
      >What pattern would work?
      >
      I use re.compile("<sc ript.*?</script>",re.DOT ALL)
      for scripts. I strip this out first since my tag stripping re will
      strip out script tags as well hope this was of help.
      This won't catch various alterations of

      <
      script
      >
      doEvil()
      <
      /
      script
      >
      which is valid html/xhtml.

      For less valid html, but still attemptable, one might find
      something like

      <scrip<script>h ah</script>t>doEvil ()</script>

      which, if you nuke your pattern, leaves the valid but unwanted

      <script>doEvil( )</script>

      I'd propose that it's better to use something such as
      BeautifulSoup that actually parses the HTML, and then skim
      through it whitelisting the tags you plan to allow, and skipping
      the emission of any tags that don't make the whitelist.

      -tkc




      Comment

      Working...