html 2 plain text

Collapse
This topic is closed.
X
X
 
  • Time
  • Show
Clear All
new posts
  • robin

    #1

    html 2 plain text

    hi,
    i remember seeing this simple python function which would take raw html
    and output the content (body?) of the page as plain text (no <..> tags
    etc)
    i have been looking at htmllib and htmlparser but this all seems to
    complicated for what i'm looking for. i just need the main text in the
    body of some arbitrary webbpage to then do some natural-language
    processing with it...
    thanks for pointing me to some helpful resources!

    robin

  • Faber

    #2
    Re: html 2 plain text

    robin wrote:
    [color=blue]
    > i remember seeing this simple python function which would take raw html
    > and output the content (body?) of the page as plain text (no <..> tags
    > etc)
    > i have been looking at htmllib and htmlparser but this all seems to
    > complicated for what i'm looking for. i just need the main text in the
    > body of some arbitrary webbpage to then do some natural-language
    > processing with it...
    > thanks for pointing me to some helpful resources![/color]

    Have a look at the Beautiful Soup library:


    Regards

    --
    Faber

    Get real-time visibility into your key parking metrics, and make data-driven decisions to optimise occupancy, pricing, and revenue. Contact our team today.


    A teacher must always teach to doubt his teaching. -- José Ortega y Gasset

    Comment

    • robin

      #3
      Re: html 2 plain text

      lucks yummy. merci beaucoup.

      robin

      Comment

      • Ravi Teja

        #4
        Re: html 2 plain text

        > i remember seeing this simple python function which would take raw html[color=blue]
        > and output the content (body?) of the page as plain text (no <..> tags
        > etc)[/color]



        Comment

        • garabik-news-2005-05@kassiopeia.juls.savba.sk

          #5
          Re: html 2 plain text

          robin <robin.meier@gm ail.com> wrote:[color=blue]
          > hi,
          > i remember seeing this simple python function which would take raw html
          > and output the content (body?) of the page as plain text (no <..> tags
          > etc)
          > i have been looking at htmllib and htmlparser but this all seems to
          > complicated for what i'm looking for. i just need the main text in the
          > body of some arbitrary webbpage to then do some natural-language
          > processing with it...
          > thanks for pointing me to some helpful resources![/color]

          text=re.sub(r'( ?s)\<.+?\>', '', html_text)
          (this will keep html entities, though)

          --
          -----------------------------------------------------------
          | Radovan Garabík http://kassiopeia.juls.savba.sk/~garabik/ |
          | __..--^^^--..__ garabik @ kassiopeia.juls .savba.sk |
          -----------------------------------------------------------
          Antivirus alert: file .signature infected by signature virus.
          Hi! I'm a signature virus! Copy me into your signature file to help me spread!

          Comment

          • Fredrik Lundh

            #6
            Re: html 2 plain text

            garabik-news-2005-05@kassiopeia.j uls.savba.sk wrote:
            [color=blue]
            > text=re.sub(r'( ?s)\<.+?\>', '', html_text)
            > (this will keep html entities, though)[/color]

            here's a variation that handles that too:



            </F>

            Comment

            Working...