ATTN : Georges ( gry@ll.mit.edu)

Collapse
This topic is closed.
X
X
 
  • Time
  • Show
Clear All
new posts
  • Guest's Avatar

    #1

    ATTN : Georges ( gry@ll.mit.edu)

    First of all thanks for helping me out.

    I have to admit I dont understand some of your suggestiosn, sorry.
    I dont know what is the "3D" thing... Is there another way to make it
    work something more simple for a newbie like me? Thanks

    What I want to do is:
    First check all the files from a folder and analyze only the one with the .Seq extension.
    What I want to do is to get the reverse complement of the DNA sequence. If their is a problem
    with some characters in the DNA Sequence I want the function to tell it to me.

    Here are the comp and iupac:

    iupac ="GgAaTtCcRrYyM mKkSsWwHhBbVvDd Nn"

    comp={"A":"T", "T":"A", "G":"C", "C":"G", "R":"Y", "Y":"R", "M":"K",
    "K":"M", "S":"W", "W":"S", "B":"V", "V":"B", "D":"H", "H":"D", "r":"y",
    "y":"r", "m":"k", "k":"m", "s":"w", "w":"s", "b":"v", "v":"b", "d":"h",
    "h":"d", "a":"t", "t":"a", "g":"c", "c":"g", "N":"N","n":"n" }

    So if a $ or Z appears in the DNA sequence, I want to know it.

    My code so far:
    # -*- coding: iso-8859-1 -*-
    import sys
    import os
    from progadn import *

    ab1seq = raw_input("Entr ez le répertoire où sont les fichiers à analyser: ") or None
    if ab1seq == None :
    print "Erreur: Pas de répertoire! \n" \
    "\nAu revoir \n"
    sys.exit()

    listrep = os.listdir(ab1s eq)
    #print listrep

    extseq=[]

    for f in listrep:
    if f[-4:]==".Seq":
    extseq.append(f )
    #print extseq

    for x in extseq:
    f=open(x, "r")
    seq=f.read()
    f.close()
    #s=seq


    def checkDNA(seq):
    """Retourne une liste des caractères non conformes à l'IUPAC."""

    junk=[]
    for c in range (len(seq)):
    if seq[c] not in iupac:
    junk.append([seq[c],c])
    #print junk
    print "ATTN: Il y a le caractère %s en position %s " % (seq[c],c)
    if junk == []:
    indinv=range(le n(seq))
    indinv.reverse( )
    resultat=""
    for i in indinv:
    resultat +=comp[seq[i]]
    return resultat

    seq=checkDNA(se q)

    -------------------------------------------------------------------------------------------------------------------------

    Path: news3!feeder.ne ws-service.com!new s.glorb.com!pos tnews.google.co m!o13g2000cwo.g ooglegroups.com !not-for-mail
    From: gry@ll.mit.edu
    Newsgroups: comp.lang.pytho n
    Subject: Re: problem with the logic of read files
    Date: 12 Apr 2005 10:47:17 -0700
    Organization: http://groups.google.com
    Lines: 104
    Message-ID: <1113328037.070 319.136110@o13g 2000cwo.googleg roups.com>
    References: <68da8$425bf954 $43461874$29054 @nf1.news-service.com>
    NNTP-Posting-Host: 129.55.200.20
    Mime-Version: 1.0
    Content-Type: text/plain; charset="iso-8859-1"
    Content-Transfer-Encoding: quoted-printable
    X-Trace: posting.google. com 1113328069 32347 127.0.0.1 (12 Apr 2005 17:47:49 GMT)
    X-Complaints-To: groups-abuse@google.co m
    NNTP-Posting-Date: Tue, 12 Apr 2005 17:47:49 +0000 (UTC)
    In-Reply-To: <68da8$425bf954 $43461874$29054 @nf1.news-service.com>
    User-Agent: G2/0.2
    Complaints-To: groups-abuse@google.co m
    Injection-Info: o13g2000cwo.goo glegroups.com; posting-host=129.55.200 .20;
    posting-account=tzIXbQw AAACT3z3X4eITVL tksgiDRxhx
    Xref: news-x2.support.nl comp.lang.pytho n:438583


    <m_t...@yahoo.c om> wrote:[color=blue]
    > I am new to python and I am not in computer science. In fact I am a[/color]
    biologist and I ma trying to learn python. So if someone can help me, I
    will appreciate it.[color=blue]
    > Thanks
    >
    >
    > #!/cbi/prg/python/current/bin/python
    > # -*- coding: iso-8859-1 -*-
    > import sys
    > import os
    > from progadn import *
    >
    > ab1seq =3D raw_input("Entr ez le r=E9pertoire o=F9 sont les fichiers =E0[/color]
    analyser: ") or None[color=blue]
    > if ab1seq =3D=3D None :
    > print "Erreur: Pas de r=E9pertoire! \n"
    > "\nAu revoir \n"
    > sys.exit()
    >
    > listrep =3D os.listdir(ab1s eq)
    > #print listrep
    >
    > extseq=3D[]
    >
    > for f in listrep:[/color]
    ###### Minor -- this is better said as: if f.endswith(".Se q"):[color=blue]
    > if f[-4:]=3D=3D".Seq":
    > extseq.append(f )
    > # print extseq
    >
    > for x in extseq:
    > f =3D open(x, "r")[/color]
    ###### seq=3D... discards previous data and refers only to that just
    read.
    ###### It would be simplest to process each file as it is read:
    @@@@@@ seq=3Df.read()
    @@@@@@ checkDNA(seq)[color=blue]
    > seq=3Df.read()
    > f.close()
    > s=3Dseq
    >
    > def checkDNA(seq):
    > """Retourne une liste des caract=E8res non conformes =E0[/color]
    l'IUPAC."""[color=blue]
    >
    > junk=3D[]
    > for c in range (len(seq)):
    > if seq[c] not in iupac:
    > junk.append([seq[c],c])
    > #print junk
    > print "ATTN: Il y a le caract=E8re %s en position %s " %[/color]
    (seq[c],c)[color=blue]
    > if junk =3D=3D []:
    > indinv=3Drange( len(seq))
    > indinv.reverse( )
    > resultat=3D""
    > for i in indinv:
    > resultat +=3Dcomp[seq[i]]
    > return resultat
    >
    > seq=3DcheckDNA( seq)
    > print seq[/color]

    ##### The program segment you posted did not define "comp" or "iupac",
    ##### so it's a little hard to guess how it's supposed to work. It
    would
    ##### be helpful if you gave a concise description of what you want the

    ##### program to do, as well as brief sample of input data.
    ##### I hope this helps! -- George[color=blue]
    >
    > #I got the following ( as you see only one file is proceed by the[/color]
    function even if more files is in extseq[color=blue]
    >
    > ['B1-11_win3F_B04_04 .ab1.Seq']
    > ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq']
    > ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
    'B1-18_win3F_D04_08 .ab1.Seq'][color=blue]
    > ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
    'B1-18_win3F_D04_08 .ab1.Seq', 'B1-18_win3R_E04_10 .ab1.Seq'][color=blue]
    > ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
    'B1-18_win3F_D04_08 .ab1.Seq', 'B1-18_win3R_E04_10 .ab1.Seq',
    'B1-19_win3F_F04_12 .ab1.Seq'][color=blue]
    > ..
    > ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
    'B1-18_win3F_D04_08 .ab1.Seq', 'B1-18_win3R_E04_10 .ab1.Seq',
    'B1-19_win3F_F04_12 .ab1.Seq', 'B1-19_win3R_G04_14 .ab1.Seq',
    'B90_win3F_H04_ 16.ab1.Seq', 'B90_win3R_A05_ 01.ab1.Seq',
    'DL2-11_win3F_H03_15 .ab1.Seq', 'DL2-11_win3R_A04_02 .ab1.Seq',
    'DL2-12_win3F_F03_11 .ab1.Seq', 'DL2-12_win3R_G03_13 .ab1.Seq',
    'M7757_win3F_B0 5_03.ab1.Seq', 'M7757_win3R_C0 5_05.ab1.Seq',
    'M7759_win3F_D0 5_07.ab1.Seq', 'M7759_win3R_E0 5_09.ab1.Seq',
    'TCR700-114_win3F_H05_1 5.ab1.Seq', 'TCR700-114_win3R_A06_0 2.ab1.Seq',
    'TRC666-100_win3F_F05_1 1.ab1.Seq', 'TRC666-100_win3R_G05_1 3.ab1.Seq'][color=blue]
    >
    > after this listing my programs proceed only the last element of this[/color]
    listing (TRC666-100_win3R_G05_1 3.ab1.Seq)[color=blue]
    >
    >[/color]
    NNNNNNNNNNNNNNN NNNNNNNNNNNNNNN TCCCGAAGTGTCCCA GAGCAAATAAATGGA CCAAAACGTTTTTAG =
    NNNNNNNNNNNNNNN NNNNNNNNNNNNNNN TCCCGAAGTGTCCCA GAGCAAATAAATGGA CCAAAACGTTTTTAG =
    AATACTTGAACGTGT AATCTCATTTTAA



    **********End Of Post*********** **




  • Pierre-Frédéric Caillaud

    #2
    Re: ATTN : Georges ( gry@ll.mit.edu)


    [color=blue]
    > My code so far:
    > # -*- coding: iso-8859-1 -*-
    > import sys
    > import os
    > from progadn import *
    >
    > ab1seq = raw_input("Entr ez le répertoire où sont les fichiers à
    > analyser: ") or None[/color]

    Ce serait mieux d'utiliser sys.argv pour spécifier le répertoire dans la
    ligne de commande du programme :
    import sys
    help(sys.argv)
    [color=blue]
    > if ab1seq == None :
    > print "Erreur: Pas de répertoire! \n" \
    > "\nAu revoir \n"
    > sys.exit()[/color]

    je propose :

    import os, os.path, sys

    def usage():
    print "documentation. .."
    sys.exit(-1)


    args = sys.argv[1:]

    if not args:
    usage()

    files = []
    for path in args:
    if os.path.isfile( path ):
    files.append( path )
    elif os.path.isdir( path ):
    files.extend( [os.path.join( path, fname ) for fname in os.listdir( path
    )] )
    else:
    print "%s n'est ni un fichier ni un répertoire..." % path
    usage()

    files = [ fname for fname in files if fname.endswith( ".Seq" ) ]
    88
    if not files:
    print "Aucun fichier a traiter."
    usage()

    print "Fichier à traiter :"
    print ", ".join( files )

    for path in files:
    print path
    checkDNA( open( path ).read() )
    [color=blue]
    > def checkDNA(seq):
    > """Retourne une liste des caractères non conformes à l'IUPAC."""
    >
    > junk=[]
    > for c in range (len(seq)):
    > if seq[c] not in iupac:
    > junk.append([seq[c],c])
    > #print junk
    > print "ATTN: Il y a le caractère %s en position %s " %
    > (seq[c],c)
    > if junk == []:
    > indinv=range(le n(seq))
    > indinv.reverse( )
    > resultat=""
    > for i in indinv:
    > resultat +=comp[seq[i]]
    > return resultat[/color]

    Je réécris un peu votre fonction d'une manière plus "python", à placer
    dans le programme avant son appel bien sûr !

    def checkDNA( seq ):
    seq = seq.strip()
    if not seq:
    print "Fichier vide."
    return
    resultat = []
    for i,c in enumerate(seq):
    try:
    resultat.append ( comp[c] )
    except KeyError:
    print "Catactère <%s> en position <%d> invalide" % (c,i)
    resultat.revers e()
    return ''.join( resultat )


    [color=blue]
    >
    > seq=checkDNA(se q)
    >
    > -------------------------------------------------------------------------------------------------------------------------
    >
    > Path:
    > news3!feeder.ne ws-service.com!new s.glorb.com!pos tnews.google.co m!o13g2000cwo.g ooglegroups.com !not-for-mail
    > From: gry@ll.mit.edu
    > Newsgroups: comp.lang.pytho n
    > Subject: Re: problem with the logic of read files
    > Date: 12 Apr 2005 10:47:17 -0700
    > Organization: http://groups.google.com
    > Lines: 104[/color]
    2> Message-ID: <1113328037.070 319.136110@o13g 2000cwo.googleg roups.com>[color=blue]
    > References: <68da8$425bf954 $43461874$29054 @nf1.news-service.com>
    > NNTP-Posting-Host: 129.55.200.20
    > Mime-Version: 1.0
    > Content-Type: text/plain; charset="iso-8859-1"
    > Content-Transfer-Encoding: quoted-printable
    > X-Trace: posting.google. com 1113328069 32347 127.0.0.1 (12 Apr 2005
    > 17:47:49 GMT)
    > X-Complaints-To: groups-abuse@google.co m
    > NNTP-Posting-Date: Tue, 12 Apr 2005 17:47:49 +0000 (UTC)
    > In-Reply-To: <68da8$425bf954 $43461874$29054 @nf1.news-service.com>
    > User-Agent: G2/0.2
    > Complaints-To: groups-abuse@google.co m
    > Injection-Info: o13g2000cwo.goo glegroups.com; posting-host=129.55.200 .20;
    > posting-account=tzIXbQw AAACT3z3X4eITVL tksgiDRxhx
    > Xref: news-x2.support.nl comp.lang.pytho n:438583
    >
    >
    > <m_t...@yahoo.c om> wrote:[color=green]
    >> I am new to python and I am not in computer science. In fact I am a[/color]
    > biologist and I ma trying to learn python. So if someone can help me, I
    > will appreciate it.[color=green]
    >> Thanks
    >>
    >>
    >> #!/cbi/prg/python/current/bin/python
    >> # -*- coding: iso-8859-1 -*-
    >> import sys
    >> import os
    >> from progadn import *
    >>
    >> ab1seq =3D raw_input("Entr ez le r=E9pertoire o=F9 sont les fichiers =E0[/color]
    > analyser: ") or None[color=green]
    >> if ab1seq =3D=3D None :
    >> print "Erreur: Pas de r=E9pertoire! \n"
    >> "\nAu revoir \n"
    >> sys.exit()
    >>
    >> listrep =3D os.listdir(ab1s eq)
    >> #print listrep
    >>
    >> extseq=3D[]
    >>
    >> for f in listrep:[/color]
    > ###### Minor -- this is better said as: if f.endswith(".Se q"):[color=green]
    >> if f[-4:]=3D=3D".Seq":
    >> extseq.append(f )
    >> # print extseq
    >>
    >> for x in extseq:
    >> f =3D open(x, "r")[/color]
    > ###### seq=3D... discards previous data and refers only to that just
    > read.
    > ###### It would be simplest to process each file as it is read:
    > @@@@@@ seq=3Df.read()
    > @@@@@@ checkDNA(seq)[color=green]
    >> seq=3Df.read()
    >> f.close()
    >> s=3Dseq
    >>
    >> def checkDNA(seq):
    >> """Retourne une liste des caract=E8res non conformes =E0[/color]
    > l'IUPAC."""[color=green]
    >>
    >> junk=3D[]
    >> for c in range (len(seq)):
    >> if seq[c] not in iupac:
    >> junk.append([seq[c],c])
    >> #print junk
    >> print "ATTN: Il y a le caract=E8re %s en position %s " %[/color]
    > (seq[c],c)[color=green]
    >> if junk =3D=3D []:
    >> indinv=3Drange( len(seq))
    >> indinv.reverse( )
    >> resultat=3D""
    >> for i in indinv:
    >> resultat +=3Dcomp[seq[i]]
    >> return resultat
    >>
    >> seq=3DcheckDNA( seq)
    >> print seq[/color]
    >
    > ##### The program segment you posted did not define "comp" or "iupac",
    > ##### so it's a little hard to guess how it's supposed to work. It
    > would
    > ##### be helpful if you gave a concise description of what you want the
    >
    > ##### program to do, as well as brief sample of input data.
    > ##### I hope this helps! -- George[color=green]
    >>
    >> #I got the following ( as you see only one file is proceed by the[/color]
    > function even if more files is in extseq[color=green]
    >>
    >> ['B1-11_win3F_B04_04 .ab1.Seq']
    >> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq']
    >> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
    > 'B1-18_win3F_D04_08 .ab1.Seq'][color=green]
    >> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
    > 'B1-18_win3F_D04_08 .ab1.Seq', 'B1-18_win3R_E04_10 .ab1.Seq'][color=green]
    >> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
    > 'B1-18_win3F_D04_08 .ab1.Seq', 'B1-18_win3R_E04_10 .ab1.Seq',
    > 'B1-19_win3F_F04_12 .ab1.Seq'][color=green]
    >> ..
    >> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
    > 'B1-18_win3F_D04_08 .ab1.Seq', 'B1-18_win3R_E04_10 .ab1.Seq',
    > 'B1-19_win3F_F04_12 .ab1.Seq', 'B1-19_win3R_G04_14 .ab1.Seq',
    > 'B90_win3F_H04_ 16.ab1.Seq', 'B90_win3R_A05_ 01.ab1.Seq',
    > 'DL2-11_win3F_H03_15 .ab1.Seq', 'DL2-11_win3R_A04_02 .ab1.Seq',
    > 'DL2-12_win3F_F03_11 .ab1.Seq', 'DL2-12_win3R_G03_13 .ab1.Seq',
    > 'M7757_win3F_B0 5_03.ab1.Seq', 'M7757_win3R_C0 5_05.ab1.Seq',
    > 'M7759_win3F_D0 5_07.ab1.Seq', 'M7759_win3R_E0 5_09.ab1.Seq',
    > 'TCR700-114_win3F_H05_1 5.ab1.Seq', 'TCR700-114_win3R_A06_0 2.ab1.Seq',
    > 'TRC666-100_win3F_F05_1 1.ab1.Seq', 'TRC666-100_win3R_G05_1 3.ab1.Seq'][color=green]
    >>
    >> after this listing my programs proceed only the last element of this[/color]
    > listing (TRC666-100_win3R_G05_1 3.ab1.Seq)[color=green]
    >>
    >>[/color]
    > NNNNNNNNNNNNNNN NNNNNNNNNNNNNNN TCCCGAAGTGTCCCA GAGCAAATAAATGGA CCAAAACGTTTTTAG =
    > NNNNNNNNNNNNNNN NNNNNNNNNNNNNNN TCCCGAAGTGTCCCA GAGCAAATAAATGGA CCAAAACGTTTTTAG =
    > AATACTTGAACGTGT AATCTCATTTTAA
    >
    >
    >
    > **********End Of Post*********** **
    >
    >
    >
    >[/color]

    Comment

    Working...