First of all thanks for helping me out.
I have to admit I dont understand some of your suggestiosn, sorry.
I dont know what is the "3D" thing... Is there another way to make it
work something more simple for a newbie like me? Thanks
What I want to do is:
First check all the files from a folder and analyze only the one with the .Seq extension.
What I want to do is to get the reverse complement of the DNA sequence. If their is a problem
with some characters in the DNA Sequence I want the function to tell it to me.
Here are the comp and iupac:
iupac ="GgAaTtCcRrYyM mKkSsWwHhBbVvDd Nn"
comp={"A":"T", "T":"A", "G":"C", "C":"G", "R":"Y", "Y":"R", "M":"K",
"K":"M", "S":"W", "W":"S", "B":"V", "V":"B", "D":"H", "H":"D", "r":"y",
"y":"r", "m":"k", "k":"m", "s":"w", "w":"s", "b":"v", "v":"b", "d":"h",
"h":"d", "a":"t", "t":"a", "g":"c", "c":"g", "N":"N","n":"n" }
So if a $ or Z appears in the DNA sequence, I want to know it.
My code so far:
# -*- coding: iso-8859-1 -*-
import sys
import os
from progadn import *
ab1seq = raw_input("Entr ez le répertoire où sont les fichiers à analyser: ") or None
if ab1seq == None :
print "Erreur: Pas de répertoire! \n" \
"\nAu revoir \n"
sys.exit()
listrep = os.listdir(ab1s eq)
#print listrep
extseq=[]
for f in listrep:
if f[-4:]==".Seq":
extseq.append(f )
#print extseq
for x in extseq:
f=open(x, "r")
seq=f.read()
f.close()
#s=seq
def checkDNA(seq):
"""Retourne une liste des caractères non conformes à l'IUPAC."""
junk=[]
for c in range (len(seq)):
if seq[c] not in iupac:
junk.append([seq[c],c])
#print junk
print "ATTN: Il y a le caractère %s en position %s " % (seq[c],c)
if junk == []:
indinv=range(le n(seq))
indinv.reverse( )
resultat=""
for i in indinv:
resultat +=comp[seq[i]]
return resultat
seq=checkDNA(se q)
-------------------------------------------------------------------------------------------------------------------------
Path: news3!feeder.ne ws-service.com!new s.glorb.com!pos tnews.google.co m!o13g2000cwo.g ooglegroups.com !not-for-mail
From: gry@ll.mit.edu
Newsgroups: comp.lang.pytho n
Subject: Re: problem with the logic of read files
Date: 12 Apr 2005 10:47:17 -0700
Organization: http://groups.google.com
Lines: 104
Message-ID: <1113328037.070 319.136110@o13g 2000cwo.googleg roups.com>
References: <68da8$425bf954 $43461874$29054 @nf1.news-service.com>
NNTP-Posting-Host: 129.55.200.20
Mime-Version: 1.0
Content-Type: text/plain; charset="iso-8859-1"
Content-Transfer-Encoding: quoted-printable
X-Trace: posting.google. com 1113328069 32347 127.0.0.1 (12 Apr 2005 17:47:49 GMT)
X-Complaints-To: groups-abuse@google.co m
NNTP-Posting-Date: Tue, 12 Apr 2005 17:47:49 +0000 (UTC)
In-Reply-To: <68da8$425bf954 $43461874$29054 @nf1.news-service.com>
User-Agent: G2/0.2
Complaints-To: groups-abuse@google.co m
Injection-Info: o13g2000cwo.goo glegroups.com; posting-host=129.55.200 .20;
posting-account=tzIXbQw AAACT3z3X4eITVL tksgiDRxhx
Xref: news-x2.support.nl comp.lang.pytho n:438583
<m_t...@yahoo.c om> wrote:[color=blue]
> I am new to python and I am not in computer science. In fact I am a[/color]
biologist and I ma trying to learn python. So if someone can help me, I
will appreciate it.[color=blue]
> Thanks
>
>
> #!/cbi/prg/python/current/bin/python
> # -*- coding: iso-8859-1 -*-
> import sys
> import os
> from progadn import *
>
> ab1seq =3D raw_input("Entr ez le r=E9pertoire o=F9 sont les fichiers =E0[/color]
analyser: ") or None[color=blue]
> if ab1seq =3D=3D None :
> print "Erreur: Pas de r=E9pertoire! \n"
> "\nAu revoir \n"
> sys.exit()
>
> listrep =3D os.listdir(ab1s eq)
> #print listrep
>
> extseq=3D[]
>
> for f in listrep:[/color]
###### Minor -- this is better said as: if f.endswith(".Se q"):[color=blue]
> if f[-4:]=3D=3D".Seq":
> extseq.append(f )
> # print extseq
>
> for x in extseq:
> f =3D open(x, "r")[/color]
###### seq=3D... discards previous data and refers only to that just
read.
###### It would be simplest to process each file as it is read:
@@@@@@ seq=3Df.read()
@@@@@@ checkDNA(seq)[color=blue]
> seq=3Df.read()
> f.close()
> s=3Dseq
>
> def checkDNA(seq):
> """Retourne une liste des caract=E8res non conformes =E0[/color]
l'IUPAC."""[color=blue]
>
> junk=3D[]
> for c in range (len(seq)):
> if seq[c] not in iupac:
> junk.append([seq[c],c])
> #print junk
> print "ATTN: Il y a le caract=E8re %s en position %s " %[/color]
(seq[c],c)[color=blue]
> if junk =3D=3D []:
> indinv=3Drange( len(seq))
> indinv.reverse( )
> resultat=3D""
> for i in indinv:
> resultat +=3Dcomp[seq[i]]
> return resultat
>
> seq=3DcheckDNA( seq)
> print seq[/color]
##### The program segment you posted did not define "comp" or "iupac",
##### so it's a little hard to guess how it's supposed to work. It
would
##### be helpful if you gave a concise description of what you want the
##### program to do, as well as brief sample of input data.
##### I hope this helps! -- George[color=blue]
>
> #I got the following ( as you see only one file is proceed by the[/color]
function even if more files is in extseq[color=blue]
>
> ['B1-11_win3F_B04_04 .ab1.Seq']
> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq']
> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
'B1-18_win3F_D04_08 .ab1.Seq'][color=blue]
> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
'B1-18_win3F_D04_08 .ab1.Seq', 'B1-18_win3R_E04_10 .ab1.Seq'][color=blue]
> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
'B1-18_win3F_D04_08 .ab1.Seq', 'B1-18_win3R_E04_10 .ab1.Seq',
'B1-19_win3F_F04_12 .ab1.Seq'][color=blue]
> ..
> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
'B1-18_win3F_D04_08 .ab1.Seq', 'B1-18_win3R_E04_10 .ab1.Seq',
'B1-19_win3F_F04_12 .ab1.Seq', 'B1-19_win3R_G04_14 .ab1.Seq',
'B90_win3F_H04_ 16.ab1.Seq', 'B90_win3R_A05_ 01.ab1.Seq',
'DL2-11_win3F_H03_15 .ab1.Seq', 'DL2-11_win3R_A04_02 .ab1.Seq',
'DL2-12_win3F_F03_11 .ab1.Seq', 'DL2-12_win3R_G03_13 .ab1.Seq',
'M7757_win3F_B0 5_03.ab1.Seq', 'M7757_win3R_C0 5_05.ab1.Seq',
'M7759_win3F_D0 5_07.ab1.Seq', 'M7759_win3R_E0 5_09.ab1.Seq',
'TCR700-114_win3F_H05_1 5.ab1.Seq', 'TCR700-114_win3R_A06_0 2.ab1.Seq',
'TRC666-100_win3F_F05_1 1.ab1.Seq', 'TRC666-100_win3R_G05_1 3.ab1.Seq'][color=blue]
>
> after this listing my programs proceed only the last element of this[/color]
listing (TRC666-100_win3R_G05_1 3.ab1.Seq)[color=blue]
>
>[/color]
NNNNNNNNNNNNNNN NNNNNNNNNNNNNNN TCCCGAAGTGTCCCA GAGCAAATAAATGGA CCAAAACGTTTTTAG =
NNNNNNNNNNNNNNN NNNNNNNNNNNNNNN TCCCGAAGTGTCCCA GAGCAAATAAATGGA CCAAAACGTTTTTAG =
AATACTTGAACGTGT AATCTCATTTTAA
**********End Of Post*********** **
I have to admit I dont understand some of your suggestiosn, sorry.
I dont know what is the "3D" thing... Is there another way to make it
work something more simple for a newbie like me? Thanks
What I want to do is:
First check all the files from a folder and analyze only the one with the .Seq extension.
What I want to do is to get the reverse complement of the DNA sequence. If their is a problem
with some characters in the DNA Sequence I want the function to tell it to me.
Here are the comp and iupac:
iupac ="GgAaTtCcRrYyM mKkSsWwHhBbVvDd Nn"
comp={"A":"T", "T":"A", "G":"C", "C":"G", "R":"Y", "Y":"R", "M":"K",
"K":"M", "S":"W", "W":"S", "B":"V", "V":"B", "D":"H", "H":"D", "r":"y",
"y":"r", "m":"k", "k":"m", "s":"w", "w":"s", "b":"v", "v":"b", "d":"h",
"h":"d", "a":"t", "t":"a", "g":"c", "c":"g", "N":"N","n":"n" }
So if a $ or Z appears in the DNA sequence, I want to know it.
My code so far:
# -*- coding: iso-8859-1 -*-
import sys
import os
from progadn import *
ab1seq = raw_input("Entr ez le répertoire où sont les fichiers à analyser: ") or None
if ab1seq == None :
print "Erreur: Pas de répertoire! \n" \
"\nAu revoir \n"
sys.exit()
listrep = os.listdir(ab1s eq)
#print listrep
extseq=[]
for f in listrep:
if f[-4:]==".Seq":
extseq.append(f )
#print extseq
for x in extseq:
f=open(x, "r")
seq=f.read()
f.close()
#s=seq
def checkDNA(seq):
"""Retourne une liste des caractères non conformes à l'IUPAC."""
junk=[]
for c in range (len(seq)):
if seq[c] not in iupac:
junk.append([seq[c],c])
#print junk
print "ATTN: Il y a le caractère %s en position %s " % (seq[c],c)
if junk == []:
indinv=range(le n(seq))
indinv.reverse( )
resultat=""
for i in indinv:
resultat +=comp[seq[i]]
return resultat
seq=checkDNA(se q)
-------------------------------------------------------------------------------------------------------------------------
Path: news3!feeder.ne ws-service.com!new s.glorb.com!pos tnews.google.co m!o13g2000cwo.g ooglegroups.com !not-for-mail
From: gry@ll.mit.edu
Newsgroups: comp.lang.pytho n
Subject: Re: problem with the logic of read files
Date: 12 Apr 2005 10:47:17 -0700
Organization: http://groups.google.com
Lines: 104
Message-ID: <1113328037.070 319.136110@o13g 2000cwo.googleg roups.com>
References: <68da8$425bf954 $43461874$29054 @nf1.news-service.com>
NNTP-Posting-Host: 129.55.200.20
Mime-Version: 1.0
Content-Type: text/plain; charset="iso-8859-1"
Content-Transfer-Encoding: quoted-printable
X-Trace: posting.google. com 1113328069 32347 127.0.0.1 (12 Apr 2005 17:47:49 GMT)
X-Complaints-To: groups-abuse@google.co m
NNTP-Posting-Date: Tue, 12 Apr 2005 17:47:49 +0000 (UTC)
In-Reply-To: <68da8$425bf954 $43461874$29054 @nf1.news-service.com>
User-Agent: G2/0.2
Complaints-To: groups-abuse@google.co m
Injection-Info: o13g2000cwo.goo glegroups.com; posting-host=129.55.200 .20;
posting-account=tzIXbQw AAACT3z3X4eITVL tksgiDRxhx
Xref: news-x2.support.nl comp.lang.pytho n:438583
<m_t...@yahoo.c om> wrote:[color=blue]
> I am new to python and I am not in computer science. In fact I am a[/color]
biologist and I ma trying to learn python. So if someone can help me, I
will appreciate it.[color=blue]
> Thanks
>
>
> #!/cbi/prg/python/current/bin/python
> # -*- coding: iso-8859-1 -*-
> import sys
> import os
> from progadn import *
>
> ab1seq =3D raw_input("Entr ez le r=E9pertoire o=F9 sont les fichiers =E0[/color]
analyser: ") or None[color=blue]
> if ab1seq =3D=3D None :
> print "Erreur: Pas de r=E9pertoire! \n"
> "\nAu revoir \n"
> sys.exit()
>
> listrep =3D os.listdir(ab1s eq)
> #print listrep
>
> extseq=3D[]
>
> for f in listrep:[/color]
###### Minor -- this is better said as: if f.endswith(".Se q"):[color=blue]
> if f[-4:]=3D=3D".Seq":
> extseq.append(f )
> # print extseq
>
> for x in extseq:
> f =3D open(x, "r")[/color]
###### seq=3D... discards previous data and refers only to that just
read.
###### It would be simplest to process each file as it is read:
@@@@@@ seq=3Df.read()
@@@@@@ checkDNA(seq)[color=blue]
> seq=3Df.read()
> f.close()
> s=3Dseq
>
> def checkDNA(seq):
> """Retourne une liste des caract=E8res non conformes =E0[/color]
l'IUPAC."""[color=blue]
>
> junk=3D[]
> for c in range (len(seq)):
> if seq[c] not in iupac:
> junk.append([seq[c],c])
> #print junk
> print "ATTN: Il y a le caract=E8re %s en position %s " %[/color]
(seq[c],c)[color=blue]
> if junk =3D=3D []:
> indinv=3Drange( len(seq))
> indinv.reverse( )
> resultat=3D""
> for i in indinv:
> resultat +=3Dcomp[seq[i]]
> return resultat
>
> seq=3DcheckDNA( seq)
> print seq[/color]
##### The program segment you posted did not define "comp" or "iupac",
##### so it's a little hard to guess how it's supposed to work. It
would
##### be helpful if you gave a concise description of what you want the
##### program to do, as well as brief sample of input data.
##### I hope this helps! -- George[color=blue]
>
> #I got the following ( as you see only one file is proceed by the[/color]
function even if more files is in extseq[color=blue]
>
> ['B1-11_win3F_B04_04 .ab1.Seq']
> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq']
> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
'B1-18_win3F_D04_08 .ab1.Seq'][color=blue]
> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
'B1-18_win3F_D04_08 .ab1.Seq', 'B1-18_win3R_E04_10 .ab1.Seq'][color=blue]
> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
'B1-18_win3F_D04_08 .ab1.Seq', 'B1-18_win3R_E04_10 .ab1.Seq',
'B1-19_win3F_F04_12 .ab1.Seq'][color=blue]
> ..
> ['B1-11_win3F_B04_04 .ab1.Seq', 'B1-11_win3R_C04_06 .ab1.Seq',[/color]
'B1-18_win3F_D04_08 .ab1.Seq', 'B1-18_win3R_E04_10 .ab1.Seq',
'B1-19_win3F_F04_12 .ab1.Seq', 'B1-19_win3R_G04_14 .ab1.Seq',
'B90_win3F_H04_ 16.ab1.Seq', 'B90_win3R_A05_ 01.ab1.Seq',
'DL2-11_win3F_H03_15 .ab1.Seq', 'DL2-11_win3R_A04_02 .ab1.Seq',
'DL2-12_win3F_F03_11 .ab1.Seq', 'DL2-12_win3R_G03_13 .ab1.Seq',
'M7757_win3F_B0 5_03.ab1.Seq', 'M7757_win3R_C0 5_05.ab1.Seq',
'M7759_win3F_D0 5_07.ab1.Seq', 'M7759_win3R_E0 5_09.ab1.Seq',
'TCR700-114_win3F_H05_1 5.ab1.Seq', 'TCR700-114_win3R_A06_0 2.ab1.Seq',
'TRC666-100_win3F_F05_1 1.ab1.Seq', 'TRC666-100_win3R_G05_1 3.ab1.Seq'][color=blue]
>
> after this listing my programs proceed only the last element of this[/color]
listing (TRC666-100_win3R_G05_1 3.ab1.Seq)[color=blue]
>
>[/color]
NNNNNNNNNNNNNNN NNNNNNNNNNNNNNN TCCCGAAGTGTCCCA GAGCAAATAAATGGA CCAAAACGTTTTTAG =
NNNNNNNNNNNNNNN NNNNNNNNNNNNNNN TCCCGAAGTGTCCCA GAGCAAATAAATGGA CCAAAACGTTTTTAG =
AATACTTGAACGTGT AATCTCATTTTAA
**********End Of Post*********** **
Comment