Commits · 55d10b825c1c695d98561d78a67d2244dbd27b89 · Mirrors / mutalyzer

May 18, 2015
- New description extractor web interface · 55d10b82
  Jeroen F.J. Laros authored 9 years ago and Vermaat committed 9 years ago
  
  We can now compare two sequences by supplying their sequence strings, accession numbers, or uploaded file.
  55d10b82
May 01, 2015
- Fix descriptionExtract webservice · 7d7cb6af
  Vermaat authored 9 years ago
  
  7d7cb6af
Apr 30, 2015
- Moved describe functionality to the extractor package. · 6c64e5ee
  Jeroen F.J. Laros authored 9 years ago and Vermaat committed 9 years ago
  
  6c64e5ee
- PEP8. · 57c55d0f
  Jeroen F.J. Laros authored 9 years ago and Vermaat committed 9 years ago
  
  57c55d0f
- Integrated the description extractor in the website. · 216146bb
  Laros authored 10 years ago and Vermaat committed 9 years ago
  
  216146bb
- Some more refactoring. · 2db722ff
  Laros authored 10 years ago and Vermaat committed 9 years ago
  
  2db722ff
- Fixed empty allele bug. · 52724cc8
  Laros authored 10 years ago and Vermaat committed 9 years ago
  
  52724cc8
- Fixed erroneous unit tests. · b0d85531
  Laros authored 10 years ago and Vermaat committed 9 years ago
  
  b0d85531
- Made the inserted and deleted sequences uniform. · 49534102
  Laros authored 10 years ago and Vermaat committed 9 years ago
  
  49534102
- Checked the generated positions. · 036fc241
  Laros authored 10 years ago and Vermaat committed 9 years ago
  
  036fc241
- Use new extract package for the description extractor · 534a41fe
  Vermaat authored 11 years ago
  
  This is a work in progress as there still seem to be some bugs. For example, some unit tests fail due to incorrect descriptions generated and others fail due to a crash.
  534a41fe
- Add some JSON and SOAP service tests · 100f53b2
  Vermaat authored 9 years ago
  
  100f53b2
Jan 30, 2015

Discard incomplete genes in genbank reference files · 73c0862f

Vermaat authored 10 years ago

Many genbank reference files contain more than one gene, especially
slices from an assembly. Some of these genes may be incomplete in
the reference file (i.e., either start or end exceeds the outer
coordinates). We cannot really do anything with these genes, so we
discard them during parsing.

73c0862f

Fix broken DMD reference in unit tests · 51d8cc50
Vermaat authored 10 years ago

51d8cc50

Add getGeneLocation webservice method · e06452a1

Vermaat authored 10 years ago

Given a gene symbol and optional genome build, this returns the location
of the gene.

Primary motivation for this is LOVD, where it will be used in combination
with sliceChromsome as an alternative for sliceChromosomeByGene which only
works on the fixed Ensembl genome build.

e06452a1

Nov 24, 2014
- Fix form buttons and general language issues · 9e6ca731
  Vermaat authored 10 years ago
  
  9e6ca731
- Many fixes in templates · 5fc78480
  Vermaat authored 10 years ago
  
  5fc78480
- New website layout by Landscape · 5010bbec
  Jeroen Laros authored 10 years ago and Vermaat committed 10 years ago
  
  5010bbec
- Check batch job input field length · b7c8fddd
  Vermaat authored 10 years ago
  
  b7c8fddd
Oct 21, 2014
- Unit tests for unicode strings · 66629914
  Vermaat authored 10 years ago
  
  66629914
Oct 20, 2014

Correctly handle batch job input and output encodings · 8acb0970
Vermaat authored 10 years ago

8acb0970

Use unicode strings · 2a4dc3c1

Vermaat authored 10 years ago

Don't fix what ain't broken. Unfortunately, string handling in Mutalyzer
really is broken. So we fix it.

Internally, all strings should be represented by unicode strings as much as
possible. The main exception are large reference sequence strings. These can
often better be BioPython sequence objects, since that is how we usually get
them in the first place.

These changes will hopefully make Mutalyzer more reliable in working with
incoming data. As a bonus, they're a first (small) step towards Python 3
compatibility [1].

Our strategy is as follows:

1. We use `from __future__ import unicode_literals` at the top of every file.
2. All incoming strings are decoded to unicode (if necessary) as soon as
   possible.
3. Outgoing strings are encoded to UTF8 (if necessary) as late as possible.
4. BioPython sequence objects can be based on byte strings as well as unicode
   strings.
5. In the database, everything is UTF8.
6. We worry about uploaded and downloaded reference files and batch jobs in a
   later commit.

Point 1 will ensure that all string literals in our source code will be
unicode strings [2].

As for point 4, sometimes this may even change under our eyes (e.g., calling
`.reverse_complement()` will change it to a byte string). We don't care as
long as they're BioPython objects, only when we get the sequence out we must
have it as unicode string. Their contents are always in the ASCII range
anyway.

Although `Bio.Seq.reverse_complement` works fine on Python byte strings (and
we used to rely on that), it crashes on a Python unicode string. So we take
care to only use it on BioPython sequence objects and wrote our own reverse
complement function for unicode strings (`mutalyzer.util.reverse_complement`).

As for point 5, SQLAlchemy already does a very good job at presenting decoding
from and encoding to UTF8 for us.

The Spyne documentation has the following to say about their `String` and
`Unicode` types [3]:

> There are two string types in Spyne: `spyne.model.primitive.Unicode` and
> `spyne.model.primitive.String` whose native types are `unicode` and `str`
> respectively.
>
> Unlike the Python `str`, the Spyne `String` is not for arbitrary byte
> streams. You should not use it unless you are absolutely, positively sure
> that you need to deal with text data with an unknown encoding. In all other
> cases, you should just use the `Unicode` type. They actually look the same
> from outside, this distinction is made just to properly deal with the quirks
> surrounding Python-2's `unicode` type.
>
> Remember that you have the `ByteArray` and `File` types at your disposal
> when you need to deal with arbitrary byte streams.
>
> The `String` type will be just an alias for `Unicode` once Spyne gets ported
> to Python 3. It might even be deprecated and removed in the future, so make
> sure you are using either `Unicode` or `ByteArray` in your interface
> definitions.

So let's not ignore that and never use `String` anymore in our webservice
interface.

For the command line interface it's a bit more complicated, since there seems
to be no reliable way to get the encoding of command line arguments. We use
`sys.stdin.encoding` as a best guess.

For us to interpret a sequence of bytes as text, it's key to be aware of their
encoding. Once decoded, a text string can be safely used without having to
worry about bytes. Without unicode we're nothing, and nothing will help
us. Maybe we're lying, then you better not stay. But we could be safer, just
for one day. Oh-oh-oh-ohh, oh-oh-oh-ohh, just for one day.

[1] https://docs.python.org/2.7/howto/pyporting.html
[2] http://python-future.org/unicode_literals.html
[3] http://spyne.io/docs/2.10/manual/03_types.html#strings

2a4dc3c1

Oct 15, 2014

Fix several error cases in LOVD2 getGS call · bcef1633

Vermaat authored 10 years ago

The `getGS` website view for LOVD2 would report "transcript not found" if
the genomic reference has multiple transcripts annotated or if the variant
description raises an error in the variant checker.

bcef1633

Oct 04, 2014
- Fix crash in position converter batch job · 55ca04e1
  Vermaat authored 10 years ago
  
  Fixes Trac#174
  55ca04e1
Sep 26, 2014
- Fix unit test for renaming in parent commit · ae685116
  Vermaat authored 10 years ago
  
  ae685116
Sep 22, 2014
- Announcement in info webservice method · 763ab1f7
  Vermaat authored 10 years ago
  
  Closes #11
  763ab1f7
Sep 19, 2014
- Upload a genbank file using the SOAP webservice · a9cb95f4
  Vermaat authored 10 years ago
  
  a9cb95f4
Aug 27, 2014
- Move from nose to pytest for unit tests · e6f19d1c
  Vermaat authored 10 years ago
  
  See http://pytest.org/
  e6f19d1c
Jun 24, 2014
- Add test case for minus in gene symbol · 86c2c143
  Vermaat authored 10 years ago
  
  86c2c143
Mar 01, 2014

Reverse complement range insertions/insertion-deletions · 57120a89

Vermaat authored 11 years ago

The name checker supports reverse complement ranges in insertions
and insertions-deletions, for example `3_4ins8_12inv'.

Reverse complement range insertions and insertion-deletions are not
part of the current HGVS nomenclature, but will be proposed.

57120a89

Feb 28, 2014

Range and compound insertions/insertion-deletions · 31b2f13a

Vermaat authored 11 years ago

The name checker supports ranges in insertions and insertion-
deletions, for example `3_4ins8_12`, and compound insertions and
insertion-deletions, for example `3_4ins[ATC;8_12]`.
The inserted sequences are accepted and concatenated before any
further processing, so reported descriptions show only the
concatenated sequences.
The support for ranges is limited to genomic descriptions.

The position converter supports compound insertions and
insertion-deletions, not ranges.

Compound insertions and insertion-deletions are not part of the
current HGVS nomenclature, but will be proposed.

31b2f13a

Feb 22, 2014
- Conveniently create tables on first use for in-memory SQLite · 6b6a846b
  Vermaat authored 11 years ago
  
  6b6a846b
Feb 17, 2014

Rename organelle_type to organelle in chromosome model · 352c590b

Vermaat authored 11 years ago

Also, the value for nuclear chromosomes is now `nucleus` instead of
`chromosome` for better alignment with the other value `mitochondrion`.

Note that I did not bother to make an Alembic migration for this, since
we don't have any installations besides my own yet anyway.

352c590b

Update BioPython dependency to 1.63 · 0de48334
Vermaat authored 11 years ago

0de48334

Jan 22, 2014

Use fixtures in the unit tests · c49d49f0

Vermaat authored 11 years ago

This is The Good Stuff. The entire test suite can now be run without
having to setup a database, running the batch checker, any of the web
services or the website. It even passes without an internet connection.
In, like, 30 seconds! Awesome!

This means tests don't randomly fail after some reference sequence
changes on the NCBI server and it doesn't take an entire configured
server with mapping database setup to run the tests. Those are things
of the past! No more frustrations, Mutalyzer is testable!

Going down now...

The mountain screamed three times today
I guess it thought it'd like to play
How much does one have to pay
To fry a peak and melt away
Launching titan's breath on mine
The sweating measure lands on time

And the old man, down by the river
Well he walks up and he walks on down
To the spaceship that's parked at your doorstep
And it's waiting to take you away now

Goin' down now
Goin' down now

Looking for the rate that crowed
He's hooked up down in Mexico
Slap my nerve now give me more
It's my disaster friend, not yours

And the old man, down by the river
Well he walks up and he walks on down
To the spaceship that's parked at your doorstep
And it's waiting to take you away now

And the last one, it's down by the river
Where he gets up and he walks on down
To the spaceship that's parked at your doorstep
And it's waiting to take you away now

It's down by the river, it's always this way now
It's down by the river, it's always this way now

Going down now
Going down now
now, now, now

down, down, down

c49d49f0

Jan 10, 2014

Remove obsolete Db module · 667f39a6

Vermaat authored 11 years ago

Now that we ported the database to SQLAlchemy, we remove the obsolete Db
module and all references to it.

667f39a6

Use Redis for stat counters · 8fa5c251

Vermaat authored 11 years ago

The Redis client automatically falls back to a mock Redis server if no
Redis server is configured. Therefore, a Redis server is not needed to
run Mutalyzer. You'll just not get any aggregate stat counts over
different runs.

8fa5c251

Port Mapping database module to SQLAlchemy · e9bf1bc9

Vermaat authored 11 years ago

This introduces a proper notion of genome assemblies. Transcript
mappings for alle genome assemblies are in the same database, which
is better for maintenance. Updating transcript mappings is also
simplified a lot, especially from NCBI mapview files where we now
require a preprocessing sort on the input file.

Overall, this port touches a lot of Mutalyzer code, so beware.

e9bf1bc9

Jan 04, 2014
- Some fixes for running the unit tests · f2a6cc59
  Vermaat authored 11 years ago
  
  f2a6cc59
- Temporarily skip tests using AL449423.14 (no longer valid) · 323a8be1
  Vermaat authored 11 years ago
  
  323a8be1