Skip to content

Tag: Information

  • Welcome to Dr. Data!

    Hello, and welcome. If you have found your way here, you are probably one of a few kinds of reader: a student who wants to follow one of my online courses, a graduate student wondering whether a method will hold up on real data, a colleague thinking about a collaboration, a practitioner trying to get a pipeline from “it works on my laptop” to something defensible, or someone simply curious what a lecturer in Information Systems actually does between lectures. Whichever it is, I am glad you are here, and I wanted this first post to explain what you will find on this site and how, if you would like to, we might work together.

    My name is Joseph Bonello. I lecture in Information Systems at the Faculty of ICT, University of Malta, and my research sits at the intersection of data science, artificial intelligence, and bioinformatics — with a particular interest in how machine learning methods can be made to work reliably on biological data, where ground truth is scarce and noise is plentiful. That last part matters more than it sounds: a great deal of applied machine learning writing assumes clean, plentiful, well-labelled data. Most of the interesting problems I encounter do not look like that, and this site is largely an attempt to write and teach honestly about the gap between the two.

    What you will find here

    The site is organised around a few kinds of content, and you are welcome to follow whichever thread is useful to you.

    • Research. Peer-reviewed papers, working drafts, datasets, and code releases, with authors, venues, and DOIs listed properly rather than buried in a CV PDF. Where a paper has an open dataset or a GitHub repository attached, I link to it directly.
    • Talks. Invited lectures, conference presentations, and the occasional podcast appearance, with recordings and slides where they exist.
    • Projects. A running account of what I am currently working on — not the tidy, finished version that eventually becomes a paper, but the state of things while they are still in progress, including what is working and what is not yet.
    • Reading. A running list of the books, papers, and essays currently on my desk, mostly so I have some record of what shaped my thinking on a given problem, and partly because I am regularly asked for reading recommendations.
    • Teaching. Live cohort courses and self-paced material on applied machine learning, aimed at graduate students, postdoctoral researchers, and practitioners working with messy, real-world data rather than benchmark datasets.
    • Writing. Longer articles, like this one, on methodology, tools, and the recurring mistakes I see — in my own work as much as anyone else’s — when applying machine learning to biological and other noisy domains.

    If you would rather not check back regularly, the quarterly digest covers all of it: what I have been reading, a short update on research in progress, one longer essay per issue, and links to anything newly published. Four letters a year, nothing in between unless something genuinely unusual happens, and no sponsors, tracking, or padding. You can subscribe from the newsletter page whenever you like.

    How we might work together

    I get in touch with people, and people get in touch with me, for a handful of fairly distinct reasons, and it is worth being explicit about what each of them looks like so you know what to expect.

    Research collaboration

    If you are working on a problem where applied machine learning meets biological or otherwise noisy real-world data, and you think there might be a shared angle worth exploring, I would like to hear about it. Tell me what you are working on, where the interesting difficulty is, and what a useful outcome would look like for you.

    Consultancy

    For research groups and small teams, I take on a limited number of engagements each year: methodology review, validation strategy, and applied machine learning work where the stakes for getting it wrong are real. These are not slide-deck engagements — they produce a written report, working code, and something you can act on. The consultancy page sets out how an engagement typically runs, and the first conversation is free and without obligation.

    Teaching and supervision

    If you are looking for structured training rather than a bespoke engagement, the live cohorts and self-paced courses cover applied machine learning for people working with real, imperfect data. I also take on a small number of dissertation and thesis supervision enquiries each year — if that is what you have in mind, include your programme, year, and a short paragraph on the topic you are considering.

    Media, press, and everything else

    For interviews, commentary, or editorial work, say so and include a deadline if there is one. And if your reason for writing does not fit neatly into any of the above, write anyway — the contact form has a general enquiry option for exactly that, and I read everything that comes through it.

    I aim to reply within five working days; during semester it can occasionally take a little longer, and if something is genuinely urgent, say so in the subject line and I will try to move it up the queue.

    A closing note

    Most of what ends up on this site is written because I could not find a clear, honest account of it elsewhere, and I suspect I am not the only person who went looking. If that describes you too, I hope something here is useful. And if you have a question, a correction, a dataset that broke your model in an interesting way, or a collaboration in mind, do not hesitate to write — that is, after all, what the site is for.

    Welcome aboard.

    — Joseph