Genomics is moving fast, but how do we get biology the hardware and software it deserves? Ben Busby, NVIDIA’s global alliances manager for omics, joins host Eleanor Howe to dig into the practical crossroads of GPU computing, open-source bioinformatics, and the next wave of precision medicine. They get specific about what GPU acceleration changes in day-to-day genomics and single-cell analysis. Faster pipelines can mean lower costs, tighter iteration loops, and newly feasible questions. They also explore the hardest bottleneck—reliable longitudinal multi-omics data—and challenge a popular assumption about AI in biomedicine: bigger isn’t always better.
GUEST BIOs
Ben Busby, Ph.D., Global Alliances Manager, Omics, NVIDIA
Ben Busby, Ph.D., is a renowned leader in biomedical informatics, computational biology, and interdisciplinary data science. With a career spanning academia, industry, and government, he is dedicated to making biomedical data science a more productive and collaborative environment for bioinformaticians and data scientists. Dr. Busby currently serves as senior alliances manager, genomics at NVIDIA, where he drives strategic collaborations and innovations in genomics, AI, and biomedical research. He is also an adjunct faculty member in the computational biology department at Carnegie Mellon University, contributing to cutting-edge research and education. He holds a doctorate in biochemistry from the University of Maryland, Baltimore, and an undergraduate degree in biochemistry from the University of Maryland, Baltimore County.
TRANSCRIPT
Welcome And Ben’s Path
Eleanor Howe
Hello
everyone,
and
welcome
to
Trends
from
the
Trenches.
My
guest
today
works
at
the
scene
where
genomics
meets
the
hardware
that
now
runs
it.
He's
one
of
the
rare
people
who
can
speak
to
both
sides
with
real
fluency.
Ben
Busby
is
the
global
alliances
manager
for
Omics
at
NVIDIA.
His
work
centers
on
the
rapid
prototyping
of
software
for
bioinformatics
and
precision
medicine,
and
on
finding
better
ways
to
teach
complex
subjects
at
the
postgraduate
level,
both
in-person
and
online.
Before
NVIDIA,
Ben
spent
years
as
a
data
scientist
at
NCBI.
He's
an
advisor
to
Johns
Hopkins,
an
adjunct
faculty
at
Carnegie
Mellon
University,
where
he
still
keeps
his
hands
in
his
own
research.
Ben,
welcome
to
the
trenches.
Thanks,
Eleanor.
I'm
glad
to
have
you
here.
So
let's
get
started
with
telling
the
audience
a
bit
more
about
your
background.
Can
you
tell
us
what
pulled
you
out
of
population
genetics
and
evolutionary
biology
and
into
hardware,
software,
and
NVIDIA?
Ben Busby
Well,
I
did
a
postdoc
in
evolutionary
biology,
and
it
was
awesome.
But
one
thing
I
realized
very
quickly
is
that
a
lot
of
people
were
writing
their
own
scripts.
Many
of
them
were
essentially
prototypes,
sort
of
pushing
the
limits.
But
a
lot
of
people
were
working
as
individuals.
So
I
got
really
interested
in
how
people
write
software
together
and
sort
of
think
as
a
scientific
team.
And
so
really
started
running
hackathons
thinking
about
how
to
push
the
limits
of
what
we
could
do
with
genomics
as
well
as
population
genetics,
etc.
etc.
Sort
of
starting
to
get
into
multi-omics
and
multimodal
data
even
then.
And
and
that
really
engaged
a
lot
of
people.
People
got
really
excited
about
that.
So
I
think
over
the
last
decade
or
so,
really
have
I've
spent
a
lot
of
time
doing
engagement
of
folks
with
computational
biology.
Eleanor Howe
Yeah,
I
think
of
you
as
a
like
a
nerd-to-nerd
translator.
Does
that
sound
fair?
Ben Busby
Thank
you
for
that.
I
that's
that's
very
nice
of
you
to
say.
From Solo Scripts To Team Frameworks
Eleanor Howe
So
you're
still
a
working
researcher
in
that
ongoing
work
with
car
Carnegie
Mellon.
What
are
you
working
on
right
now
and
what's
at
stake
there?
Ben Busby
So
there's
two
things
I'm
particularly
excited
about
in
the
research
world.
One
is
thinking
about
really
sort
of
data
federation
between
biomanks.
And
that's
been
really
exciting,
and
that's
something
I've
worked
on
with
Carnegie
Mellon
quite
a
bit
per
se.
And
then
the
other
thing
is
I'm
really
interested
in
being
able
to
process
millions
of
genomes
at
the
scale
of
biology,
but
not
just
in
and
of
themselves,
being
able
to
mix
with
phenotype
as
well
as
transient
data
types
like
transcriptomics
and
proteomics,
really,
and
and
really
sort
of
build
models
that
are
really
specific
and
answer
real-world
questions
in
this
sense.
Eleanor Howe
So,
you
know,
not
a
big
deal
at
all,
not
a
big
problem.
Ben Busby
It
just
depends
on
what
your
motivations
are,
right?
Like,
I
mean,
if
you
want
sort
of
genomics
to
inform
healthcare,
then
you
need
to
do
work.
If
you
don't
care
about
that,
then
of
course
you
can
work
on
other
things.
Eleanor Howe
So
then
when
did
you
realize
that
the
infrastructure
and
the
hardware
was
just
as
important,
or
at
least
similarly
important
to
the
biology
here?
Ben Busby
Well,
I
think
one
thing
I
realized
is
that,
you
know,
I
mean,
again,
there
were
a
lot
of
individuals
doing
brilliant
work,
but
there
were
very
few
frameworks,
and
this
led
to
somewhat
of
a
lack
of
reproducibility,
as
well
as,
you
know,
there
was
a
lack
of
folks
being
able
to
pick
up
on
other
people's
code
bases.
And
so
now,
of
course,
we
have
agents
and
models
to
help
us
with
that.
But
fundamentally,
I
think
it's
important
to
have
frameworks
for
standardization,
particularly
when
we're
talking
about
fundamental
types
of
data
analysis.
Eleanor Howe
Yeah,
yeah.
And
now
that
you're
in
NVIDIA,
you
get
to
see
up
close
the
free
software
and
the
tools
that
they
release
to
you
know
enable
people
to
use
their
chips
better.
So,
you
know,
how
well
adopted
are
these
by
the
community?
And
NVIDIA Genomics Tools People Miss
Eleanor Howe
maybe
you
could
tell
us
a
little
about
them.
Ben Busby
Well,
yeah,
so
I
mean,
I
think
a
lot
of
people
are
quite
familiar
with
NVIDIA
in
terms
of
modeling,
for
example,
protein
structure,
as
well
as
sort
of
NVIDIA
as
a
whole
with
its
footprints
in
AI.
But
there
are
still
a
lot
of
researchers
that
aren't
aware
that
we
accelerate
genomics
tools,
make
those
open
source.
That's
called
Parabricks,
if
you're
interested.
And
then
we
have
lots
of
tools
for
single
cell
as
well
as
as
well
as
really
just
accelerating
fundamental
data
science.
If
if
you
use
Scikit
Learn,
for
example,
check
out
check
out
Rapids.
Eleanor Howe
Okay.
Yeah,
and
the
single
cell
stuff,
like
my
team
has
seen
how
much
you
can
accelerate
single
cell
analysis
using
some
of
the
NVIDIA
tools.
That
is
a
real
thing
there.
Ben Busby
Yeah,
no,
I
mean
it's
it's
you
know
a
couple
hundred
times
faster.
So
really
you
can
be
talking
about,
you
know,
up
to
a
hundredfold
cost
reduction
if
you're
doing
a
lot
of
analysis.
Eleanor Howe
So
yeah,
that
cost
reduction
and
also
just
being
able
to
do
the
work
in
the
first
place
at
all.
Sometimes
you
just
you
can't
wait
that
long
to
get
your
analysis.
You
have
to
get
you
have
to
get
it
tomorrow.
Ben Busby
Yeah,
and
that's
that's
something
I'm
hoping
we'll
talk
about
a
little
bit
later,
is
really
sort
of
being
able
to
push
into
new
fields
of
analysis,
really
at
this
sort
of
new
scale
of
biology
that
we're
living
in
now.
Eleanor Howe
So
then
what
about
do
you
want
to
talk
a
little
bit
about
like
what
happens
when
people
don't
like
aren't
aware
of
the
tools
that
are
published
by
other
researchers,
by
companies
like
NVIDIA?
Like
what
happens
to
them?
Ben Busby
Well,
I
mean,
I
think
it's
always
challenging
to
keep
up,
right?
I
mean,
this
is
a
a
battle
that
scientists
have
fought
for
the
last
hundred
years
at
least,
you
know,
if
not,
if
not
more
than
that,
of
course.
So
that
said,
I
think
things
are
obviously
accelerating
faster
and
faster.
And,
you
know,
I
was
in
an
interesting
meeting
on
Friday
and
you
know,
with
a
large
group
of
people,
and
they
were
asking
how
to
keep
up
on
things.
I
think,
I
think
it's
it's
challenging,
but
I
think,
you
know,
really
thinking
about
how
we
take
really
sort
of
mixed
data
in
in
relatively
flexible
ways
and
put
it
into
models.
I
think
that's
really
important.
And
I
think
NVIDIA's
pivoted
a
little
bit
in
the
last
few
years
in
this
sense.
We're
really
starting
to
distribute
sort
of
libraries
and
fundamental
tools
for
this
kind
of
thing.
I
mean,
I
think
in
many
senses,
NVIDIA
is
really
a
sort
of
a
picks
and
shovels
company.
And
particularly
now
in
the
biological
sciences,
we're
pivoting
towards
that,
making
things
more
flexible.
For
example,
a
lot
of
people
are
aware
of
a
platform
called
BioNemo,
where
we
host
a
whole
bunch
of
open
models,
and
now
we're
starting
to
distribute
recipes
for
those
open
models,
which
I'm
particularly
excited
about.
And
then
sort
of
the
vision
is
to
enable
the
community
to
pass
recipes
around
for
tuning
some
of
these
models.
And
I
think
to
me,
that's
really
cool.
The
idea
that
we
can
enable
scientists
to
sort
of
really
be
part
of
an
ecosystem.
And
I
think
that's
the
kind
of
thing
that
will
really
enable
awareness
of
new
tools.
Eleanor Howe
Right.
I
love
that.
Yeah.
And
free
tooling
is
what
bioinformatics
was
built
on
from
the
very
beginning
was
people
building
these
systems
and
sharing
them,
and
then
they
become
de
facto
standards
because
they
were
good,
they
worked,
they
did
the
job.
So,
do
you
want
to
talk
about
like
a
real
workflow
change
that
maybe
you've
seen
HaploBlocks And Local Ancestry Context
Eleanor Howe
where
you
know
new
tooling
just
like
made
a
huge
difference?
Maybe
some
of
your
CMU
work?
Ben Busby
Sure.
So
I've
been
working
on
sort
of
a
side
project,
and
you
can
find
it
at
the
haploblocks.org
website.
It's
something
that
I
think
is
particularly
cool,
and
it's
it's
really
just
a
local
ancestry
tool.
So,
so
what
it
does
is
it
puts
SNPs
and
other
variants
in
a
local
genomic
context.
And
so
I
think
this
is
particularly
important
because
local
genomic
context
could
be
quite
nuanced.
You
know,
for
one,
say,
50
to
100,000
base
pair
chunk
of
the
genome
between
recombination
sites,
you
might
be
Portuguese
type
four
on
one
copy
and
German
type
eight
on
the
other
copy.
We've
simply
built
a
software
to
elucidate
that.
So
then
what
you
can
do
is
take
SNPs
in
their
particular
genomic
background
and
mix
them
with
other
SNPs
in
their
genomic
backgrounds
and
then
put
those
into
models.
So,
and
I
think
this
is
particularly
important
when
we
look
at
complex
diseases
and
particularly
complex
diseases
in
admixed
individuals.
And
we're
all
admixed
to
to
some
extent.
Eleanor Howe
Right.
And
so
this
is
so
this
is
this
is
something,
is
this
something
that
any
like
individual
could
use?
Like,
well,
okay,
let's
say
an
individual
who
codes
like
me,
could
I
put
my
ancestry
information
in
there
and
just
get
a
readout
or
put
my
SNP
data
from
a
SNP
chip
or
something
in
there
and
get
a
readout
on
like
where
all
my
blocks
are?
Ben Busby
So
it
depends
on
what
you
mean
by
a
readout.
Like,
yes,
in
in
theory,
you
could,
I
mean,
if
especially
if
you
can
code,
you
could
certainly
put
your
own
genome
in
there
and
then
compare
to
thousand
genomes.
Uh
and
you
could
see,
you
know,
what
your
background
looks
like
in
the
context
of
thousand
genomes,
and
then
you
could
take
all
the
SNPs
that
you
have
in
ClinVar
and
see,
you
know,
how
those
SNPs
are
contextualized,
etc.,
etc.
So
that
could
be,
you
know,
that's
something
that
you
could
do
that
that
might
be
interesting.
But
I
think
again,
the
more
interesting
thing
you
can
do
is
then
you
could
take
those
and
and
we
have
a
way
to
translate
these
SNPs
in
context
or
just
the
context
themselves
into
binary
strings,
and
those
are
really
easy
to
put
into
models.
So
that's
something
that
we're
particularly
excited
about.
And
one
thing
that
I
don't
think
we
have
a
whole
lot
of
time
to
talk
about,
but
we're
we're
really
interested
not
only
in
using
this
for
humans
and
understanding
human
biomedical
things,
but
also
agriculture.
So
thinking
about
taking
these
types
of
approaches
to
to
plants
as
well.
Eleanor Howe
That
would
be
really
incredible.
Actually,
so
it
can
it
can
deal
with
the
multiploidy
and
the
complexity
of
plant
genomes?
Is
or
is
is
that
where
you
where
you
came
up
with
this
system
in
the
first
place?
Is
because
plants
are
really
difficult.
Ben Busby
Well,
sort
of
both
at
the
same
time.
So
we've
had
to
make
a
few
modifications
in
in
collaboration
with
a
company
called
Vail
Genomics
to
think
about
sort
of
the
nuances
of
plant
genomes.
But
also,
you
know,
actually,
this
is
a
quote
from
somebody
at
Vail,
and
he
says,
you
know,
for
humans,
genome
graphs
are
the
future.
And
I
and
I
believe
that,
but
I
think
for
plants,
he
said,
graphs
are
the
now.
And
so
that's
that's
something
that
kind
of
resonated
with
me.
And
interestingly
enough,
I
mean,
what
we're
doing
with
these
hashes
is
sort
of
making
approximate
dimensionally
reduced
graphs,
but
we're
seeing
that
with
graph
structures
all
over
the
place,
not
just
in
genome
graphs,
but
in
knowledge
graphs.
And
I
think
that's
that's
sort
of
something
that
you
know
we're
we're
sort
of
moving
towards
in
kind
of
an
information
theory
sense
in
biology.
Eleanor Howe
And
you've
been
you've
been
a
knowledge
graph
advocate
for
a
long
time.
I
remember
that.
Ben Busby
Uh
yeah,
I
I'm
a
particular
fan
of
knowledge
graphs,
and
I
I
think
a
lot
of
people
think
that
they
have
a
particularly
sort
of
pointed
relevance
now
in
the
age
of
in
the
age
of
models.
Eleanor Howe
Okay.
Well,
for
people
who
are
interested,
the
haplobox
haploblocks
is
published
on
haploblocks.org.
You
can
check
out
the
paper
and
the
data
there.
The Real Bottleneck Is Longitudinal Data
Eleanor Howe
So
then,
okay,
moving
on,
let's
talk
about
bottlenecks.
Every
time
a
new
technology
is
advanced,
the
bottleneck
just
moves.
It
moves
somewhere,
it
doesn't
disappear.
So,
where
would
you
say
the
bottlenecks
are
in
genomic
workflows
today?
Compute,
data
movement,
something
else?
What
do
you
think?
Ben Busby
Um
I
I'd
say
it's
it's
data
and
particularly
longitudinal
data
right
now.
So,
for
example,
you
know,
I
mean,
I
can
make
a
really
cool
knowledge
graph,
for
example,
with
,
you
know,
HaploBlocks
and
whatnot,
or
or
hash
things
into
a
model
with
HaplaBlocks
from
genomic
data.
But
what
I
really
want
to
do
is
be
able
to
layer
on
using
a
model
transcriptomic,
proteomic
imaging
data,
and
getting
sort
of
reliable
longitudinal
data
for
groups
of
individuals,
typically
in
a
disease-specific
way,
is
a
very
challenging
thing
to
do
right
now.
And
that's
one
of
the
reasons
I'm
so
interested
in
data
federation
from
biobank
style
data,
because
I
think
where
we
really
want
to
be
is
thinking
about
being
able
to
move
data
from
one
hospital
to
another,
from
one
cancer
center
to
another,
from
one
biobank
to
another,
but
not
raw
data
being
able
to
move
weights
and
embeddings
and
models
such
that
we
can
have
more
power
for
this
type
of
specific
longitudinal
analysis.
Eleanor Howe
Okay.
And
so
that
those
federation
decisions
are
made,
you
know,
by
more
often
by
leadership
teams
than
by
the
scientists
who
want
to
access
the
data.
So,
like
where
do
you
think
these
leadership
teams
are
making
good
decisions
about
you
know
connecting
science
to
the
hardware,
about
investing
in
infrastructure,
investing
in
federation?
Like,
where
do
you
think
people
are
doing
the
right
things?
And
where
do
you
think
that
they're
really
not
very
much?
Ben Busby
Well,
I
mean,
I
think
there's
a
there's
an
onus
on
the
the
scientific
and
technical
community
first.
Now,
I
don't
want
to
sort
of
you
know
wash
everything
with
rose-colored
glasses
and
say
if
we
build
a
scientific
infrastructure,
then
they
will
come.
I
I
don't
think
that's
true.
There
still
needs
to
be
driving
forces
from
a
bunch
of
places.
That
said,
one
of
the
reasons
I'm
so
excited
about
the
NV
Flare
platform
that
I
work
with
a
lot
is
because
basically
we
can
make
the
model
embedding
simple
enough
such
that
sort
of
the
policy
people,
the
legal
people,
you
know,
that
that
think
hard
about
the
social
and
ethical
aspects
of
distributing,
you
know,
model
weights,
model
embeddings,
can
understand
exactly
what's
going
in
there
and
understand
exactly
what
facets
of
the
data
are
being
shared.
And
that
to
me
is
really
exciting.
And
there's
there's
other
people
that
I
work
with
that
other
people
work
with
doing
some
really
cool
stuff
with
latent
space
representation,
same
basic
ideas.
So
I
think
what
we
need
to
do
is
present
a
range
of
viable
options
to
the
folks
that
do
the
sort
of
legal
and
ethical
type
of
stuff
with
the
biobanks
and
also
hospital
data.
But
I
think
there's
another
onus
upon
us,
which
is
also
to
show
them
that
there's
a
there
there,
right?
And
this
is
the
biggest
thing
that
we
are
missing
in
biomedical
science
is
showing
that
we
can
build
effective
models.
And
we're
starting
to
see
really
effective
models
come
out
of
a
number
of
shops.
They're
being
licensed
to
pharma,
they're
being
licensed
to
diagnostics,
they're
being
developed
by
diagnostics
companies.
But
I
would
say
often
these
models
come
from
counterintuitive
places.
So,
for
example,
often
we
see
small
disease-specific
models
working
very,
very
effectively.
And
I
think
this
is
something
that
is
somewhat
specific
to
biomedical
science
and
is
kind
of
counterintuitive
to
some
sort
of
prevailing
logic
that,
you
know,
sort
of
longer
context,
larger
and
larger
models
are
always
better.
So,
and
that
to
me
is
really
interesting.
But
when
you
go
back
and
you
think
about
the
genomics
of
disease,
when
you
think
about
the
genomics
of
something
like
schizophrenia
versus
cystic
fibrosis,
those
are
really
two
very
different
diseases
on
a
genomic
level
and
may
require
a
fundamentally
different
model
structure.
Eleanor Howe
Right.
So
you're
saying
that
the
one
size
fits
all
foundational
models
that
are
supposed
to
model,
quote,
diseases
don't
work.
It's
that
we
have
to
be
specific
about
what
you're
trying
to
model.
Ben Busby
I
would
say
that
I
am
saying
that,
but
not
so
deterministically.
I
think
probabilistically
there's
like
most
of
the
evidence
is
pointing
to
the
efficacy
of
swarms
of
small
models,
but
I
am
not
saying
it's
impossible
that
in
six
months
Anthropic
Why Swarms Of Small Models Win
Ben Busby
will
blow
us
all
out
of
the
water
and
we'll
be
working
on
something
completely
different
because
they
will
have
figured
all
of
this
out.
Eleanor Howe
Good
luck
to
them.
That
would
be
terrifying
but
useful.
Yes.
Ben Busby
There
you
go.
Yeah,
terrifying
but
useful
is
a
really
great
way
to
say
that
amazing.
Eleanor Howe
So
then
where
do
you
see,
let's
take
the
opposite
thing,
some
of
these
tools
potentially
really
powerful.
Where
do
you
where
do
you
see
it
happening
that
access
to
better
tools
is
not
translating
into
better
outcomes?
Ben Busby
Oh,
that's
a
great
question.
And
I
think
the
answer
is,
and
this
is
almost
an
old
adage,
but
when
people
use
fancier
and
fancier
tools
to
ask
the
same
questions,
they've
been
asking
for
a
long
time.
And
I
think
really
that
happens
when
people
fail
to
do
the
work
to
contextualize
biology,
right?
If
you
have
a
particular
SNP
in
Klinbar
and
that
there's
a
penetrance
measured
in,
you
know,
a
particular
population,
you
know,
in
Germany
or
the
United
States
or
something
like
that,
one
thing
that
has
been
very
clear,
particularly
with
papers
just
in
the
last
few
months,
is
that
may
not
be
the
same
level
of
penetrance.
You
may
not
get
the
same
presentation
with
folks
that
are
in
Pakistan,
for
example,
at
least
on
a
sort
of
social
societal
level.
And
then
I
would
say
that,
you
know,
we
we
often
completely
ignore,
you
know,
things
that
have,
you
know,
two,
three,
four,
five
genomic
loci
that
are
really,
you
know,
important
in,
you
know,
manifestation
and
presentation
of
disease.
So
I
think
those
are
things
that
we
really
need
to
spend
time
thinking
about
and
do
the
work
to
address.
Eleanor Howe
Okay,
great.
Subscribe And Send Topic Ideas
Announcement
Are
you
enjoying
the
conversation?
We'd
love
to
hear
from
you.
Please
subscribe
to
the
podcast
and
give
us
a
rating.
It
helps
other
people
find
and
join
the
conversation.
If
you've
got
speaker
or
topic
ideas,
we'd
love
to
hear
those
too.
You
can
send
them
in
a
podcast
review.
Eleanor Howe
So
then
sometimes
what
we
see
is
that
companies
will
buy
huge
server
farm,
they'll
buy
a
whole
bunch
of
GPUs,
they'll
license
a
whole
bunch
of
time
on
clouds
and
provide
their
teams
suddenly
a
bunch
more
compute
than
they
had
before.
So
when
that
happens,
Compute Portfolios: On-Prem Plus Cloud
Eleanor Howe
what
actually
changes
for
them?
Like,
do
they
do
they
actually
make
good
decisions
about
how
to
use
that?
Have
we
what
are
you
seeing?
Ben Busby
Well,
so
as
the
scale
of
biology
changes,
right,
people
often
want
more
and
more
compute.
And
typically
we
don't
see
researchers
getting
flooded
with
compute.
I
mean,
if
you're
in
an
academic
institution
and
and
they
they
buy
a
bunch
of
GPUs,
that's
great.
But
if
you
as
a
biologist
don't
sort
of
lay
your
claim
to
some
of
them,
your
physics
colleagues
are
probably
going
to
gobble
all
of
them
up.
So
that
said,
I
think
it's
important
to
do
creative
things,
right?
With
this
amount
of
compute,
we're
able
to
do
things
like
look
at
five
genomic
sites
simultaneously
and
integrate
phenotype
and
integrate
proteomics.
And
these
are
the
kinds
of
creative
things
we
should
be
doing
with
the
new
compute
we
have,
right?
I
mean,
data
preparation
is
one
thing,
and
that
is
super
important.
And
data
generation
is
gonna
be
a
huge
bottleneck.
So
we're
gonna
need
automated
labs
for
that
kind
of
thing,
et
cetera,
et
cetera.
Ben Busby
But
beyond
sort
of
data
generation
and
and
preparation,
which
should
be
really
reproducible,
we
need
to
really
be
creative
about
the
kinds
of
questions
we're
asking.
And
we
really
need
flexible
frameworks
to
be
able
to
ask
those
questions.
And
I'm
just
gonna
be
blunt
about
it.
I
mean,
when
I
see
either
academics,
you
know,
startup
companies
being
really
successful,
they
usually
have
a
small
amount
of
on-premise
compute,
and
then
an
ability,
or
in
the
case
of
a
large
academic
institution,
a
fairly
large
amount
of
on-premise
compute,
and
then
being
able
to
burst
up
into
cloud
using
small
cloud
providers
as
well
as
large
cloud
providers.
And
so
I
think
really
being
able
to
diversify
compute
portfolios,
thinking
about
what
you
need,
having
a
strategy
there,
talking
to
folks
about
your
IT
strategy,
really
doing
the
math,
doing
the
homework.
That's
what
is
gonna
make
sense.
And
that's
what's
gonna
make
a
competitive
advantage
for,
say,
diagnostics
companies
moving
forward.
Eleanor Howe
So,
what
I'm
hearing
from
you
is
that
science
is
not
dead
and
we
still
need
scientists.
Would
you
say
that
that's
true?
Ben Busby
Yeah,
no,
I
mean,
we
we
really
need
people
to
ask
creative
questions.
Again,
if
you're
if
you're
just
kind
of
phoning
it
in
at
the
end
of
the
day
and
you're
doing
technical
things
with
GPUs,
but
you're
not
asking
the
questions
that
are
pushing
the
envelope
in
science
and
biomedicine,
you're
not
changing
the
world,
right?
And
so
I
think
that's
the
thing.
We
need
scientists
to
ask
creative
questions.
And
then
we
also
need
to
figure
out
how
to
help
models
communicate
with
professionals
in
the
field.
And
then
nowhere,
I
can't
think
of
any
other
field
where
that's
more
important
than
biology,
both
medicine
as
well
as
agriculture,
because
you
have
experts
in
the
field.
Case
of
agriculture,
literally,
but
in
in
medicine,
you
know,
you
have
people
that
have
been
oncologists
for
40
years,
they
have
a
lot
of
model
Turning Expert Judgment Into Models
Ben Busby
weights
in
their
head,
right?
And
so
figuring
out
how
to
translate
the
brain
of
somebody
who's
been
an
oncologist
for
40
years
into
a
model,
I
think,
is
one
of
our
next
frontiers
that
we
want
to
be
working
on.
Eleanor Howe
Wow.
Okay,
that
is
a
big
project.
How,
how,
where,
how
would
you
even
start
with
that?
Ben Busby
So,
well,
for
one
thing,
I
think
we
we
need
really
better
software
interfaces
for
professionals
who
are
non-technical
to
work
with
models
and
ways
to
to
collect
those
sorts
of
weights.
And
and
those
are
things
you
know,
some
people
in
the
community
are
working
on,
we're
working
on,
but
I'm
really
excited
to
see
where
that
field,
where
that
field
evolves.
Eleanor Howe
And
when
you
say
different
interfaces
for
folks,
are
you
talking
about
the
chatbot
LLM
interfaces
or
do
you
mean
something
else?
Ben Busby
Well,
I
mean,
I
think
chatbot
LLM
interfaces
were
were
a
watershed
moment,
right?
I
mean,
it
got
sort
of
the
the
world
and
the
internet
a
lot
of
sort
of
bulk
human
knowledge,
you
know,
sort
of
into
the
world
of
AI.
We're
gonna
need
really
sort
of
more
nuanced
interfaces
for
thinking
about
how
to
get
medical
professionals,
agricultural
professionals,
other
scientific
professionals
to
get
sort
of
their
knowledge
into
these
models.
And
they're
probably
going
to
be
sort
of
small
specific
models.
So
we
won't
be
able
to,
they're
more
likely
to
be
small
specific
models,
so
we
won't
be
able
to
rake
information.
We'll
have
to
be
very
selective
of
it
about
information,
and
then
also
do
pruning
on
information
that's
not
relevant.
And
that's
another
thing
that
we
really
need
to
be
putting
some
work
into,
and
something
that's
enabled
specifically
by
knowledge
graphs.
Eleanor Howe
Okay.
So
it
sounds
like
that
you
think
that
you
know
the
future,
the
next
three
to
five
years,
maybe
knowledge
graphs
are
going
to
Knowledge Graphs As Model Guardrails
Eleanor Howe
be
a
big
player.
Ben Busby
Yeah,
I
think
knowledge
graphs
are
likely
to
be
a
very
big
player.
I
mean,
they're
they're
really
a
sensible
way
to
organize
relatively
static
biological
information
and
sort
of
give
checkpoints
with
models
as
well
as
both
validate
and
prune
information.
One
thing
that
I'm
working
on
very
actively
is
the
idea
of
democratizing
knowledge
graphs.
I
think
you
know
it's
really
important
to
enable
biological
researchers
everywhere
to
build
knowledge
graphs.
And
we're
we're
actually
running
a
competition
with
the
AWS
Open
Data
Program
in
the
fall
on
building
agents
to
build
knowledge
graphs.
So
if
that's
something
you're
interested
in,
you
please
keep
an
eye
out
for
that.
It'll
be
all
over
social
media,
et
cetera.
Eleanor Howe
And
what
is
the
name
of
that
again?
Can
you
say
that
again?
Ben Busby
Thanks,
Ellen.
That's
called
the
Bioagent
Knowledge
Graph
Construction
Challenge.
Eleanor Howe
Okay,
great.
All
right.
So
then
let's
say
you're
advising
a
biotechnology
leadership
team
for
a
new
startup.
Yeah.
You
want
to
talk,
you
want
to
tell
them
like
what
do
you
tell
them
about
their
computing
strategy?
What
is
the
piece
of
advice
that
is
different
now
than
it
was,
say,
a
year,
two
years
ago,
that
you
would
give
them?
Ben Busby
Yeah.
So
for
one,
you
need
a
computing
strategy.
And
two,
I
think
it's
important
to
do
the
math.
But
as
I'll
say
again,
I
mean,
I
think
companies
are
most
successful
when
they
have
some
on-prem
balance
with
being
able
to
burst
onto
cloud.
This
is
where
we
see
people
being
very
successful
and
being
selective
about
what
they
buy.
For
example,
in
the
biological
sciences,
the
RTX
6000
chip
is
amazing.
I
mean,
it
does
so
Democratizing Knowledge Graph Construction
Ben Busby
much
of
all
of
this
stuff
we
talked
about,
you
know,
from
processing
of
genomics
data,
being
able
to
run
models,
et
c., etc.,
being
able
to
do
some
tuning,
all
of
that
stuff.
That's
fantastic.
So,
for
example,
you
could
have
sort
of
a
workforce,
you
know,
four
to
eight
RTX
6000s.
And
then
if
you
need
to
do
some
training
of
a
model,
you'd
want
to
burst
up
onto
cloud.
But
but
I
think
there's
something
more
fundamental
here,
which
is
so
we're
starting
to
see
this
emerging
economy
of
licensing
models,
right?
I
mean,
if
you
look
at
you
know
big
business
news,
et
cetera,
et
cetera,
you
see,
you
know,
large
companies
signing
licensing
agreements
with
model
building
shops.
So,
what
I
would
say
to
startup
companies
is
even
if
you
plan
to
keep
the
whole
thing
proprietary
and
develop
a
drug
yourself
or
whatever
it
is,
build
models
that
other
people
would
want
to
license,
or
build
models
that
you
would
want
to
license.
I
mean,
really,
you
know,
you
should
be
asking
questions
that
other
people
will
be
so
excited
about
in
terms
of
them
moving
the
needle
that
they
would
be
willing
to
pay
you
for
them.
Eleanor Howe
And
so
that
implies
that
you
understand
the
landscape
of
the
science
well
enough
to
know
what's
missing.
Ben Busby
Absolutely.
Again,
I
mean,
I
think
maybe
this
is
a
little
bit,
I
don't
know,
iconoclastic
or
something.
But
in
fact,
I
think
for
computational
scientists,
right,
the
next
three
to
five
years
is
gonna
be
an
incredibly
busy
time
because
we're
gonna
be
working
on
a
lot
of
these
things.
We
have
a
lot
to
work
out
in
terms
of
doing
modeling,
but
more
importantly,
validating
models,
right?
And
that
that
comes
with
characterizing
benchmarks,
and
that
takes
a
lot
of
scientific
work
to
come
up
with
benchmarks
that
make
sense.
And
really
at
this
point,
to
benchmark
models
specifically
and
effectively,
we
need
clever
biologists
who
can
come
up
with
benchmarks
that
not
only
make
sense,
but
also
push
the
limit
of
science.
And
that's
a
really
tricky
thing
to
do,
right?
Not
just
come
up
with
benchmarks
for
the
science
we
did
10
years
ago,
but
come
up
with
benchmarks
for
the
science
of
the
future.
Eleanor Howe
And
I
think
one
of
the
things
that
is
really
important
is
is
understanding
that
there's
also
other,
you
know,
we
talk
about
models,
there's
old
technology
that
still
works.
And
one
of
the
baseline
things
we
need
to
do
is
to
make
sure
that
any
new
model
that
is
computationally
expensive
and
hard
to
make
is
actually
better
than
the
existing
models
that
we
have.
This
is
this
is
a
beef
of
mine,
is
that
people
jump
straight
to
a
complex
neural
network
when
an
actual
linear
model
would
work
just
fine.
And
a
lot
of
the
time,
my
team
has
done
these
benchmarks
showing
that
that
that's
the
case.
And
so
I'm
super
into
benchmarks
as
well.
I
agree
completely.
Ben Busby
So
I
would
say,
I
mean,
XG
Boost
is
still
alive
and
well.
And
at
the
beginning,
I
mentioned
this
thing.
Check
out
XG
Boost
on
Rapids,
it's
about
45
times
faster
on
GPU
than
CPU.
That
said,
I
mean,
just
to
echo
your
point,
we
need
benchmarks
that
will
compare
a
linear
model
versus
XG
Boost
plus
a
small
gen
model
or
versus
a
small
gen
model
versus
a
large
gen
model,
right?
You
should
absolutely
be
able
to
compare
apples
to
apples
there
and
say,
I'm
gonna
use
the
most
efficient
model
that
allows
me
to
push
science
forward.
Eleanor Howe
Yeah,
yeah,
yeah,
absolutely.
Okay,
we're
on
the
same
page
here.
I
love
it.
So
then
what
what
do
you
wish
every
bioinformatician
now
understood
about
the
you
know,
the
latest
hardware
and
infrastructure
that
could
change
how
they
work?
I
think
you
touched
on
it
a
bit
already.
Maybe
this
is
a
little
redundant.
Ben Busby
What
do
I
wish
every
bioinformatician
understood?
I
mean,
I
I
would
hope
that
every
bioinformatician
would
think
about
how
to
take,
and
I
think
they
many
are
how
to
take
multiple
complex
data
sites,
the
data
types,
and
use
them
in
concert
to
be
able
to
peel
the
onion
of
a
scientific
problem,
right?
We
know
that
all
ICD
10
codes
typically
are
umbrellas
for
a
whole
variety
of
etiologies
that
lead
to
fairly
consistent
symptomology,
right?
And
so
Benchmarking: Proving New Beats Old
Ben Busby
being
able
to
come
up
with
a
model
that
starts
to
peel
those
layers
of
the
onion
back
and
point
towards
pipelines
where
things
are
treatable,
I
think
that's
that's
where
I
wish
every
bioinformatician's
head
was
at.
Eleanor Howe
Yeah,
makes
a
lot
of
sense.
I
don't
what
do
you
think
of
my
hypothesis
that
the
the
rise
of
coding
agents
to
assist
in
coding
just
means
that
bioinformaticians
get
to
be
biologists
again?
Ben Busby
Wow,
I
love
the
question.
You
know,
I
saw
this
quote
somewhere
on
social
media
that
I
thought
was
great,
and
,
and
I'm
sad
I
don't
have
attribution
for
it.
How
useful
you
think
language
models
are
really
often
depends
on
how
much
you
code.
And
I
think
that's
something
that
I've
I've
seen
be
consistent,
and
so
many
of
my
friends
who
are
professional,
like
you
know,
I
mean
PhD
computer
scientist
type
people
have
just
dove
in,
right?
They
love
this
stuff.
I
love
this
stuff.
Oh
my
gosh,
it's
so
great,
you
know,
being
able
to
prototype
so
fast.
And
and
honestly,
I
think
you're
absolutely
right.
And
I
would
go
even
a
step
further
and
say
it
doesn't
even
allow
us
to
be
biologists,
it
forces
us
to
be
biologists
because
if
we
can't
use
all
this
cool
stuff
to
ask
better
questions,
then
the
world
is
gonna
go
work
on
something
else.
And
you
know,
I
mean,
don't
get
me
wrong,
I
like
animals,
but
I
hope
that's
not
cat
videos.
Eleanor Howe
Yeah,
that
that's
what
I've
observed
from
my
team
is
that
they
just
get
to
do
more
science
in
the
same
amount
of
time.
You
when
with
the
coding
agents
helping
them
write
the
code
they
need
to
write,
they
can
focus
on
learning
the
biology
rather
than
fussing
with
the
details
of
some
irritating
syntax
of
a
poorly
written
language.
Ben Busby
Yeah,
and
I
think
agents
are
just
gonna
amplify
that
trend,
right?
And
yeah,
so
like
then
for
basic
processes,
not
the
the
ones
that
require
the
creativity
I
was
talking
about,
not
the
big
combinatorial
things,
but
but
basic
things,
right?
You
know,
like
BWA
to
deep
variant
or
G
A
T
K.
I
mean,
we
already
have
agents
that
wrap
Parabricks
tools
to
do
that,
and
so
anyone
should
be
able
to
do
the
sort
of
basic
foundational
bioinformatics
processes
on
their
own
at
this
point.
Eleanor Howe
Yeah,
well.
So
then
what
do
you
say?
One
last
question,
which
is
you
know,
is
there
something
that
NVIDIA
is
building
that
people
don't
know
about,
that
everyone
is
sleeping
on,
that
you
think
they
should
know?
Ben Busby
Well,
there's
there's
a
number
of
things,
really.
One
thing
I'm
particularly
excited
about
from
a
hardware
point
of
view
is
the
rise
of
CPUs
that
are
that
that
go
with
the
GPU,
and
really
there's
zero
latency
in
terms
of
transfer
between
those
CPUs
and
and
GPUs.
So
you
really
can
have
your
cake
and
eat
it
too.
Um
and
then
also
those
GPUs
tend
to
have
a
lot
of
RAMs.
So
you
can
really
get
deep
into
biological
questions.
I
think
what
I'm
most
excited
about
in
the
the
the
software
sense,
although
I'm
hugely
biased
because
I
work
on
it,
is
the
idea
of
model
validation
and
thinking
about
all
right,
how
do
we
build
frameworks
that
allow
people
to
compete
models
against
each
other
in
a
fairly
flexible
way
such
that
they
know
how
to
compare
A
to
B,
whether
it's
the
same
model,
tuned
to
some
extent,
whether
they're
totally
different
models,
et
cetera,
et
cetera,
or
different
models
that
special
specialize
in
different
etiologies
of
disease.
Eleanor Howe
So
if
people
can
run
their
own
pipelines
without
necessarily
knowing
how
to
code
in
the
first
place,
like
I
agree,
that
is
absolutely
a
thing.
But
biologists
are
now
much
more
empowered
to
do
things
on
their
own.
So
what
happens
to
all
of
us
Coding Agents And Bioinformatics Careers
Eleanor Howe
bioinformaticians
who
spend
so
much
time
learning
to
code?
What
are
we
good
for?
Ben Busby
Actually,
I
think
counterintuitively,
over
the
last
few
months,
I'm
spending
more
and
more
time
referring
companies
to
small
bioinformatics
engineering
shops,
because
they
might
have
one
or
two
bioinformaticians
who
can
use
agents,
spin
up
prototypes,
do
a
proof
of
concept.
That's
amazing,
right?
But
for
hardcore
engineering
and
efficiency,
particularly
when
like
somebody
like
NVIDIA
gets
involved,
and
we
can
give
great
advice,
but
we
best
our
advice
is
best
given
to
professional
engineers.
There
need
to
be
a
cadre
of
working
people
that
are
working
engineers,
working
bioinformaticians
who
can
develop
a
stable
reproducible
code
base,
particularly
in
biomedicine,
because
we
are
talking
at
the
end
of
the
day
about
giving
reproducible
clinical
decision
support
that
is
used
to
treat
patients.
Eleanor Howe
Amazing.
So
we
have
jobs
still.
Is
that
what
you're
saying?
Ben Busby
There's
a
useful
thing.
I'm
not
saying
you
have
jobs
still.
I'm
saying
those
jobs
should
expand.
Ah
my
kids
aren't
quite
old
enough
yet,
but
if
my
kids
were,
you
know,
at
you
know,
sort
of
finishing
college
and
thinking
about
computational
biology,
I
mean,
I
see
this
as
a
growth
industry.
Eleanor Howe
Amazing.
Final Takeaways And Farewell
Eleanor Howe
Thanks
so
much
for
joining
me,
Ben.
Um,
I
really
appreciate
you
taking
the
time
and
it's
been
a
really
interesting
conversation.
It's
so
good
to
hear
your
perspective.
Ben Busby
Absolutely.
Yeah.
Thanks,
,
thanks
for
inviting
me.
And
,
yeah,
I
look
forward
to
to
talking
soon
and
seeing
everybody
at
various
meetings
over
the
next
year.
Eleanor Howe
It's
been
a
pleasure
having
you
on
the
trenches.
Ben Busby
Thanks
a
lot.