So, if only 2% of the genome consists of genes, what is the difference between a gene and the rest of the genome? Is this a "junk DNA" type thing? If so, I'm not asking about a preference for non-junk vs. junk DNA.
If we take a phenotype based view of the gene as an allele, then the gene would could include the upstream promoters and binding domains, sequence that affect mRNA editing, etc. You can consider the gene to be more than just the DNA between the start and stop codons for protein translation.
DNA that does not directly code for proteins can still be functional DNA. DNA that does not affect the fitness of the organism in any discernable manner is junk DNA, and I think almost all of that is non-coding DNA. About 90% of the human genome is accumulating mutations at a rate consistent with not having function that impacts fitness.
To further confuse this situtation, scientists like the ENCODE consortium come in and try to redefine what function is. What they tried to do is conflate the terms "functional" and "does something". Those are not the same thing. The trash in your kitchen trash can releases odor molecules into the air. It does something. However, it is still trash. The same for junk DNA. It is DNA that would not affect the fitness of the organism if it were removed just as your kitchen continues to work just fine after you dump the trash.
What is a histone and a transcription unit?
A histone is a protein that has DNA wrapped around it. Think of it like the sticks at a library that they wrap maps or newspapers around. It is a way of physically compacting DNA.
Transcription units are areas with a lot of genes in them. Think of it as the Earth when you fly over in an airplane. The transcriptional units are the towns where there are more lights, with each house being analogous to a gene.
From the sound of the term, the "transcription unit" may be closer to what I'm asking. From that term it sounds as if there is some means of dividing the genome into units. I'm asking if insertion prefers the start of a unit or if it will insert anywhere in the unit.
When MLV does insert into genes, it tends to insert at the beginning of the gene, as shown in Figure 2 of the paper I mentioned before:
Retroviral DNA Integration: ASLV, HIV, and MLV Show Distinct Target Site Preferences
Even then, if only 45% of the genome is made up of these units, is the rest of the DNA just a random string of nothing? With no discernible difference between one part and other?
It is made up of DNA bases where the retrovirus can insert. However, it probably has very little to no function as it relates to fitness.
Interesting, but basically irrelevant to my question. Again, I'm asking if, by chance, the insertion happened in the usable part of DNA, is the insertion most likely to happen at the start/end of a "unit", or does it still just happen randomly anywhere?
With MLV it tends to happen at the beginning of a gene if it does insert into a gene. However, the bulk of insertions happen outside of genes, and there are also tens of thousands of genes. IOW, this mechanism isn't able to create hundreds of thousands of orthologs between chimps and humans.