Transients in the Palomar Observatory Sky Survey

has Villarroel indicated at any point that they have reached out to the "custodians of the plates" to ask for access and if so have they reported the response?

I don't know. Maybe others do, but I don't remember Beatrix Villarroel suggesting this (which doesn't mean she hasn't).

I guess many other astronomers have done work based on/ informed by DSS and/or SSS, and probably the earlier glass POSS1 copies (perhaps more so pre-digitisation). If this is the case, it might be argued that those astronomers were not obliged to refer back to the original material (conjecture, maybe some were, I don't know). But I feel B.V. et al.'s transient papers are perhaps different; their findings are extraordinary and if accepted their implications are far-reaching. Hambly and Blair have raised what might be plausible concerns about reprographic artefacts, which might be tested by examining the POSS1 plates. Maybe other relevant insights would be made while examining the originals.

I don't know anything about access to the POSS1 plates. Even if allowed, a re-examination of the 1st generation plates (presumably "in situ") might be a much more time-consuming, inconvenient and in effect expensive process than the study of the digitised material.
 
Last edited:
I don't think they have actually examined the original plates. However I think they would say they have shown that the transients can't be emulsion artifacts because their distribution has been shown to correlate with the Earth's shadow. But even that logic has been shown to have some questionable assumptions and tenuous rationale.

Snips from my post #72:
To me having not looked at these papers in 2 years nothing seems to have moved on. Like you said all the assumptions made around why these may not be plate artifacts seem flawed either,

1-address the issues the authors themselves highlighted in the 2021 paper
or
2-clearly state the problem or the assumptions being made on those issues in future papers

Completely pointless discussing/building on the initial work without this. Go make some plates again and do a study on that even to further the conversation.
 
Last edited:
I don't know. Maybe others do, but I don't remember Beatrix Villarroel suggesting this (which doesn't mean she hasn't).
to me, it seems a major oversight to not examine the originals, and if she had tried and was rejected that would be an important point to make in order to show that one is trying in good faith to rule out an instrumental artifact. It seemed like she just said "I don't have access to the originals" and then carried on leaving the uncertainty just hanging over the entire work.
 
to me, it seems a major oversight to not examine the originals, and if she had tried and was rejected that would be an important point to make in order to show that one is trying in good faith to rule out an instrumental artifact. It seemed like she just said "I don't have access to the originals" and then carried on leaving the uncertainty just hanging over the entire work.
Properly assessing the original plates -- especially dealing with any significant number of suspected transients -- would have to be done by someone with expertise in glass-plate astrophotography. Probably someone working at the one of the several astronomical archives holding glass plate collections, like Harvard College Observatory or the Carnegie Plate Archives.
 
Properly assessing the original plates -- especially dealing with any significant number of suspected transients -- would have to be done by someone with expertise in glass-plate astrophotography. Probably someone working at the one of the several astronomical archives holding glass plate collections, like Harvard College Observatory or the Carnegie Plate Archives.
And that could in principle be done in collaboration with Villarroel but there's no indication that I have seen that this was attempted.
 
Properly assessing the original plates -- especially dealing with any significant number of suspected transients -- would have to be done by someone with expertise in glass-plate astrophotography. Probably someone working at the one of the several astronomical archives holding glass plate collections, like Harvard College Observatory or the Carnegie Plate Archives.
And may have been a bit of a wild goose chase before. But her most recent paper has identified 203 transients on plate(s) from 1 particular day, 5/3/1954:

External Quote:

To maximize likelihood of interpretable results, we again focused on the highest probability transients (≥ 0.90 probability of being a real transient per the ML model) [10].There were 231 transients meeting this criterion observed within +/- 1 day of a nuclear test(for nuclear testing details, see https://nnss.gov/wp-content/uploads/2023/08/DOE_NV209_Rev16.pdf). Of these, 203 transients (87.9%) were associated with a single nuclear test:Castle Yankee conducted on 5/4/54.
In addition, these transients are supposedly located in particular spots over the Eastern Pacific, off the coast of Mexico:

External Quote:
Cluster 1 was centered in the Pacific ocean to the southwest of Baja California. Clusters 5 and 6, while statistically distinct, were located in two adjacent regions of the Pacific further south than Cluster 1, with both being to the west of southern Mexico and Central America. These two clusters combined represented the location of nearly 65% of all high probability transients identified.
1789752453498.png

https://arxiv.org/pdf/2609.09461

So, there is a plate(s) from a particular day that has 203 transients in a particular area on it. That would substantially narrow down what to look for on the specific original plate(s). As there is a specific plate(s), it should also be easier to track the copy(s) from original to digitization thus providing a complete pipeline that can be checked for transients. Right?
 
What does it mean for a transient to be "over" a particular location on Earth? Aren't these supposed to be in space? Is that point on Earth just the "sub-satellite" point (i.e., draw a line from the supposed location of the transient to the center of the earth and where it intersects the surface)? It seems odd to geolocate a transient happening in space, especially when there's no direct measurement of the 3-d position of the supposed object. A transient by its nature is not repeating, which means no ephemeris can be calculated and an altitude would have to be assumed, no?

sorry, I am not well versed in all of this research, so perhaps these questions have been addressed in the papers and/or this thread. I can be ignored if I am rehashing old ground.
 
What does it mean for a transient to be "over" a particular location on Earth? Aren't these supposed to be in space? Is that point on Earth just the "sub-satellite" point (i.e., draw a line from the supposed location of the transient to the center of the earth and where it intersects the surface)? It seems odd to geolocate a transient happening in space, especially when there's no direct measurement of the 3-d position of the supposed object. A transient by its nature is not repeating, which means no ephemeris can be calculated and an altitude would have to be assumed, no?

sorry, I am not well versed in all of this research, so perhaps these questions have been addressed in the papers and/or this thread. I can be ignored if I am rehashing old ground.

The geolocation thing, as well as the supposed altitude of transient being calculated from the POSS1 plates is from her new pre-print that I linked to. Whether it works or not, I have no idea. But I would think one could try to follow the logic and find a particular area on the plate that the authors believed was above certain location on the Earth. In this case, a number of points in the Eastern Pacific:

External Quote:

Because the shadow deficit is consistent with transients representing reflective objects in Earth orbit and recent findings suggest transients may be due to brief flashes like tumbling objects in orbit exhibit [12.14], we began exploring their potential orbital parameters. We recently developed a modelling approach that permits estimating the altitude of POSS-I transients based on shadow deficit characteristics [12]. Results using two complementary statistical approaches both suggested that transients were between 26,371 - 35,786 km above Earth.

After filtering for high probability transients, results were most consistent with the higher estimate, suggesting geosynchronous orbit. Notably, a significant excess of transients was observed at low declinations [12], potentially indicating orbits near the celestial equator where modern geosynchronous satellites are typically observed [15].

External Quote:

Building on these recent findings, the current work explored the Earth-projected locations over which transients were observed. Specifically, given an assumed orbital altitude of 35,786 km, combined with information regarding the Mount Palomar observatory location and observation parameters (e.g., right ascension, declination), we estimated the latitude and longitude over which transients were positioned at the time POSS-I transient images were taken.
At least it narrows down a particular date and possible location on a plate(s) from that date for examination.

The new pre-prints may need a new thread.
 
In addition, these transients are supposedly located in particular spots over the Eastern Pacific, off the coast of Mexico:
can I ask the transients on what plate?, how was that plate produced? was it a first generation negative that was used to make a positive and that positive was used to make how many copies? these copies seen to have produced in batches and record keeping of copies seems to have been poor on the whole.

maybe link the paper that you quoted from in this instance-were any of the plates issuses that have been mentioned in the paper or why they might not apply to this plate in particular?
 
But her most recent paper has identified 203 transients on plate(s) from 1 particular day, 5/3/1954... ...In addition, these transients are supposedly located in particular spots over the Eastern Pacific, off the coast of Mexico

Rough map showing the location of the Castle Yankee nuclear test, 05 May 1954, and proposed transient "hotspots"
(Fig. 2, "Earth-Projected Clustering of Historical Optical Transients in the Palomar Observatory Sky Survey-I (POSS-I)", Bruehl, Villarroel et al.)

Not all contributing transients would be on this date, but Bruehl, Villarroel et al. write
External Quote:
Of these, 203 transients (87.9%) were associated with a single nuclear test: Castle Yankee conducted on 5/4/54.
The map is very approximate as the projections used differ. Point C4 in SW Canada is not shown.

gm5.png


Bruehl, Villarroel et al:
External Quote:
It is notable that these 203 transients were all observed the night immediately before this test and were located exclusively over Pacific hotspot regions west of Mexico and Central America identified in hostpot analyses described above. None were observed over the southwestern U.S. or in the Gulf of Mexico region.
If the Oschin Schmidt telescope was not facing East to take photos for POSS1 on 04 May, the lack of transients over the Gulf of Mexico would not be a surprise. Transients will only be apparent in areas of sky being imaged at the time.

I wonder if more transients at lower latitudes (and higher, e.g. C4 in British Columbia) might correlate with positions of apparent transients on the plates? IIRC significantly more transients were detected near plate edges.
 
Last edited:
The geolocation thing, as well as the supposed altitude of transient being calculated from the POSS1 plates is from her new pre-print that I linked to. Whether it works or not, I have no idea. But I would think one could try to follow the logic and find a particular area on the plate that the authors believed was above certain location on the Earth. In this case, a number of points in the Eastern Pacific:
Sure. But if I were an alien in a spacecraft at geostationary altitude where I could see almost the entirety of the Pacific at once I question the meaningfulness of assigning a specific latitude and longitude as being "under" the spacecraft.
 
can I ask the transients on what plate?, how was that plate produced? was it a first generation negative that was used to make a positive and that positive was used to make how many copies? these copies seen to have produced in batches and record keeping of copies seems to have been poor on the whole.

maybe link the paper that you quoted from in this instance-were any of the plates issuses that have been mentioned in the paper or why they might not apply to this plate in particular?

Link to the paper I quoted from is right there under the graph, but I'll link again:

https://arxiv.org/pdf/2609.09461

As to which plate(s) and how they were produced, it's sorta been covered. I think you read the first paper on this subject where they explain they used digitized scans that trace back to the original POSS 1 photographic plates produced at the Palomar Observatory. This new paper uses the same sources.

To be clear, Villarreol et al. don't mention or specify any specific plate in the paper, which is consistent with all of the related papers as far as I know. They used various techniques to analyze digitized versions of the POSS 1 plates and identified "transients" as has been discussed. But it was never clear what transients are found on which digitized plates. That data seems to be unavailable or hard to understand.

As you pointed out up-thread, the digitized versions were often from at least 2nd or more generations of photographic copies, so not only should the originals be examined, but each successive copy between the originals and the digitizing process should also be examined. This becomes quite a daunting task when there is no clear explanation as to which original plates and subsequent copies, supposedly contain which transients. Add to this the delicate nature of the originals which makes them difficult to access.

However, in the new paper a specific date is given for a digitized plate(s) that contains 203 transients related to a nuclear test on May 4, 1954. It appears plates were produced on May 3 and May 5 of 1954. With the test happening on the other side of the date line, the May 3 digitized plate(s) is likely the one.

So, someone now has a specific plate to look at. Even if the original plates are hard to access, there is potentially 1 specific plate(s) to look at. That's 1 plate(s) from 1 night that supposedly contains 203 transients. It narrows down where to look for transients on the originals. And, with one specific plate(s) it might be easier to trace the copies from that plate that led to the digitized versions used in the paper.

In addition, if one can understand their geo-location strategy, regardless if it's wrong or right, it can tell someone where to look for the 203 transients on the 1 specific plate(s). Someone now knows where to look for 203 transients.
 
there is potentially 1 specific plate(s) to look at. That's 1 plate(s) from 1 night that supposedly contains 203 transients.

Probably a modest number of plates.
Each POSS1 plates covered approx. 6° x 6° of sky, but we know there was some overlap at plate edges.

External Quote:
The Samuel Oschin Telescope's corrected field of view measures 6 degrees by 6 degrees, equivalent to about 12 full Moons side by side, which allows it to image vast sky regions efficiently for survey purposes.
StudyGuides.com, Oschin Telescope (Astronomical Telescope), https://studyguides.com/study-methods/study-guide/cml27gd8hb0zs01925v85ttul

External Quote:
The survey utilized 14 inches (36 cm) square photographic plates, covering about 6° of sky per side (approximately 36 square degrees per plate).
Wikipedia, https://en.wikipedia.org/wiki/National_Geographic_Society_–_Palomar_Observatory_Sky_Survey

Exposure times were 45-50 minutes, so not many plates per night (Emergent Mind website, Palomar Observatory Sky Survey (POSS1),
https://www.emergentmind.com/topics/palomar-observatory-sky-survey-poss1).


I think it's interesting that Bruehl, Villarroel et al. (in "Earth-Projected Clustering of Historical Optical Transients in the Palomar Observatory Sky Survey-I (POSS-I)") link their reported transient clusters to a widening category of things:

(1) the Trinity test site: The test was over four years before the first POSS1 exposure.
External Quote:
The nearby White Sands, NM region hotspot (identified via excess analysis) was the location of the first nuclear weapon test (the Trinity test); a transient hotspot in this area might be predicted given observed associations between transients and nuclear testing
-but the association reported by V. et al. is +/- 1 day of a test date, not four years. Admittedly that association was between number of transients regardless of location within a +/- 1 day window; the supposed Trinity cluster is an association with location regardless of time window. There were no other nuclear tests at White Sands IIRC.

(2) The Travis Walton abduction- the same proposed cluster that the authors link to the Trinity test site,
External Quote:
This cluster is centered within Sitgreaves National Forest, the location of the famous Travis Walton UAP sighting and alleged abduction [24].
Travis Walton's abduction narrative is questionable.
The area of Walton's claimed abduction and the Trinity test site are very roughly 400 km apart.

(3) The Chicxulub asteroid impact site,
External Quote:
Hotspots 15 identified in the southern Gulf of Mexico and the Yucatan were noted to be in the same general area as the Chicxulub asteroid impact, a region with documented geomagnetic and gravitational anomalies [20,22]. How such geological features might be linked to groupings of transients in Earth orbit is however unknown.
The Chicxulub site is not a large geomagnetic anomaly. ETI would be aware that planetary bodies get hit by meteorites and sometimes comets or asteroids. Chicxulub is mainly of interest to most of us because of its role in the extinction of the dinosaurs.

Many places might be within 400 km of something of technological interest, a geophysical feature, a military base, nuclear establishment or well-known UFO report.
 
to me, it seems a major oversight to not examine the originals, and if she had tried and was rejected that would be an important point to make in order to show that one is trying in good faith to rule out an instrumental artifact. It seemed like she just said "I don't have access to the originals" and then carried on leaving the uncertainty just hanging over the entire work.
I've never worked with astronomical images, but I have worked with important and valuable documents, and the cameras used for digitisation of same are incredibly high quality (and ludicrously expensive, but I guess in today's money I was working on a project with a hefty 11-digit budget, so 6 digits per camera was small potatoes). Digital copies not accurately representing the originals would be one of the last places I would look for issues. Gels having flaws is an order of magnitude more likely, and if they have them, the photos will show them.
 
I've never worked with astronomical images, but I have worked with important and valuable documents, and the cameras used for digitisation of same are incredibly high quality (and ludicrously expensive, but I guess in today's money I was working on a project with a hefty 11-digit budget, so 6 digits per camera was small potatoes). Digital copies not accurately representing the originals would be one of the last places I would look for issues. Gels having flaws is an order of magnitude more likely, and if they have them, the photos will show them.
the issues is in the making the plates and maybe their storage not the digitization from my understanding. In a very analogy process the negative is used to make a positive that is then used to make plates. These plates seem to have been made in batches and record keeping of their making and distribution seems to have been poor
 
Last edited:
We recently developed a modelling approach that permits estimating the altitude of POSS-I transients based on shadow deficit characteristics [12]. Results using two complementary statistical approaches both suggested that transients were between 26,371 - 35,786 km above Earth.
I'm completely unable to understand this statement; apologies if I'm being dense. These images were taken at Mount Palomar, at 33°21′23″N 116°51′54″W. If the height of each 'transient' is only known with a nine thousand km error bar, then its location above the Earth will also be uncertain with error bars of hundreds, perhaps thousands of kilometres.

Additionally, a satellite with a height above the Earth of 26,371 km would not be in a truly geosynchronous orbit; it would either be travelling too fast to be synchronous, or it would be an eccentric orbit, so that it would only be synchronous for a brief period each orbit. Unless I am misunderstanding this altogether, of course.
 
Link to the paper I quoted from is right there under the graph, but I'll link again:

https://arxiv.org/pdf/2609.09461
This was a great post, Dave, and thanks for taking the time to cover the main points in a way that helps those of us who aren't familiar with the field or paper. Sorry, I didn't realise that was a link to the paper. I'm not familiar with that journal.

I initially misunderstood which earlier dataset was being used here. The paper states that the source is the much larger 107,875-transient dataset derived from Solano et al. (2022) and further detailed in later work. It is not the nine-transient April 1950 case from the 2021 paper. In Methods it states:

"The initial source dataset was a sample of n = 107,875 transients derived by Solano et al. [1] and further detailed in Villarroel et al. [9]. For each analysis described below, we then selected from this full dataset only transients determined in our prior work [10] to have a probability of ≥0.70 (excess analysis) or ≥0.90 (cluster analysis) of being a real transient rather than a plate defect."
The 2021 "nine simultaneously occurring transients" paper is separately listed as reference [3], whereas [1] is Solano et al. (2022). So the 203 Castle Yankee candidates are being selected from this much larger dataset, not from those nine 1950 objects.

So I agree it may be possible to work backwards from the date, POSS-I logs and coordinates to identify the plate(s), but that seems unnecessarily difficult when the analysis itself already knows which plate each transient came from.

I initially thought the "reasonable request" referred to access to physical plates, but on rereading it explicitly says the dataset:

"The final dataset reflecting all analyzed variables will be made available by the authors upon reasonable request to Dr. Beatriz Villarroel."
I'm unsure what the convention is in this field, but as an outsider it seems odd not to provide the plate IDs when plate-level analysis was performed.

I also agree about the assumed altitude. The paper states that "all projected locations depended on the assumed fixed altitude" of 35,786 km. I understand there is another paper Villarroel et al. preprint Modelling Palomar Transients: Constraints from Reflection Geometry and Orbital Altitude (September 2026) that addresses this novel approach of establishing the altitude, but why not also show the projected locations at, say, 20,000, 35,786 and 50,000 km, simply to demonstrate how sensitive the claimed geographical hotspots are to that assumption?

The significance the authors attach to the locations also seems important here. In the Discussion they say:

"Nearly 90% of the highest probability transients associated with nuclear testing (n = 203) were observed exclusively over Pacific region hotspots identified in this work, all immediately prior to the second largest ever US thermonuclear detonation"
But this raises another question for me: why should the phenomenon occur before the second-largest nuclear test rather than the largest, and why before rather than during or after? Their association is defined using a ±1-day window. For Castle Yankee all 203 happened the night before, whereas for the Nevada tests they report transients both on the night of a test and, notably, 23 on the night after one test.

The authors go considerably further when discussing the "before" result:

"one tentative interpretation of our findings could be that transients display characteristics suggesting possible intelligence, that is, a preference for specific locations over the Earth that may fit UAP lore and behavior suggesting possible awareness of imminent nuclear tests."
They immediately acknowledge:

"While this hypothesis may fit available data, it is unfortunately not falsifiable."
That seems important. If the proposed significance of the preceding observations ultimately requires something to somehow anticipate an imminent test, I'd first want to understand whether there is a more ordinary reason for the temporal pattern and why the ±1-day window is the appropriate one.

As an outsider to the field, and someone more interested in what the papers actually establish than in the surrounding drama, I'm finding it increasingly difficult to assess this later work when the fundamental problem of possible plate defects still seems unresolved. More assumptions and unsupported hypotheses are being built on top of detections whose physical origin has yet to be firmly established. I may not look at these papers or anything surrounding them for a long while.
 
As an outsider to the field, and someone more interested in what the papers actually establish than in the surrounding drama, I'm finding it increasingly difficult to assess this later work when the fundamental problem of possible plate defects still seems unresolved. More assumptions and unsupported hypotheses are being built on top of detections whose physical origin has yet to be firmly established. I may not look at these papers or anything surrounding them for a long while.
For whatever reason, these papers have not made much of an impression on the astronomical community, other than the Watters, et. al. critique; most of the citations are their later papers citing their earlier papers.

A lot of the substance is "if X is true, then Y is interesting" and then speculating about Z while undercooking the "if" part of the equation. And the speculations seem more based on received UFO lore than any particular physics or global context; they're reminiscent of how Avi Loeb talked about how the interstellar comet 3I/Atlas could be an alien spaceship, where any sort of sci-fi technology could be invoked.
 
Probably a modest number of plates.

Yes, I was unclear when looking through a list of dates for plate production as to exactly how many plates were made on a giving night. We know there were at least 2, a red sensitive plate and a blue sensitive plate for each session, but the authors, going back to the original paper, only use the red plates. I assumed it was probably more than 1 for each color on each night, but wasn't sure, so I fudged by using the (s) after the word plate.

The main point being, the new paper identifies 203 transients, that they say are 90% likely real objects, from a single night. And again, if one can understand their geo-location system, a location on the plates to look for transients.

(2) The Travis Walton abduction- the same proposed cluster that the authors link to the Trinity test site,

I will confess to missing that one.

Their association is defined using a ±1-day window. For Castle Yankee all 203 happened the night before, whereas for the Nevada tests they report transients both on the night of a test and, notably, 23 on the night after one test.

As @John J. pointed out above, this new paper contains a number of "associations" between the supposed transients and what @jdog referred to as "UFO lore". The original paper associating transients with nuke tests referred to the self published book by Robert Hastings that concerned Robert Salas' claims of nuclear ICBMs going off line at Malmstrom AFB due to the presence of UFOs in 1967. The association is tenuous at best, as the Malmstrom case had nothing to due with nuclear testing and was 10 years out of their self-appointed window. Link to the thread on the very complicated and convoluted Malmstrom event below.

In addition to the association of transients with nuke tests and UFO sightings in the previous paper, this new pre-print (that's why it's on arxiv and not in a journal...yet) goes further and identifies transients in particular geographic locations associated with a number of UFO/paranormal events and areas.

As mentioned above, besides the cluster of transients in the Eastern Pacific associated with nuclear testing ~5000 miles away, there is the problematic Travis Walton abduction case (link to thread below) as well as UFO/New Age/paranormal hotspot Sedona AZ and the Gulf of Mexico near Tampico:

External Quote:

Cluster 2, reflecting nearly a quarter of high probability transients, was centeredin the Gulf of Mexico just off the coast of the Tampico-Vera Cruz area. Cluster 3 was locatedj ust southeast of Sedona, AZ, in the Sitgreaves National Forest.
External Quote:

The Gulf of Mexico hotspot identified using both methods offshore from the Tampico-Vera Cruz region is an area with so many UAP sightings that many local residents believing there is an underwater "UFO base" offshore [28,29].
And a cluster kinda off the coast of Baja the authors suggest is associated with the 2004 Nimitz encounter with the now famous TicTac UFO:

External Quote:

Finally, the hot spot closest to Baja California (E9, identified only using the excess method) is just south of the southern boundary of the US Navy W-291 training range where the famous"tic-tac" UAP incident occurred [30]. It is unknown whether the UAP-related location profiles above associated with these hotspots are meaningful or simply coincidental.
While the authors say these associations may be coincidental, they make a point of talking about them. The transient cluster from 1954 identified as E9 that might be associated with the 2004 Nimitz event is maybe ~1000 miles south and 50 years later from the area the event happened (red circle):

Screenshot 2026-09-19 8.23.48 AM.png


Seems a bit of a stretch.

Travis Walton thread that also links to @Charlie Wiser blog and all of her extensive research:

https://www.metabunk.org/threads/travis-walton-case-crew-boss-confesses-hoax.11878/

Long thread with lots of resources about the confusing Malmstrom AFB story:

https://www.metabunk.org/threads/uf...mstrom-eagle-flight-skeptical-resources.3284/
 
I have some critiques of Villarroel's paper Use of machine learning to enhance detection of transient astronomical phenomena in historical observatory images. (For convenience, I will refer to Dr. Villarroel as BV, the convention she uses in the paper.)

I've said this many times but I'll repeat it again, I am NOT accusing BV of deliberate fraud or deception. Every critique I am making here can be the result of honest mistakes and/or subconscious biases. I am not making any claims on motivation.

The first critique actually applies to most of her papers - a very high rate of self-citation. For example, 7 of the first 10 citations of this machine learning paper are to her previous work. Earth-Projected Clustering of Historical Optical Transients in the Palomar Observatory Sky Survey-I (POSS-I) has 12 self-citations in the first 15 sources. This is not inherently a bad thing and is not too surprising given the niche nature of the topic. However, the lack of replication by other, independent researchers and the extreme reliance on her own work creates the risk of mistakes and biases propagating throughout her research.

Back to the ML paper specifically. I am a data scientist/actuary so I know the models used extremely well. Their training set was 250 samples, which is quite small. Only 250 samples when using 23 features makes that sample size even more problematic. That's a small number of samples in general and very small relative to the number of features. The next issue is the target, i.e. what the model is trying to predict. The model is a probabilistic classification model, meaning it outputs the probability that a transient is "real". How were samples determined to be real or defects?
This training set was manually inspected by an astronomer with expertise in the transient phenomenon (BV); 134 were labelled likely real transients and 116 as plate defects.
Dr. Villarroel herself decided! But there was a second reviewer...
To enhance classification consistency over time, a second reviewer (SB) who was trained by the primary evaluator periodically reviewed a subset of the images assigned to the primary reviewer, withany disagreements discussed with the primary reviewer and resolved.
The second reviewer was trained by Dr. Villarroel herself! Any disagreement was discussed with BV. Why does this matter? The goal of this paper was to
External Quote:
enhance transient identification accuracy and validate the phenomenon.
But since BV determined the training data, this model isn't enhancing accuracy or validating the phenomenon. It is a model that just predicts BV's opinion. This exercise can be summarized as "the model I built on data I classified based on my own assumptions validates my other set of claims that are also based on the same assumptions." Since the original analysis and the ML analysis are all predicated on the same assumptions, none of the potential sources of bias are actually addressed. All she's really testing is whether the data used to build the ML model is similar to the data used for the transient analysis, which is of course true. It comes from the same source!

The modeling procedure and its description is quite odd. Saying "model" is somewhat inaccurate since BV actually trained 4 models and a final meta model (aka ensemble models, model stacking, etc). She trained 4 models and then the final prediction is just the average of those 4. That makes the 250 sample size even more insufficient. It's also just completely unnecessary and leads to a mountain of possible issues, such as overfitting, despite the use of cross-validation. The modeling procedure also suggests a lack of experience and expertise on machine learning. See for example:

External Quote:
The ensemble ML classifier detailed in this study combined four tree-based models(XGBoost, Random Forest, Gradient Boosting, LightGBM), each trained with 300 trees and identical hyperparameters, with final classification predictions based on the unweighted mean ofthe four models' predicted probabilities.
The above sentence doesn't actually make sense to someone who uses these models (such as me). She says the four models are XGBoost, Random Forest, Gradient Boosting, and LightGBM. XGBoost and LightGBM are both gradient boosting decision tree models (hence the GB in both names). It's not even clear what the third model is referring to. "Gradient boosting" is a specific modeling technique, not a particular model itself. It might be the implementation of gradient boosted trees from the python package she used (the paper says sci-kit learn which includes GradientBoostingClassifier). Random Forest, while not gradient boosted, is another type of ensemble of decision trees. Which random forest implementation was used was also not specified. Again, I'll just assume sci-kit learn RandomForestClassifier.

So 3 of the four models are just different implementations of the same model type and the other model is extremely similar. There is no prima facie reason to use an ensemble of four extremely similar models. It actually overcomplicates the analysis, significantly increasing variance while likely having no significant decrease in model bias (see Bias–variance tradeoff. Model stacking is a useful technique when the models aren't highly correlated. Think of a case where you have 2 models that give the exact same predictions. Averaging them gives you no benefit. Model stacking is useful when you have say, model 1 that performs on some subset of data and poorly on the rest, and another model 2 that does the opposite. Stacking those models is useful because the meta model will learn when to use model 1 and when to use model 2.

Saying all 4 were trained with "identical hyperparameters" is also simply impossible. "Hyperparameters" are essentially the settings determining how the model will learn from the data. Despite the similarity between XGBoost, LightGBM, and Gradient Boosting, each of these implementations have different hyperparameters. You can view their hyperparameters at the following links: XGBoost, LightGBM, GradientBoostingClassifier, RandomForestClassifier. Not just different default values for hyperparameters, they all have some hyperparameters that are completely unique from each other, due to the difference in how they build their trees. Random Forest is a different algorithm altogether so it simply cannot have identical hyperparameters to the others. XGBoost and LightGBM can be explained simply as combining the predictions of a bunch of decision trees. However, XGBoost builds its trees "depth-wise" and LightGBM builds its trees "leaf-wise". The meaning of those terms is out of scope here, but the point is that since they build their trees using different algorithms, it is not possible to have identical hyperparameters.

The features themselves used in the model also may not be what many would have expected. My immediate assumption when seeing the title of the paper was that BV used a computer vision model of some sort. That is, a model that actually "looks" at the images on the plate (using the pixel values). The most common CV model type has been convolutional neural networks (or at least traditionally most common until transformers arrived). These models actually look at the images by taking their pixel values as inputs. From there, the basic explanation is they learn relationships and structures of things in the image. BV's models took a different approach:
External Quote:
The ML model included 23 predictors extracted from the red FITS images and the VASCO v4 catalog. Seven catalog-level morphometric features were included: signal-to-noise ratio (SNR), point spread function (PSF) ratio, elongation, compactness, sharpness, number of comparison stars, and candidate score (described in Solano et al.1). The ML model also included 6 plate-aggregate features: a plate quality indicator, plate-level SNR fraction, SNR standard deviation, mean elongation, and high-SNR source count. Finally, the model included 10 morphometric features identified by the ML model in the red FITS images themselves: PSF Full Width at Half Maximum (PSF FWHM), ellipticity, sharpness, connected pixel count, aperture flux, distance to plate edge, symmetry score, gradient magnitude, proximity to bright star, and FITS-measured SNR. The ML model did not include any spectral features or red-blue plate comparison.
Fewer than half of the features actually relate to any given transient itself. 13 of the 23 features are either catalog-level (7) or plate-level (6).

We can also take a look at Figure 2 to see how these features contribute to the overall prediction. The figure visualizes SHAP values, which are a measure of how a feature contributed to the overall prediction (the name deriving from Game Theory's Shapley Values). We can see which features were the most influential, on average. Here are the top 6 features (some formatting added in by me):
External Quote:

  1. plate_low_snr_frac = Plate- level: fraction of candidates on the same plate with SNR < 5;
  2. plate_quality_GOOD = Plate-level: binary, plate rated as GOOD (one-hot encoded);
  3. plate_quality_MODERATE = Plate-level: binary, plate rated as MODERATE (one-hot encoded);
  4. plate_elongation_mean = Plate-level: mean elongation of all candidates on the same plate;
  5. plate_snr_std = Plate-level: standard deviation of SNR across all candidates on the same plate;
  6. candidate_score = Candidate-level: original pipeline quality score from Solano et al.
The top 5 features are all plate-level. The 6th most important feature is "candidate score", which is a value derived in another paper authored by, you guessed it, BV.
External Quote:
Solano, E., Villarroel, B., & Rodrigo, C. Discovering vanishing objects in POSS I red images using the Virtual Observatory, Monthly Notices Royal Astron Soc. 515, 1380–1391 (2022).
This is just another source of bias being introduced into the model. The most important features being plate-level means that very little of a prediction for a given transient is actually a characteristic of that specific transiet.

Lastly, one question I have is "why not start with a simple logistic regression generalized linear model?" It is best practice in ML to start with the simplest models first (kind of analogous to Occam's razor). Basically, start simple, and then look for evidence of interactions and non-linear effects that would be handled by more complex models. Getting even more specific, catalogs and plates should likely be treated as "random effects" because multiple transients on a plate violate the independence assumption of most statistical analyses. That topic is again out of scope for this comment.

Ok, so what's the point of all of the above? Returning to the abstract, the goal of the paper was improving identification and validating the phenomenon.
External Quote:
These findings remain debated; some argue transients identified via existing automated pipelines are simply plate defects. Machine learning (ML) was used to enhance transient identification accuracy and validate the phenomenon.
This ML paper does neither. Due to the fact that every step of the ML process was determined by BV, the ML approach hasn't done anything to actually remove that bias. The common saying is "garbage in, garbage out", but I don't consider the work to be "garbage", so a more fair saying would be "biased or invalid data in, biased or invalid data out".

There's quite a lot of jargon in this comment, and I tried to balance providing sufficient detail without being too dense. Please feel free to ask for any clarifications!

One last thing I forgot to mention. Despite the paper referring to a GitHub repo for the code, it doesn't appear to exist, so all of the above is based off the paper itself. I would have preferred the code itself, but I'm working with what I can.
 
Last edited:
The original paper associating transients with nuke tests referred to the self published book by Robert Hastings that concerned Robert Salas' claims of nuclear ICBMs going off line at Malmstrom AFB due to the presence of UFOs in 1967
what a great post again very informative
next time maybe lead with the Robert Hastings and others connections. That would of got we where I needed to be with this collection of work a lot sooner :p

thankfully I only read those 3 papers and someone else may find your post informative also-10x good work
 
Back
Top